Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals

未关闭
#1,940 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
需要澄清
活跃度
活跃
技术栈
python

调研方向

首先跟踪 run_message_listeners 和 SocketModeRequest.from_dict,重点关注在消息监听器接收帧之前,hello 元数据和断开连接原因是如何处理的。如果可能,复现重新连接行为,然后定义诊断信息,清晰地显示服务器端连接数和原因,以便对不健康的连接池发出警告。

由索引模型根据 Issue 内容生成。

描述

auto-triage-skip discussion

Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT

Summary

This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.

Wire evidence

(from a run_message_listeners wrapper logging hello and disconnect frames)

  • fresh start: hello reports num_connections=1
  • ~1.5 h later, on a reconnect: 3x disconnect reason=too_many_websockets,
    then hello num_connections=10 — while the process verifiably held ONE
    established TCP connection to Slack the whole time
  • ghost registrations age out at roughly one per 30-45 minutes
  • while num_connections is high, most button clicks never arrive on any
    connection we hold; with a clean pool, every click arrives (tested across
    message sizes 0.5-5 KB — size is irrelevant)
  • reproduced on a SECOND app in the same workspace: first hello after a
    process restart reported num_connections=7 for an app that also runs as a
    single instance
  • observed approximate_connection_time (inside hello.debug_info) is
    consistently 18060 (~5 h), which sets the ghost age-out horizon

Why this is hard to see with the current SDK

  1. The hello envelope (carrying num_connections) never reaches
    message_listenersSocketModeRequest.from_dict requires
    type+envelope_id+payload, so apps cannot observe the most important signal
    without wrapping internals.
  2. disconnect frames (including too_many_websockets) are handled by
    run_message_listeners before the listener loop and only visible at debug
    logging.
  3. A degraded connection still passes is_connected() / ping-pong checks, so
    client-side health monitoring cannot detect the server-side pool state.

Proposals (any subset would help)

  1. Surface hello metadata (num_connections, approximate_connection_time,
    host) and disconnect reasons via a public callback or at INFO logging.
  2. Emit a loud warning when num_connections in hello exceeds a threshold
    (e.g., 4) while the client manages fewer connections — this is direct
    evidence of ghost registrations and imminent interactive-payload loss.
  3. Consider make-before-break reconnects with close-confirmation, or documenting
    that reconnect-heavy operation behind NAT can poison the server-side pool.
  4. Still-open PRs #1914 / #1926 address orphaned client-side sessions in the
    same failure family — this issue is their server-side counterpart, and since
    neither is merged/released (latest release is 3.43.0), apps currently have no
    upstream remedy at all; that raises the priority of surfacing the diagnostics
    from proposals 1-2.

Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.

主要语言
Python
星标
4k
派生
857
平均合并
22 小时 21 分钟
30 天内合并 PR
16

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

slackapi/python-slack-sdk 的其他 Issue

查看 slackapi/python-slack-sdk 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。