Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 35/100
- Issue 类型
- 缺陷
- 描述清晰度
- 需要澄清
- 活跃度
- 活跃
- 技术栈
- python
- 领域
- backend, networking
调研方向
首先跟踪 run_message_listeners 和 SocketModeRequest.from_dict,重点关注在消息监听器接收帧之前,hello 元数据和断开连接原因是如何处理的。如果可能,复现重新连接行为,然后定义诊断信息,清晰地显示服务器端连接数和原因,以便对不健康的连接池发出警告。
由索引模型根据 Issue 内容生成。
描述
Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT
Summary
This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.
Wire evidence
(from a run_message_listeners wrapper logging hello and disconnect frames)
- fresh start:
helloreportsnum_connections=1 - ~1.5 h later, on a reconnect: 3x
disconnect reason=too_many_websockets,
thenhello num_connections=10— while the process verifiably held ONE
established TCP connection to Slack the whole time - ghost registrations age out at roughly one per 30-45 minutes
- while
num_connectionsis high, most button clicks never arrive on any
connection we hold; with a clean pool, every click arrives (tested across
message sizes 0.5-5 KB — size is irrelevant) - reproduced on a SECOND app in the same workspace: first
helloafter a
process restart reportednum_connections=7for an app that also runs as a
single instance - observed
approximate_connection_time(insidehello.debug_info) is
consistently18060(~5 h), which sets the ghost age-out horizon
Why this is hard to see with the current SDK
- The
helloenvelope (carryingnum_connections) never reaches
message_listeners—SocketModeRequest.from_dictrequires
type+envelope_id+payload, so apps cannot observe the most important signal
without wrapping internals. disconnectframes (includingtoo_many_websockets) are handled by
run_message_listenersbefore the listener loop and only visible at debug
logging.- A degraded connection still passes
is_connected()/ ping-pong checks, so
client-side health monitoring cannot detect the server-side pool state.
Proposals (any subset would help)
- Surface
hellometadata (num_connections,approximate_connection_time,
host) anddisconnectreasons via a public callback or at INFO logging. - Emit a loud warning when
num_connectionsinhelloexceeds a threshold
(e.g., 4) while the client manages fewer connections — this is direct
evidence of ghost registrations and imminent interactive-payload loss. - Consider make-before-break reconnects with close-confirmation, or documenting
that reconnect-heavy operation behind NAT can poison the server-side pool. - Still-open PRs #1914 / #1926 address orphaned client-side sessions in the
same failure family — this issue is their server-side counterpart, and since
neither is merged/released (latest release is 3.43.0), apps currently have no
upstream remedy at all; that raises the priority of surfacing the diagnostics
from proposals 1-2.
Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.
- 主要语言
- Python
- 星标
- 4k
- 派生
- 857
- 平均合并
- 22 小时 21 分钟
- 30 天内合并 PR
- 16
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
slackapi/python-slack-sdk 的其他 Issue
-
needs info server-side-issue
难度 4/5 3-5 天 新手友好度 35/100
slackapi/python-slack-sdk#1961 · 3 条评论 ·
-
Use logger.isEnabledFor(logging.DEBUG) instead of logger.level <= logging.DEBUG for debug guards 未关闭auto-triage-skip bug
难度 4/5 3-5 天 新手友好度 55/100
slackapi/python-slack-sdk#1957 ·
-
chat_postMessage silently forwards thread_id to the API, so a threaded reply posts to the channel 未关闭auto-triage-skip enhancement
难度 4/5 3-5 天 新手友好度 48/100
slackapi/python-slack-sdk#1923 · 2 条评论 ·
-
auto-triage-skip bug socket-mode
难度 3/5 1-2 天 新手友好度 72/100
slackapi/python-slack-sdk#1922 · 2 条评论 ·
-
auto-triage-skip bug python web-client
难度 3/5 1-2 天 新手友好度 52/100
slackapi/python-slack-sdk#1853 · 2 条评论 ·
查看 slackapi/python-slack-sdk 的全部 Issue
相似的 Issue
-
area: harness bug status: needs-triage
难度 2/5 1-3 小时 新手友好度 75/100
Human-Agent-Society/reef#625 ·
-
难度 2/5 1-3 小时 新手友好度 70/100
-
难度 1/5 1 小时以内 新手友好度 80/100
learningequality/kolibri#15351 · 2 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
-
Name consistency 未关闭
难度 2/5 1-3 小时 新手友好度 75/100
eellak/triplestore#65 · 1 条评论 ·