SocketModeClient.connect() retries forever against a permanently closed aiohttp ClientSession
还没有人认领这个 Issue。
评估
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 新手友好度
- 72/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
- 技术栈
- python
- 领域
- backend, networking
调研方向
从 slack_sdk/socket_mode/aiohttp/init.py 中的会话创建部分(~L130)、is_connected()(~L321-334)、connect()(~L347-409)和 close()(~L446-457)开始。重现会话被强制关闭的场景并跟踪重试循环;完成的标准是客户端不再针对已关闭的会话无限重试,并且健康检查能够反映失败,或者调用方收到明确的终止信号。
由索引模型根据 Issue 内容生成。
描述
Title
SocketModeClient.connect() retries forever against a permanently closed aiohttp ClientSession
Environment
slack_sdk3.43.0- File:
slack_sdk/socket_mode/aiohttp/__init__.py - Relevant locations: session creation ~L130,
is_connected()~L321-334,connect()~L347-409,close()~L446-457
Description
SocketModeClient (aiohttp implementation) keeps a single aiohttp.ClientSession for the client's entire lifetime — a reasonable design, per the comment at its creation: "it is suggested you use a single session for the lifetime of your application, to benefit from connection pooling." The bug isn't that design choice; it's that the unbounded retry loop doesn't handle that session ever entering a closed state.
connect() wraps reconnection attempts in while True: (~L352). On exception it logs "Failed to connect (error: {e}); Retrying..." (~L408) and loops again, reusing self.aiohttp_client_session. If that session itself has been closed (not just the individual WebSocket tracked as self.current_session), every subsequent retry fails identically forever with RuntimeError: Session is closed — nothing inside this loop ever recreates the session; that only happens when a new SocketModeClient instance is constructed from scratch.
Separately, is_connected() checks self.current_session/ping-pong state but does not check self.aiohttp_client_session.closed, so downstream consumers building their own health checks or watchdogs on top of this client have no way to detect this specific failure mode without reaching into a private-ish attribute themselves.
Impact observed
In production, a consumer application's own reconnect watchdog (checking is_connected()) never detected this state, because the WebSocket layer could still look "connected enough" (some event types were still being delivered) while the HTTP session underneath was permanently dead. The client was stuck retrying every ~10 seconds for over 24 hours with no self-healing, until the whole process was restarted externally.
Likely forced repro (not yet reduced to a minimal script)
- Construct a
SocketModeClientwith a valid app token;await client.connect(). await client.aiohttp_client_session.close()(or otherwise force it closed) while the connection is established.- Trigger a reconnect attempt (e.g. disconnect the network, or otherwise cause
connect()'s loop to retry). - Observe it fails forever:
Failed to connect (error: Session is closed); Retrying... - Observe
is_connected()may not reflect the problem ifself.current_session(the WebSocket) hasn't itself been marked closed/None.
Suggested fixes (ranked)
- Preferred: in
connect()'s retry loop, checkself.aiohttp_client_session.closedat the top of each iteration; if closed, recreate it (e.g. re-instantiateaiohttp.ClientSession) before retryingws_connect. - Minimum: if recreating isn't desired,
raiseorbreakout of the loop when the session is closed, so the caller can rebuild the whole client instead of retrying forever against a dead one. - Consumer-facing improvement: reflect
aiohttp_client_session.closedinis_connected()(or an equivalent health-check property), so consumers can detect this without reaching into a non-public attribute.
Notes
We don't have a minimal standalone repro yet — this was diagnosed from production logs plus reading this source after a ~39 hour incident where a single client instance got stuck in this state. Happy to help characterize a repro further if useful. We've worked around this downstream by having our own consumer explicitly check aiohttp_client_session.closed before deciding whether to rebuild — happy to link that once it's merged, in case the pattern is useful context here too.
- 主要语言
- Python
- 星标
- 4k
- 派生
- 857
- 平均合并
- 22 小时 21 分钟
- 30 天内合并 PR
- 16
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
slackapi/python-slack-sdk 的其他 Issue
-
needs info server-side-issue
难度 4/5 3-5 天 新手友好度 35/100
slackapi/python-slack-sdk#1961 · 3 条评论 ·
-
Use logger.isEnabledFor(logging.DEBUG) instead of logger.level <= logging.DEBUG for debug guards 未关闭auto-triage-skip bug
难度 4/5 3-5 天 新手友好度 55/100
slackapi/python-slack-sdk#1957 ·
-
auto-triage-skip discussion
难度 5/5 一周以上 新手友好度 35/100
slackapi/python-slack-sdk#1940 · 2 条评论 ·
-
chat_postMessage silently forwards thread_id to the API, so a threaded reply posts to the channel 未关闭auto-triage-skip enhancement
难度 4/5 3-5 天 新手友好度 48/100
slackapi/python-slack-sdk#1923 · 2 条评论 ·
-
auto-triage-skip bug python web-client
难度 3/5 1-2 天 新手友好度 52/100
slackapi/python-slack-sdk#1853 · 2 条评论 ·
查看 slackapi/python-slack-sdk 的全部 Issue
相似的 Issue
-
area: harness bug status: needs-triage
难度 2/5 1-3 小时 新手友好度 75/100
Human-Agent-Society/reef#625 ·
-
难度 2/5 1-3 小时 新手友好度 70/100
-
难度 1/5 1 小时以内 新手友好度 80/100
learningequality/kolibri#15351 · 2 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
-
Name consistency 未关闭
难度 2/5 1-3 小时 新手友好度 75/100
eellak/triplestore#65 · 1 条评论 ·