AgentCoreMemorySessionManager with batch_size > 1 restores agents without the changes made in their previous invocation
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
调研方向
从 src/bedrock_agentcore/memory/integrations/strands/session_manager.py 开始:read_agent (~L474) 取最新 AGENT 事件的 payload[0],而 flush (L1137-L1182) 按最早优先的顺序追加 SessionAgents,因此最后一次保存永远不会被读取。粘贴 issue 中的复现脚本(in-memory data plane,batch_size=10)以复现 model_calls 卡在 1 的情况,然后修改索引并确认恢复后的 agent 读取到 3。完成的定义是:该复现在 state、sliding-window history、interrupt resume 和 overflow 上全部通过,并在现有 session_manager 测试旁边补充一个回归测试。
由索引模型根据 Issue 内容生成。
描述
With batch_size > 1, an agent restored by AgentCoreMemorySessionManager does not get the changes made in its previous invocation. This happens when the app builds a new Agent and session manager for each request, the usual way to serve many sessions since session_id is fixed in AgentCoreMemoryConfig. The restore raises and logs nothing: the agent just starts from an older state.
What breaks, with batch_size=10:
- Values written to
agent.stateduring an invocation are missing in the next one. In the repro below, a counter kept inagent.statenever goes past 1. - Since 1.23.1, Strands' default conversation manager,
SlidingWindowConversationManager, stops limiting the history across requests. Each restore comes back withremoved_message_count=0and every message of the session, so each request sends the whole conversation to the model. Withwindow_size=4, eight requests sent 1, 3, 5, 7, 9, 11, 13 and 15 messages. On 1.20.0 they sent 1, 3, 5, 5, 5, 5, 5 and 5. - Since 1.23.1, an interrupt cannot be resumed from a new agent: the call with the
interruptResponseraisesValueError: Received interrupt responses but agent is not in interrupt state. On 1.20.0, or withbatch_size=1, it resumes. - After a context-window overflow, Strands removes the oldest messages, but the next restore brings them back. This one also happens on 1.20.0.
Cause (links to v1.24.0):
Strands saves the agent as a SessionAgent, which holds agent.state, the conversation manager state and the interrupt state. It calls sync_agent after each new message and once more at the end of the invocation. With batch_size > 1, each SessionAgent is appended to a buffer (L400). The buffer is flushed as one AGENT event whose payload holds those SessionAgents, oldest first (L1137-L1182). On restore, read_agent takes the newest AGENT event and reads payload[0] (L474), the oldest SessionAgent in it.
The first sync_agent after a restore always writes, so each invocation's AGENT event starts with the SessionAgent the agent was restored from, and the changes saved after it are never read back.
Up to 1.23.0, the SessionAgent saved at the end of the invocation stayed in the buffer until the next flush, so it started a later event and was restored. Since 1.23.1 (#664), the flush at AfterInvocationEvent runs after that last sync_agent (L948-L954), so the end-of-invocation SessionAgent lands in the same event and is lost too. That is why the sliding-window and interrupt cases only fail since 1.23.1.
Repro, with no AWS account: the data plane is an in-memory fake that returns events newest first, which is what read_agent expects from ListEvents. Each turn builds a new agent on the same session, as a server does for each request, and a hook counts model calls in agent.state.
import copy, os
os.environ.update(AWS_ACCESS_KEY_ID="x", AWS_SECRET_ACCESS_KEY="x", AWS_DEFAULT_REGION="eu-west-1", AWS_ENDPOINT_URL="http://127.0.0.1:9")
from bedrock_agentcore.memory.integrations.strands.config import AgentCoreMemoryConfig
from bedrock_agentcore.memory.integrations.strands.session_manager import AgentCoreMemorySessionManager
from strands import Agent
from strands.hooks import BeforeModelCallEvent
from strands.models.model import Model
class InMemoryDataPlane:
def __init__(self):
self.events = []
def create_event(self, **params):
self.events.append(copy.deepcopy(params))
return {"event": {"eventId": f"e{len(self.events)}"}}
def list_events(self, **params):
def matches(event):
metadata = event.get("metadata") or {}
return all(
metadata.get(f["left"]["metadataKey"], {}).get("stringValue") == f["right"]["metadataValue"]["stringValue"]
for f in params.get("filter", {}).get("eventMetadata", [])
)
return {"events": [e for e in reversed(self.events) if matches(e)][: params["maxResults"]]}
class Session:
region_name = "eu-west-1"
def __init__(self, plane):
self.plane = plane
def client(self, name, **kwargs):
return self.plane
class EchoModel(Model):
def update_config(self, **kwargs): pass
def get_config(self): return {}
def structured_output(self, *args, **kwargs): raise NotImplementedError
async def stream(self, messages, *args, **kwargs):
yield {"messageStart": {"role": "assistant"}}
yield {"contentBlockDelta": {"contentBlockIndex": 0, "delta": {"text": "ok"}}}
yield {"contentBlockStop": {"contentBlockIndex": 0}}
yield {"messageStop": {"stopReason": "end_turn"}}
def count_model_calls(event):
event.agent.state.set("model_calls", (event.agent.state.get("model_calls") or 0) + 1)
def new_agent(plane):
manager = AgentCoreMemorySessionManager(
AgentCoreMemoryConfig(memory_id="m", session_id="s", actor_id="a", batch_size=10),
region_name="eu-west-1",
boto_session=Session(plane),
)
agent = Agent(model=EchoModel(), session_manager=manager, callback_handler=None)
agent.hooks.add_callback(BeforeModelCallEvent, count_model_calls)
return agent, manager
plane = InMemoryDataPlane()
for turn in (1, 2, 3):
agent, manager = new_agent(plane)
with manager:
agent(f"question {turn}")
print(f"after turn {turn}: model_calls={agent.state.get('model_calls')}") # expected 1, 2, 3
restored, _ = new_agent(plane)
print("restored:", restored.state.get("model_calls")) # expected 3
Expected: model_calls is 1, 2 and 3 after the three turns, and the restored agent reads 3, which is what batch_size=1 gives. Actual, on every version below: model_calls=1 after each turn, then restored: None. The newest AGENT event holds two SessionAgents whose state is {} and {"model_calls": 1}, and read_agent returns the first.
Suggested fix: read_agent reads payload[-1] instead of payload[0]. With that one-line change, the repro restores 3, and the sliding-window, interrupt and overflow cases behave as with batch_size=1. Sessions stored by affected versions then restore the last SessionAgent they saved, which keeping only the latest SessionAgent in the buffer would not do.
Versions tested: bedrock-agentcore 1.20.0 with strands-agents 1.45.0, 1.23.1 with 1.57.2, and 1.24.0 with 1.57.2 and 1.58.0, on Python 3.13. main has the same code as 1.24.0.
- 主要语言
- Python
- 星标
- 776
- 派生
- 153
- 平均合并
- 1 天 8 小时
- 30 天内合并 PR
- 15
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
aws/bedrock-agentcore-sdk-python 的其他 Issue
-
AgentCoreMemorySessionManager discards the two boto3 clients MemoryClient builds; boto_session and boto_client_config are not passed through可能已有人在做 @avneetbansal-aws 于 10 天前认领。 未关闭bug
难度 2/5 1-3 小时 新手友好度 84/100
aws/bedrock-agentcore-sdk-python#681 · 1 条评论 ·
维护者通常 1 天内回复
-
bug high-severity
难度 2/5 1-3 小时 新手友好度 70/100
aws/bedrock-agentcore-sdk-python#680 ·
维护者通常 1 天内回复
-
[Bug] update_message fails with parameter validation error when SessionMessage.message_id is a Strands positional integer index可能已有人在做 @citizen204 于 100 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 68/100
aws/bedrock-agentcore-sdk-python#556 ·
维护者通常 1 天内回复
-
CodeInterpreter should not require a region argument可能已有人在做 @citizen204 于 115 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 68/100
aws/bedrock-agentcore-sdk-python#511 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
aws/bedrock-agentcore-sdk-python#496 · 1 条评论 ·
维护者通常 1 天内回复
查看 aws/bedrock-agentcore-sdk-python 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
SR_SECURITY_DESCRIPTOR.fromString drops the SACL when no DACL is present可能已有人在做 @paul7436 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 85/100
equinor/fmu-sumo-uploader#302 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 75/100
modelscope/evalscope#1821 ·
维护者通常 1 天内回复
-
Sanity on ansible-core devel fails: ignore-2.23.txt references the removed import-3.9 test可能已有人在做 @yurnov 今天认领。 未关闭needs_triage
难度 1/5 1 小时以内 新手友好度 91/100
ansible-collections/kubernetes.core#1275 ·
维护者通常 1 天内回复