bug(router): relay circuit limits are hardcoded to go-libp2p defaults, and a relay reset surfaces as a bare 502
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 70/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- go
- Domain
- backend, networking
Research direction
Start by reading the router setup in internal/router/router.go:410 and the relay resource limits in go-libp2p's p2p/protocol/circuitv2/relay/resources.go:63-66. Then examine the egress proxy in internal/node/sidecar.go:690 to see where network.ErrReset is handled. The fix involves adding configurable limits to the relay setup and an ErrorHandler to the reverse proxy. Test by setting up a relayed connection and triggering the limits as described in the repro steps.
Written by the indexing model from the issue text.
Description
Problem
Nodes behind NAT reach each other through circuit-relay-v2 on the router. The
router sets up the relay with go-libp2p's default resources
(relay.New(hostNode, relay.WithACL(...)), internal/router/router.go:410),
so every relayed connection is reset after 2 minutes or 128 KiB per
direction, whichever comes first (go-libp2p v0.49.0,
p2p/protocol/circuitv2/relay/resources.go:63-66). Both limits count over the
lifetime of the relayed connection, not per request.
Nothing in SAM exposes these limits, so on a mesh where peers only talk
through the relay (every node behind NAT, e.g. a sam-one on Cloud Run with
laptop, phone and VM nodes):
- Any request that stays open longer than 2 minutes dies mid-flight. A
blocking A2ASendMessageto an agent that runs an LLM and delegates to
another agent routinely exceeds this. - Any exchange that moves more than 128 KiB over one relayed connection dies,
e.g. file artifacts or a long MCP session.
What the caller sees
The egress proxy (createEgressProxy, internal/node/sidecar.go:690) is an
httputil.ReverseProxy with no ErrorHandler. When the relay resets the
stream, the transport returns network.ErrReset and the caller gets a plain
502 Bad Gateway with an empty body. The callee keeps working: the remote
agent carries on, delegates to a third agent and finishes, while the caller
has already reported failure. Nothing in the 502 points at the
relay, so the natural reading ("the remote service crashed") is wrong.
Repro
- Run the router (or
sam-one) with two nodes that cannot connect directly,
so traffic between them is relayed. - Register a service on node B whose backend sleeps 150 s before responding.
- Call it from node A through the sidecar:
curl -H 'X-Sam-Authentication: Bearer <token>' http://127.0.0.1:<port>/sam/<peer-B>/<type>/<name>/ - After ~120 s the call returns
502, while B's backend is still running.
Proposal
- Configurable relay limits on
sam-routerandsam-one: a duration
and a data limit for relayed connections, passed to
relay.WithResources. Keep the go-libp2p defaults when unset. Allow0
to mean "no limit" (a nilResources.Limitmakes the relay unlimited,
relay.go:473). Operators of a relay-only mesh can then size it for
their workloads. - Explicit error on relay reset: give the egress proxy an
ErrorHandlerthat recognisesnetwork.ErrReseton a limited (relayed)
connection and answers with a body that names the cause, e.g.
502 relayed connection to <peer> was reset by the relay (limit: 2m / 128KiB),
and logs it at warn level on the caller's node. The status can stay 502;
the point is a body and a log line that name the relay limit.
Out of scope: clients that avoid long-lived requests (A2A ReturnImmediately
plus polling GetTask) already survive the resets; that is a per-client
choice and does not remove the 128 KiB cap.
- Dominant language
- Go
- Stars
- 926
- Forks
- 138
- Avg merge
- 11h 27m
- Merged PRs (30d)
- 146
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from google/sam
-
debug network-info: observed_addresses reports announced addresses, not observed onesPossibly taken @Mukezh claimed this 5 days ago. Opengood first issue
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
google/sam#484 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
-
Observability storyOpen
Difficulty 5/5 Over a week Newbie friendliness 25/100
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
Maintainers usually reply within 1 day
-
Can we replace the MCP to A2A with the official CLI https://github.com/a2aproject/a2a-cli ?Possibly taken @kaisoz claimed this 20 days ago. Open
google/sam#378 · 3 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 55/100
Maintainers usually reply within 1 day
Similar issues
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 92/100
open-telemetry/opentelemetry-go-compile-instrumentation#1417 ·
Maintainers usually reply within 2 days
-
agent-research-finding agent-research-recommend chore ready-for-agent
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
jordansmall/spindrift#4068 · 1 comment ·
Maintainers usually reply within 1 day
-
Type/Task
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
OpenNSW/nsw-srilanka#537 ·
Maintainers usually reply within 1 day
-
security
Difficulty 2/5 1-2 days Newbie friendliness 62/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
Maintainers usually reply within 1 day