feat(routing): share session affinity across gateway replicas
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- redis, rust
Research direction
Start by reading Praxis's request-selection and current in-memory affinity paths, then inspect Grid's routing overlay and Helm/Forge-backed qualification setup. Define the shared-store binding and failure semantics before implementation; done means deterministic tests cover replica sharing, eligibility changes, expiry, outage and recovery, plus Grid qualification evidence that does not expose session values.
Written by the indexing model from the issue text.
Description
Problem
Praxis can keep a session on the same provider by reading a configured request header or cookie and storing a provider binding in memory. The binding expires after a configurable TTL, and Praxis replaces it when the selected provider is no longer eligible.
Because this state is local to one Praxis process, affinity is lost when later requests reach another gateway replica or when the original replica restarts. Horizontally scaling a consumer gateway can therefore move one conversation between providers, reducing conversation continuity and KV-cache locality.
Desired behavior
Add an optional shared affinity store backed by Valkey or Redis-compatible storage so every replica of a logical consumer gateway resolves the same session to the same stable provider candidate.
The binding should remain advisory rather than overriding routing safety: health, authorization, admission, provider removal, and hard policy constraints must still be able to invalidate it. A provider in existing_only should remain eligible for a session already bound to it but must not receive a new session.
Architecture considerations
- Use the provider candidate's stable identity, not a pod address or list position.
- Namespace bindings by the logical gateway or routing scope so unrelated gateways and tenants cannot collide.
- Store only an opaque hash of the session identifier; do not put raw cookies, conversation IDs, credentials, or request content in keys, logs, or metrics.
- Apply a configurable idle TTL and refresh it only according to a documented access policy.
- Make concurrent first-request binding deterministic or atomic so replicas do not establish conflicting bindings.
- Define atomic compare-and-replace behavior when a bound provider becomes ineligible.
- Preserve current in-memory affinity as the default for backward compatibility.
- Define explicit behavior when Valkey is unavailable. Do not silently claim shared affinity while using divergent replica-local bindings.
- Keep shared affinity out of Grid's CRDT/SWIM state; this is request-path data-plane state, not converged Grid control-plane state.
Grid integration
Grid should continue publishing stable candidate identities and admission states in the routing overlay. Praxis owns lookup and mutation of the affinity binding during request selection. Grid's responsibility is to provide a qualification topology and prove that drain, failure, restoration, and overlay changes interact correctly with shared bindings.
Acceptance criteria
- Two Praxis replicas configured with the same shared store route one session to the same stable provider when requests alternate between replicas.
- Different session identifiers can be distributed normally across eligible providers.
- Restarting either gateway replica does not lose the shared binding.
- A provider in
existing_onlycontinues serving sessions already bound to it but receives no new sessions. - An unhealthy, removed, unauthorized, or
noneprovider causes a bounded rebind to an eligible provider. - Concurrent first requests for one session cannot leave conflicting bindings.
- Binding expiry and refresh behavior are covered with deterministic tests.
- Valkey/Redis outage and recovery behavior is explicit, observable, and tested.
- Raw session identifiers and credentials do not appear in storage keys, metrics, logs, or qualification evidence.
- Existing deployments without a shared-store configuration retain the current in-memory behavior.
- A Helm/Forge-backed Grid qualification exercises two gateway replicas, at least two providers, restart persistence, drain, provider failure, expiry, and store outage/recovery.
- Evidence records the serving overlay revision and opaque binding/provider identities without exposing the session value.
Out of scope
- Replicating session bindings through SWIM or Grid CRDTs.
- Persisting model conversation content or KV-cache data.
- Allowing affinity to bypass health, authorization, or hard routing policy.
- Dominant language
- Rust
- Stars
- 10
- Forks
- 24
- Avg merge
- 1d 9h
- Merged PRs (30d)
- 79
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from praxis-proxy/grid
-
triage/needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
praxis-proxy/grid#294 ·
Maintainers usually reply within 1 day
-
triage/accepted
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
praxis-proxy/grid#275 ·
Maintainers usually reply within 1 day
-
triage/needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
praxis-proxy/grid#222 ·
Maintainers usually reply within 1 day
-
triage/accepted
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
praxis-proxy/grid#200 ·
Maintainers usually reply within 1 day
-
triage/needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
praxis-proxy/grid#134 ·
Maintainers usually reply within 1 day
All issues in praxis-proxy/grid
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
pact-foundation/pact-cli#154 ·
Maintainers usually reply within 3 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
antithesishq/bombadil#361 ·
Maintainers usually reply within 1 day
-
test(executor_l0): assert execute() TaskOutcome, not only bus events / 断言 execute() 返回的 TaskOutcomeOpentype:debt
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
skaiy/wild_agentos#425 ·
Maintainers usually reply within 1 day
-
Default-import note suggests `import * as process` for velt:process, which does not name the builtinOpen
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
bug ticket
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
cratestack/cratestack#1154 ·
Maintainers usually reply within 1 day