rust: in-process transport frees the connection callback state immediately when `copilot_runtime_connection_open` returns 0 (possible use-after-free)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 2/5
- Tiempo estimado
- 1-3 horas
- Aptitud para principiantes
- 76/100
Línea de trabajo
Comienza en rust/src/ffi.rs, en FfiHost::start_blocking, y compara la rama de apertura fallida con FfiShared::close y release_callback_state. Verifica que la ruta de apertura fallida ya no reclame CallbackState prematuramente y, a continuación, usa el shim descrito con ASan o una herramienta similar para confirmar que un callback retrasado no puede acceder al estado liberado.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
With the in-process transport (in-process / bundled-in-process features), FfiHost::start_blocking boxes a CallbackState and passes the raw pointer as user_data to copilot_runtime_connection_open. If the call returns 0, the SDK reclaims the pointer right away with Box::from_raw, drops it, and then calls host_shutdown.
Elsewhere, the SDK treats user_data as reclaimable only once the runtime has signalled quiescence. FfiShared::close / release_callback_state free the state only after copilot_runtime_connection_close returns true, and retry on a background thread until it does (the fix from #2610 / #2622 in v1.0.14). The failed-open path was not covered by that fix.
The C ABI as documented in this repo does not justify the immediate free. The prototype comment in go/internal/ffihost/ffihost.go and ADR-007 (java/docs/adr/adr-007-native-bundling-strategy.md) say only that connection_open "registers the on_outbound callback" and "Returns a connection handle (0 = failure)". Neither promises that a failed open did not retain user_data or will not invoke the callback with it. A failed open also returns no connection id, so the host has nothing to close and no quiescence signal to wait for.
If the runtime has already set up outbound delivery before the open fails, a callback can arrive after the SDK freed the state. on_outbound would then dereference a freed CallbackState and send on a dropped tx.
The Go SDK is not exposed to this: it passes an opaque token as user_data and removes it from its lookup map on failure, so a late callback is a harmless miss. The Rust SDK passes a raw heap pointer, so it is exposed.
Affected versions
- SDK rust v1.0.14 and v1.0.15. The failed-open arm in
rust/src/ffi.rsis byte-identical in both. - Observed against CLI runtime libraries 1.0.84-5 through 1.0.89.
Code location
rust/src/ffi.rs, impl FfiHost { fn start_blocking }, at rust/v1.0.15:
let state_ptr = Box::into_raw(Box::new(CallbackState { tx, closing: AtomicBool::new(false) }));
let connection_id = unsafe { (self.connection_open)(server_id, on_outbound, state_ptr as *mut c_void, /* … */) };
if connection_id == 0 {
drop(unsafe { Box::from_raw(state_ptr) });
unsafe { (self.host_shutdown)(server_id) };
return Err(Error::with_message(ErrorKind::InvalidConfig, "copilot_runtime_connection_open failed"));
}
Compare FfiShared::close / release_callback_state in the same file. They free only after connection_close returns true.
Minimal reproduction
This is a timing-dependent memory-safety issue, so the reliable repro is a shim library plus a sanitizer:
- Build a host with
features = ["in-process"]against a shim library. The shim'scopilot_runtime_connection_openspawns a thread that callson_outbound(user_data, bytes, len)after a short delay, then returns 0. The other exports forward to a real runtime library. - Start the client and observe
Err("copilot_runtime_connection_open failed"). - Under ASan or Miri-style tooling, the delayed callback reads the freed
CallbackStateand sends on a droppedtx.
Expected vs actual
- Expected: after a failed open, the SDK does not free
user_dataunless the C ABI guarantees it was not retained. - Actual: it is freed synchronously, while the ABI as documented leaves open whether the runtime still holds it and may invoke the callback with it.
Proposed fix
Minimal fix, which we carry as a local patch: on the connection_id == 0 arm, do not reclaim state_ptr. Deliberately leak the CallbackState, which is one small struct plus an unbounded-channel sender per failed open. In practice only host boot retries produce failed opens. Keep host_shutdown and the error return unchanged. Keep release_callback_state as the single place that calls Box::from_raw.
Alternatives:
- Pass an opaque token as
user_dataand look it up in a registry, as the Go SDK does. A late callback then finds no entry. - Document in the shared C ABI that a
0return fromconnection_opennever retains or invokesuser_data, with the runtime guaranteeing it. The current free would then be sound as written.
Patch diff summary
rust/src/ffi.rs: 1 hunk. It removes 1 line (drop(unsafe { Box::from_raw(state_ptr) });) and adds a comment explaining the intentional leak. No API change. Afterwards Box::from_raw appears exactly once in the file, inside release_callback_state.
- Lenguaje dominante
- TypeScript
- Estrellas
- 10.5k
- Forks
- 1.5k
- Merge medio
- 1 d 5 h
- PR fusionados (30 d)
- 108
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de github/copilot-sdk
-
agentic-workflows
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
github/copilot-sdk#2782 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
github/copilot-sdk#2781 ·
Los mantenedores suelen responder en 1 día
-
agentic-workflows
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
github/copilot-sdk#2779 ·
Los mantenedores suelen responder en 1 día
-
agentic-workflows
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
github/copilot-sdk#2760 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
github/copilot-sdk#2759 ·
Los mantenedores suelen responder en 1 día
Todos los issues de github/copilot-sdk
Issues similares
-
Mend: dependency security vulnerability untriaged
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
opensearch-project/security-dashboards-plugin#2545 ·
Los mantenedores suelen responder en 1 día
-
Add: Dream TR SDAbiertocheck:passed streams:add
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
doctor integrity sample scans soft-deleted pages on Postgres (batch path has no deleted_at filter)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
SocialGouv/egapro#4672 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
area:agents area:tui bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
anthropics/claude-code#98358 ·
Los mantenedores suelen responder en 1 día