Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

deploy: atomically claim and own pending Agent Runtime operations

Cerrado
#76 1 comentario 0 reacciones 1 asignado Ver en GitHub

@asrujana-44 ya está trabajando en esto.

Desde el 17/9/2026.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Tranquilo
Stack tecnológico
python
Área
cli, cloud

Línea de trabajo

Comience en la ruta de despliegue de Agent Runtime del baseline y compárela con la rama Batch 4 del fork enlazado, centrándose en las transiciones de read/modify/write de deployment_metadata.json. Ejecute las regresiones de race y fallos de thread/subprocess mencionadas y, a continuación, verifique exactamente un owner y una mutation, una restauración byte-identical para los fallos previos al submit, un estado inicial fail-closed para resultados inciertos y una finalización y limpieza owner-safe.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Summary

Agent Runtime deployment operation state in v1.3.1 is a non-atomic read/modify/write record without an owner token. Two CLI processes can both pass the pending-operation check, start remote mutations, and overwrite or clear each other's deployment_metadata.json state.

Baseline: 5a306f8956cb1eeae69f9709de0e4d61b44e11e7 (v1.3.1).

Reproduction

  1. In one agent project, start two agents-cli deploy --no-wait processes at the same time.
  2. Both read the metadata before either writes its pending operation.
  3. Observe that both may submit a remote create/update, while the last metadata write wins.
  4. Also simulate:
    • local SDK request-config construction failing before remote submission;
    • remote create/update submission raising with an unknown outcome;
    • a delayed status/cleanup process running after a newer operation replaced the record.

Actual behavior

  • More than one remote mutation can start.
  • One process can overwrite or clear another process's pending operation.
  • A crash between remote submission and operation-name persistence is not represented safely.
  • Failure cleanup cannot distinguish known pre-mutation failure from outcome-uncertain remote submission.

Expected behavior

  • Exactly one process owns the right to start a mutation.
  • Ownership spans target revalidation, remote submission, status recording, completion, and cleanup.
  • Local preparation failure restores prior metadata byte-for-byte.
  • Once remote submission or identity creation may have happened, retain a fail-closed starting claim for manual reconciliation.
  • A stale owner must never clear a replacement owner's record.
  • Successful completion should merge current sibling metadata and remove the owned claim in one atomic transition.

Minimal fix

Use a cross-process lock around metadata read/modify/write, add a random claim ID, claim before remote mutation, revalidate the selected Runtime while holding ownership, and make clear/finish operations owner-aware. Split local request preparation from remote submission so only errors proven to precede mutation restore the previous bytes.

Reference implementation and regressions: fork Batch 4 branch.

Verification evidence

  • Thread and subprocess races each produced exactly one owner and exactly one mutation.
  • Replacement-owner, legacy-operation, atomic-finish, and sibling-metadata regressions passed.
  • Local request-config failure: zero mutation, byte-identical restoration, retry claim succeeds.
  • Remote submit and identity-create outcome-uncertain failures retain state=starting.
  • Fresh targeted review: 35 tests passed; full local suite: 92 passed.
  • ruff check src tests, ty check src, build, and Python 3.11/3.13 installed-wheel smoke tests passed.
Lenguaje dominante
Python
Estrellas
6k
Forks
686
Métricas de merge de PR
Sin PR fusionados en 30 d

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de google/agents-cli

Todos los issues de google/agents-cli

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.