Allow backup: on workload volumes, not only managed services
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 38/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- docker, go
- Área
- devops, infrastructure
Línea de trabajo
No se nombran archivos de implementación ni pruebas; empieza inspeccionando la ruta del esquema de onebox.run-v1 para workloads.*.volumes y muestreando otra implementación, tal como se solicita en la pregunta abierta. Determina si la copia de seguridad de volúmenes es representativa y si las semánticas propuestas de destino, quiesce, retención, exclusión y drill encajan con la maquinaria de copias de seguridad existente del servicio; terminado significa una decisión de plataforma acotada y un plan de implementación.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
backup: is available only on managed services — the entries that declare a driver, such as postgres. A workload that keeps its own state in a bind-mounted host directory has no declarative backup path at all.
Checked against the onebox.run-v1 schema at v2026.8.17. Workload volumes has no backup property:
.properties.workloads.additionalProperties.properties.volumes
-> "Managed named volumes or bind mounts. Relative bind sources are read-only
release content; absolute sources are external host state."
The schema calls these "external host state" and then offers nothing to protect them.
How this looks in a real app
From labstack/monk's ob.yml, the stateful surfaces:
| path | workload | contents | backup |
|---|---|---|---|
/data/qdrant |
qdrant | research embeddings | none |
/data/monk |
tape / server | market bar store | none |
/data/kestra/storage |
kestra | flow storage | none |
/data/cipher |
cipher | trading journal and order ledger | none |
| postgres | service | application database | PITR, 5m max loss, nightly, weekly drill |
| redis | service | cache | ephemeral by choice |
Four of six stateful things in the app have no path to the backup machinery, and the platform already has every piece they would need: a configured backup_targets entry with endpoint, region and failure domain; sops credential resolution; retention.keep; and drill.schedule with max_age.
The workaround, and why it is only adequate sometimes
schedule: on a role: job workload is a genuine answer for simple cases. For an append-only directory of a few hundred kilobytes, a nightly job that uploads a tarball plus a weekly job that extracts it and asserts it parses is proportionate, and arguably better than a platform feature because the verification is specific to the data.
It stops being adequate on three axes:
Consistency. A directory being written during the copy yields a torn read. Postgres avoids this because the driver knows how to quiesce it. A job-based copy of a live vector store or a bar store has no such story, and nothing in the platform signals that the resulting artifact is crash-consistent at best.
Drill. drill.schedule with max_age is the property that separates a backup from a hopeful upload. A job can perform a restore check, but nothing notices when the job stops running. Every app that hand-rolls this will write the upload; few will write the drill, and none will get staleness alerting for free.
Credentials. To write to the configured target, a job workload needs the backup credentials injected as env. That hands an application container the platform's own object-storage write credentials — a materially wider grant than any workload holds today, repeated once per stateful workload. The alternative is a separately provisioned scoped key per app, which is real operational work that exists only because the platform will not perform the write itself.
Proposal
Allow backup: on a workload volume, reusing the existing backup_targets, retention and drill blocks unchanged:
cipher:
role: worker
volumes:
- source: /data/cipher
path: /data/cipher
backup:
target: offsite
recovery_kind: snapshot
exclude: [sessions/]
quiesce:
command: [sh, -c, "sync"]
schedule:
cron: "0 17 * * 1-5"
timezone: America/New_York
retention:
keep: 14
drill:
schedule:
cron: "0 6 * * 0"
max_age: 8d
Design intent:
recovery_kind: snapshot, neverpitr. The semantics are weaker than a driver-managed backup and the name should say so, so that nobody plans a recovery around a guarantee this does not provide.quiesce— an optional command run in the workload before the copy, and its counterpart after. Trivial or absent for append-only data; meaningful for a store that can be asked for a consistent snapshot.drillverifies a restore into a scratch location and reports staleness the same way the service drill does. This is the main reason to build the feature rather than leave it to each app.- The platform performs the write, so no application container needs backup credentials.
excludekept deliberately minimal — a list of path prefixes, nothing resembling a sync DSL.
Explicit non-goals
This should not grow into a general backup product. No deduplication, no incremental chains, no arbitrary include/exclude expressions, no restore-to-arbitrary-target. That is restic or borg, and pulling it into onebox trades a small well-defined capability for a large one that competes with mature tools.
The scope is: copy a declared volume to a declared target on a schedule, expire it on the declared retention, and prove it restores on the declared drill.
Open question
The evidence here is one application. Whether four unbacked volumes is representative of onebox's install base or an artifact of how this app was built is the thing that decides whether this is a platform feature or an app-level pattern. Worth sampling another deployment before committing to it.
Context
labstack/monk#938 tracks the app-level cron job for the case that prompted this. That job is going ahead on its own merits and does not depend on this issue.
- Lenguaje dominante
- Go
- Estrellas
- 3
- Forks
- 0
- Merge medio
- 2 h 46 min
- PR fusionados (30 d)
- 38
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de labstack/onebox
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 80/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 65/100
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
Los mantenedores suelen responder en 1 día
Todos los issues de labstack/onebox
Issues similares
-
enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
AOSSIE-Org/DebateAI#611 ·
Los mantenedores suelen responder en 3 días
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
MHSanaei/3x-ui#6737 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
terraform-provider
Dificultad 2/5 1-3 horas Aptitud para principiantes 73/100
ClickHouse/terraform-provider-clickhousedbops#281 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
open-telemetry/opentelemetry-go-compile-instrumentation#1450 ·
Los mantenedores suelen responder en 2 días