Development deployment recovery: retired sandbox push auto-deploy replaced by a dev.relay.datatalks.club workflow
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 20/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- aws, github-actions, yaml
- Domain
- devops, infrastructure
Research direction
Read .github/workflows/deploy-sandbox.yml and docs/relay-deployment.md to confirm the committed target and deployment flow; do not run cloud queries without infra-owner authorization. Follow OPERATOR-PREFLIGHT.md v2 for the bounded reads and archive the evidence under .tmp/42-preflight/. Done means the owner-confirmed cause is recorded and the appropriate reviewed verification or verification-contract decision is completed, while preserving the failed run and #37.
Written by the indexing model from the issue text.
Description
Development deployment recovery: retired sandbox push auto-deploy replaced by a dev.relay.datatalks.club workflow
Status: groomed (v4) — owner resolved the blocker decisions 2026-10-05; implementation-ready behind two named prerequisite gates; v4 is a consistency amendment only (see changelog at the end — no gate weakened, no scope added)
Tags: (standard relay process labels do not exist in this repo — recorded 2026-10-05, none created)
Depends on: DataTalksClub/relay#43 — enforced RELAY_DRY_RUN must be on the deployed head and active before any development worker starts (not merely before explicit mail testing) · DataTalksClub/aws-infra#73 — scoped GitHub OIDC deploy role + applied-state/SSM-Online evidence (operator-supplied; agents perform no AWS reads; canonical infra path — never the website #58/#71 operator plan) · #37 stays OPEN: only the already-accepted calendar transport, integrated at 925286f4365887e6ff4663fbf529be2cbe758bfc, is a dependency of this issue; the rest of #37 must not create a deployment dependency cycle
Blocks: first development deployment of an exact accepted head; any change to deploy-sandbox.yml triggers
Owner-resolved decisions (2026-10-05, recorded by root delegation — do not re-ask)
- Retirement confirmed (failure branch E). The old sandbox host
i-03b7cd5de0a10a889/relay.dtcdev.clickwas intentionally stopped during the dtcdev→Relay transfer. No reconciliation of that target is attempted; the failed run 37253575983 stays immutable as the record of why. - Replacement development verification target. The dedicated existing main/dev Relay surface
https://dev.relay.datatalks.club— main account387546586013, workload regioneu-west-1— is the development deployment target. Committed contract: aws-inframain/devatf40e3a13798bce488a4d0b7822361cccbb24aa80(privatet4g.microARM64 host, shared-ALB target on port 8000, CloudFront front door atdev.relay.datatalks.club, log group/relay/dev/host). - Production untouched.
relay.datatalks.club, themain/relayleaf, hosti-044716a7d6b053eef,deploy-prod.yml, and allRELAY_PRODUCTION_*variables are out of scope.
Scope (relay repository only; no application code changes)
- New ordinary development workflow
.github/workflows/deploy-development.yml:testjob identical to the currentdeploy-sandbox.ymltest job:uv run ruff check .,uv run python manage.py makemigrations --check --dry-run,uv run python manage.py check,uv run pytest. No gate weakened or dropped.deployjob: OIDC via GitHubdevelopmentenvironment; assumevars.RELAY_DEPLOYMENT_DEPLOY_ROLE_ARN(required — fail-fast check step, production pattern, no committed instance/role fallbacks until aws-infra exports them); one SSMsend-commandto the requiredvars.RELAY_DEVELOPMENT_INSTANCE_IDineu-west-1, deploying the exactGITHUB_SHA(git fetch <sha>+checkout --force <sha>+bash scripts/deploy_relay_sandbox.sh --environment development <sha>); public health verificationhttps://dev.relay.datatalks.club/health/ready.- Triggers:
pushtomain(samepaths-ignoreas sandbox) +workflow_dispatch;concurrencygrouprelay-development,cancel-in-progress: false.
- Deploy script
scripts/deploy_relay_sandbox.shgains a third--environment developmentbranch. Sandbox and production branches keep their semantics; production's invocation fromdeploy-prod.ymlis unchanged. deploy-sandbox.ymlpush trigger removed. Manualworkflow_dispatchretained (rollback/manual behavior with explicit safe control — the operator must consciously choose to touch the stopped sandbox). Nothing else in the file changes; no target substitution.- Docs:
docs/relay-deployment.mdgains the development section (target, account/region, log group, dev runtime contract, rollback);README.mdsandbox lines marked stopped/retired with the decision date.
Development runtime contract (dev branch must satisfy; committed aws-infra main/dev + host user_data)
- Database: no local PostgreSQL, no
relay-postgrescontainer.DATABASE_URLfetched production-style from Secrets Manager viaRELAY_DATABASE_URL_SECRET_ARN(shared dev RDS databasedtc_relay_dev). Deploy must fail loudly if the secret is missing. - Edge: no Caddy; gunicorn binds
0.0.0.0:8000(the shared-ALB target port); public TLS terminates at the dev CloudFront/ALB path. - Containers: every container started with an explicit
--memorylimit (1 GiB host + 2 GB swap — host README obligation), same limits table as production. - Outbound only: no
inbound-emaildrain; the only SQS drain isses-webhooksfromSQS_SES_WEBHOOKS_QUEUE_URL(dev-owned queue pair). Norelay-sandbox-*drains (their queue-URL keys are absent on dev). - Runtime env:
DEBUG=False;ALLOWED_HOSTS/CSRF_TRUSTED_ORIGINSfordev.relay.datatalks.club(devinfrastructure.envdoes not carry them);PUBLIC_BASE_URL,DEFAULT_FROM_EMAIL([email protected]),AWS_SES_*,RELAY_EMAIL_SEND_ROLE_ARNflow frominfrastructure.env. Client-scope + API-key bootstrap provisioned in the dev branch (theRELAY_BOOTSTRAP_API_KEYpattern) so the release gate (required containers+/health/ready+scripts/smoke_test_relay.pysystem.echo) passes. Sandbox-only steps (dtc-courses senders, CMP template upsert, public-list env) must not run on dev. - Known source-level gap (offline proof at
925286f): the application has noRELAY_DRY_RUNsupport — the send path is direct boto3 SES (mailing/aws.py:ses_client,mailing/ses.py:_send_raw→send_raw_email), so the infra-sideRELAY_DRY_RUN=1contract is not app-enforced, andRELAY_API_KEYS_SECRET_ARN/RELAY_BIND_PORTare not read by the app (the bind is decided by the deploy script; the API-keys container is operator-facing). Consequences, binding on this issue: enforced dry-run is a hard prerequisite for starting any development worker — the follow-up enforcement issue is now filed as relay#43 (groomed 2026-10-05; canonical transport-boundary enforcement with truthful skip semantics and offline synthetic tests). No development worker may boot on a head lacking #43's enforcement; no email may be submitted or exercised on dev in this issue's deployment or verification; the host README's "turn dry-run off for one deliberate test" is forbidden until a separately authorized deliberate dry-run change lands after #43 (mail-feature code is a protected write path — #43 owns enforcement only, not dry-run deactivation).
Infra prerequisite (evidence gate — cite, never claim applied)
Owned by aws-infra#73 (groomed 2026-10-05): a least-privilege Relay GitHub OIDC deployer role in the aws-infra main/dev state root, trusted for exactly repo:DataTalksClub/relay:…:environment:development, scoped to SSM SendCommand/bounded command-result reads on the exact dev Relay host. Before the first development deployment, the infra operator posts on relay#73 and cross-posts here, from applied terraform -chdir=main/dev output on the exact aws-infra revision:
- the exact dev Relay host instance id (
relay_hostoutput) and deploy IAM role ARN (relay_github_deployeroutput) trusted forrepo:DataTalksClub/relay:…:environment:development(the committedmain/devroot has no Relay-repo OIDC deployer —deployment_iam.tfcovers only the website ECS deployer — so that role is precisely the aws-infra#73 change; it is not part of the website #58/#71 operator's saved-plan work and must not be absorbed there); - live confirmation that this exact instance id is SSM-
Onlineineu-west-1, account387546586013.
Repository variables RELAY_DEVELOPMENT_INSTANCE_ID and RELAY_DEPLOYMENT_DEPLOY_ROLE_ARN are then set from that evidence. No agent performs AWS reads or substitutes targets.
Acceptance Criteria
- Relay-side change implemented by SWE and landed through the full pipeline (separate SWE → independent Tester → PM acceptance → single commit with
Closes #42); write set limited to:deploy-sandbox.yml(trigger block + comment), newdeploy-development.yml,deploy_relay_sandbox.sh(development branch + help text),docs/relay-deployment.md,README.mdsandbox lines, optionaldocs/audits/2026-10-05-development-deployment-recovery.md. No protected files (mailing/,relay/,taskdeck/,jobs/,templates/,tests/app suites, migrations,deploy/Caddyfile,deploy-prod.yml,ci.yml,cli.yml,load-transfer-prod.yml). - Committed source proof:
deploy-sandbox.ymlhas nopushtrigger and retainsworkflow_dispatchwith unchanged target semantics;deploy-development.ymlcarries the identical four-step test contract, fail-fast variable checks,environment: development, no hardcoded instance id or role fallback. - Infra prerequisite evidence (above) archived on aws-infra#73 and cross-posted on this issue before the first deployment.
- App prerequisite: relay#43's enforced dry-run is on the deployed head and active before any development worker starts; no development worker boots on a head lacking it.
- One newly authorized ordinary development deployment of an exact accepted head (
925286f4365887e6ff4663fbf529be2cbe758bfcor a later reviewed head): Deploy Developmenttest✅,deploy✅, public health200athttps://dev.relay.datatalks.club/health/ready, host-sidegit -C /opt/relay rev-parse HEADequals the deployed SHA, running container set matches the dev contract (norelay-postgres, norelay-caddy, no inbound drains), observed by a sole exact-head OnCall. Failed run 37253575983 preserved; no rerun, retry, or cancellation of it; no rerun of the verification deployment. - Post-change push-to-main evidence: CI
testruns and no sandbox deployment is attempted against the retired host. - No email send exercised on dev anywhere in this issue's verification;
RELAY_DRY_RUNremains1/unset-off in the host runtime, and deactivating it is a separately authorized post-#43 change only. - #37 remains OPEN; existing full-suite acceptance (919 tests) and calendar/UI/features untouched — no application code changed by this issue; #37's remaining lifecycle work never gates this deployment.
- relay#43 (app-side
RELAY_DRY_RUNenforcement, mail-feature protected path) landed before worker boot — filed and groomed 2026-10-05; this issue dispatches it as a separate SWE role, not as its own work.
Test Notes
- No application behavior changes → no new app tests; the full existing suite must stay green.
- Offline (Tester, no cloud):
uv sync --frozen --dev;uv run ruff check .;uv run python manage.py makemigrations --check --dry-run;uv run python manage.py check;uv run pytest;bash -n scripts/deploy_relay_sandbox.sh; YAML parse of both workflows; static assertions: sandbox workflow has nopush:key; development workflow containsenvironment: development, required-variable fail-fast,--environment development,/health/readyondev.relay.datatalks.club; script development branch: memory limits applied, norelay-postgres/relay-caddy/inbound-email,DATABASE_URLfrom Secrets Manager, log group/relay/dev/host,DEBUG=False, does not writeRELAY_DRY_RUN. An optional offline test harness for the deployment-only script is allowed if it followsdocs/testing-guidelines.md(deterministic, no network, no AWS); it may add coverage but must never weaken or replace the required four-step test contract or the static assertions above. - Live (single OnCall, only after prerequisite evidence): the one exact-head dispatch; verify the three green gates, health, host-side SHA and container evidence, and the no-sandbox-on-push behavior. Source proof (committed files) and live proof (run logs, health, SSM/host evidence) are distinct; neither substitutes the other.
Stop conditions (any one halts the issue and routes back — no workaround, no gate weakening)
- Infra prerequisite evidence missing, or mismatching the committed contract (wrong account/region/hostname/role subject) → infra operator path; do not invent targets.
- Dev target not SSM-
Onlineat first deploy (InvalidInstanceIdagain) → infra operator; do not substitute instances or weaken the gate. - Host drift:
RELAY_DRY_RUNoff without a recorded deliberate decision; queue URLs not the dev-owned pair; missing required secret ARNs → operator review before any retry. - Secret containers unfilled (
DATABASE_URL/secret key) → deploy fails at migrate → operator; agents never populate secrets. - Any test-job failure → fix through the SWE pipeline first; test-only success is never reported as delivery.
- Deploy green but health/smoke fails → treat as failed deployment; diagnose via
/relay/dev/hostlogs; no acceptance. - Any push to
mainstill attempting a sandbox deploy after the change → reject the change. - Any production-surface change (prod workflow/host/variables/
relay.datatalks.club) or any protected-file edit → out of scope; route via root. - Any development worker started on a head lacking relay#43's enforced dry-run, any email send exercised on dev, or any attempt to disable dry-run expectations before a separately authorized post-#43 change → forbidden.
- aws-infra changes beyond the cited prerequisite role are owned by aws-infra#73's own scope discipline, never this issue; and this role must never be silently absorbed into the website #58/#71 operator plan — any such coupling is rejected at review.
Historical receipts (preserved)
Raw intake as filed (2026-10-05) — preserved verbatim from v1/v2
## Raw intake: exact-main sandbox deployment refused by SSM
Needs grooming. Preserve the failed run and reviewed application behavior.
### Observed evidence
- Accepted Relay main 925286f4365887e6ff4663fbf529be2cbe758bfc, tree f7278c8810e51d12e566de80343ebca7f96ade1c, focused source303041ea (Refs #37).
- [CI37253575966](https://github.com/DataTalksClub/relay/actions/runs/37253575966): completed success, required test job111585956863 success.
- [Deploy Relay Sandbox37253575983](https://github.com/DataTalksClub/relay/actions/runs/37253575983): test job111585957187 success, deploy job111587083812 failure.
- Failed step Deploy the complete Relay release: SSM SendCommand returned InvalidInstanceId / Instances not in a valid state for account, native exit254 at 2026-10-05T02:05:29Z. Public health verification was skipped.
- The committed workflow selects a sandbox target using repository variables with a committed fallback. This failure does not establish whether the target was retired, is unregistered/unavailable in SSM, or has an account/region/configuration mismatch. No cause is assumed.
- The sole observer archived exact-head push job metadata and one sanitized failed-log read. No rerun, retry, cancellation, source edit, provider query or production promotion occurred.
### Required investigation and outcome
1. Coordinate with the Relay migration/transfer owner and infrastructure owner to establish whether the sandbox target was intentionally retired during migration.
2. Define the smallest bounded, authorized sandbox-only read preflight to verify target/account/region/SSM readiness and the effective deployment variables. No secret values, production data or credentials inspection.
3. If sandbox remains required, reconcile only its approved target/configuration through the normal reviewed pipeline, then verify one newly authorized ordinary deployment of an exact accepted head with all three required jobs successful. Do not silently redirect sandbox deployment to production.
4. If it was intentionally retired, obtain an owner decision on the development verification contract before changing the workflow. Do not weaken or skip a required gate merely to report green.
5. Preserve failed run37253575983 and existing source/QA/PM acceptance. Full #37 remains OPEN; application CI success is not deployment or migration/cutover acceptance.
Related #23 and #24 are duplicate earlier disk-hygiene/documentation-drift intakes; this failure is a new SSM target refusal, not a disk-full failure. Open/all issue searches in Relay and aws-infra found no existing matching SSM refusal intake. DTC aws-infra#58/#71 are separate task-secret-reference reconciliation issues.
This intake grants no actual AWS/provider/credentials/SSM commands, Terraform plan/apply, production/API/mail/database operation, reset, rerun, deployment-target substitution or source changes. Offline PM grooming and owner coordination can proceed.
- v1 body sha256
eca12bbf…(replaced by v2 under a fresh exit-before-side-effect guard). - v2 body (branch taxonomy A–E, preflight
OPERATOR-PREFLIGHT.mdv2, pre-decision acceptance) sha256ae426299403647e918c12b3b36f2cd930b5a0aee03c8bda91808988867fe24f0(archived asISSUE-BODY-v2-archive.mdin the grooming task directory); raw intake body sha256bcf906450ca3292a4550146c1323f86fa24c95d459c3f62096ccf1ebf1f5e64d. - v2 grooming comments 5987032059 and 5987032061 (both 2026-10-05T02:23:53Z, identical, body sha256
d452b96e9c0b94b252c178fd5b945419222634f5a01ec2687a09eb3d704fbf6c) — retained, nothing deleted. - v3 resolution of the v2 taxonomy: branch E (intentional retirement) confirmed by owner decision; the v2 "verification-contract decision" required for E is now this body's dev.relay.datatalks.club contract. Branches A/B/C/D are moot for the old target; the preflight-of-the-old-target acceptance items are closed by the owner decision, not executed (no AWS reads were performed).
- v4 consistency amendments (2026-10-05, groomed alongside #43/#73; no gate weakened, no scope added):
- Prerequisite structure made explicit: relay#43 (enforced dry-run) must be on the deployed head and active before any development worker starts — strengthened from v3's "prerequisite for any dev mail testing"; the follow-up issue v3 asked the orchestrator to file is now filed and groomed as #43.
- The infra prerequisite role is owned by aws-infra#73 (groomed 2026-10-05) with its own offline gate, separate exact saved-plan review, and applied-evidence gate — replacing v3's "#58 operator path" phrasing; explicit guard added against absorbing the role into the website #58/#71 operator plan.
- #37 dependency narrowed truthfully: only the accepted calendar transport (integrated at
925286f) is a dependency; the rest of #37 stays OPEN and must never create a deployment dependency cycle. - Optional offline test harness for the deployment-only deploy script explicitly allowed per
docs/testing-guidelines.md, without weakening the required four-step test contract or static assertions.
Groomed by PM session authorized-relay42-development-pm-20261005 (zai model via zcy), 2026-10-05, at relay source 925286f4365887e6ff4663fbf529be2cbe758bfc and aws-infra f40e3a13798bce488a4d0b7822361cccbb24aa80, offline analysis only. Body-file receipt: sha256 recorded in the v3 grooming comment on this issue.
v4 amendment by PM session relay43-infra73-prerequisites-pm-20261005-zcy (zai/glm via zcy), 2026-10-05: the four consistency amendments recorded above, applied to the v3 body verbatim otherwise. v4 body-file sha256 recorded in the v4 grooming comment on this issue. Offline analysis only; v3 baseline snapshot sha256 c94d59969ab7d6463b0301e53c1131801b03580ec2ed0592090c3c8c24dcc9da verified immediately before publication.
- Dominant language
- Python
- Stars
- 0
- Forks
- 0
- Avg merge
- 7h 12m
- Merged PRs (30d)
- 12
Getting set up
- Ships a Dockerfile or Docker Compose file
- No pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from DataTalksClub/relay
-
Difficulty 5/5 Over a week Newbie friendliness 15/100
DataTalksClub/relay#41 · 9 comments ·
Maintainers usually reply within 1 day
-
community-base enhancement
Difficulty 5/5 Over a week Newbie friendliness 25/100
DataTalksClub/relay#39 · 6 comments ·
Maintainers usually reply within 1 day
-
community-base enhancement
Difficulty 5/5 Over a week Newbie friendliness 15/100
DataTalksClub/relay#37 · 15 comments ·
Maintainers usually reply within 1 day
-
community-base enhancement
Difficulty 5/5 Over a week Newbie friendliness 15/100
DataTalksClub/relay#35 ·
Maintainers usually reply within 1 day
-
community-base enhancement
Difficulty 5/5 Over a week Newbie friendliness 18/100
DataTalksClub/relay#34 ·
Maintainers usually reply within 1 day
All issues in DataTalksClub/relay
Similar issues
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitOpenneeds-triage
Difficulty 2/5 1-3 hours Newbie friendliness 77/100
krkn-chaos/krkn#1627 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NousResearch/hermes-agent#136483 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last elementPossibly taken @peterdsharpe claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
pytorch/tensordict#2307 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day