Master Plan: Get Clustrix Production-Ready
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 15/100
Direzione di ricerca
Questo è un piano generale, non un singolo punto di ingresso. Inizia da Phase 0 ed esamina tests.yml, fast_ci.yml, real-world-tests.yml e real_world_tests.yml, quindi scegli un'attività con un ambito separato. Il lavoro è completato quando i criteri e i test della fase pertinente hanno esito positivo, ma l'issue non definisce una modifica delle dimensioni adatte a chi contribuisce per la prima volta.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Master Plan: Get Clustrix Production-Ready
Status: Planning · Opened: 2026-08-17 · Target: v0.2.0 (honest beta) → v1.0.0 (production)
This is the umbrella issue for taking Clustrix from its current state to a package that does what it advertises, is tested against real infrastructure, and has been adversarially reviewed. It is grounded in a five-track audit performed 2026-08-17 covering the full source tree, the entire test suite (actually executed), all 29 open issues, all documentation, and a security review of the working tree and git history.
Every claim below is cited to file:line or to verbatim command output. Where something could not be verified it is marked UNVERIFIED rather than assumed.
1. Executive summary: what is actually true today
| Question | Audited answer |
|---|---|
| Does the advertised feature set work? | Partially. SSH is solid. SLURM/SGE/Kubernetes are partial. PBS is broken. The entire cloud path has never executed successfully. |
| Is it tested? | No, in the way that matters. CI runs 15 of 2,280 tests (~0.7%). When the full suite is actually run: 127 failed, 1738 passed, 26 skipped, 8 errors. |
| Is the coverage number real? | No. Issues #61/#86/#98–#106 all reason from "74%". Measured full-suite coverage is ~56.14%; CI reports 8.21%; the committed artifact says 5.69%. Four different numbers, none of which was the planning input. |
| Can a new contributor run the tests? | No. The documented pip install -e ".[dev]" produces an environment where pytest cannot start. |
| Is it safe to run the tests? | No. The documented unit-test command provisions real, billable AWS EKS clusters. |
| Any leaked secrets? | Not on GitHub. Issue #107 was a false positive. Two real HF tokens exist in local-only git objects and should be rotated. |
| Biggest security defect? | Unauthenticated pickle deserialization of remote data — remote-to-local RCE. No issue existed for it. |
The single most important structural finding
Roughly every "completed" item in this project is half done in the same specific way: the scaffolding landed, the seam did not.
- Closure variables are detected and injected as parameters (
function_flattening.py:623,654-695) — but never passed at the call site (:749-751). - The REPL limitation is documented (
README.md:554) — but the code still fails opaquely (utils.py:118-119). - Kubernetes provisioners exist for all six targets — but GCP's
_assign_iam_role()assigns nothing (kubernetes/gcp_provisioner.py:356) and Azure's_create_service_principal()creates nothing (kubernetes/azure_provisioner.py:331). - AWS cleanup scripts exist — but have no dry-run, confirmation, or scoping (
grep -rn "dry_run\|confirm\|input(" scripts/aws/*.py→ 0 hits). - The cloud executor serializes
{"function": ...}(utils.py:151) and deserializesfunc_data['func'](executor_cloud.py:390) — a guaranteedKeyErroron the first real invocation.
That last one is diagnostic. It is not a regression; it is unconditional. Its survival proves the path was never executed. This is what heavy mocking produces: coverage rises while every integration seam stays untested. The fix is not "more tests" — it is tests that actually execute.
2. Root causes (fix these, not the symptoms)
RC1 — Tests mock the thing under test. 60 of 240 test files, 2,513 occurrences of unittest.mock|MagicMock|@patch|monkeypatch. Representative: tests/test_executor_schedulers.py:73-131 stubs execute_remote_command to return "Submitted batch job 12345", then asserts job_id == "12345" — it tests a regex against a string the test itself supplied. No SLURM is involved. This directly violates the project's own rule: "Do not use mock services for anything ever."
RC2 — Production code is mock-aware. clustrix/executor_scheduler_status.py:89 imports unittest.mock and branches on isinstance(self.connection_manager.ssh_client, Mock), with the comment "Use robust checking only if we have a real SSH connection (not unit tests)". The shipped code path _check_slurm_job_status_robust is structurally unreachable from any unit test. Separately, clustrix/notebook_magic_mocks.py ships fake widgets (_MockDropdown, observe(): pass) imported by six production modules.
RC3 — CI does not run the tests. tests.yml and fast_ci.yml both run only tests/unit/ = 5 files, 391 LOC, 15 tests. real-world-tests.yml has all three jobs gated if: false. real_world_tests.yml sets continue-on-error: true, as does the mypy step in tests.yml. ~99.3% of the suite never executes in CI, which is why all 127 failures are invisible.
RC4 — Planning ran against stale snapshots. #102 targets "notebook_magic.py, 1,236 statements"; that file is now 90 lines split across nine modules. #104 targets "executor.py, 511 lines"; that is now a 39-line re-export shim (executor.py:19-26). The work items were decomposed from one snapshot and the code was refactored underneath them.
RC5 — Failure is silent by default. 40+ except Exception: → pass/return None sites. executor_kubernetes.py:304-308 returns "completed" on any API exception. config.py:206 swallows all cloud-dependency install errors during __init__. Errors surface as wrong results rather than as errors.
RC6 — Local work never reached the remote. master is 30 commits ahead of origin/master (git rev-list --left-right --count origin/master...master → 0 30). All 30 implement issues #99–#106. quality_gates.yml is committed locally but returns HTTP 404 on GitHub — it has never run.
3. Backend reality check
| Backend | Status | Evidence |
|---|---|---|
| ssh | COMPLETE | Most mature path; real runner writes result.pkl (utils.py:1330-1452) |
| slurm | PARTIAL | Env-setup failures downgraded to logger.warning (executor_schedulers.py:105-116); status checker branches on Mock (executor_scheduler_status.py:89-105) |
| sge | PARTIAL | Implemented (executor_schedulers.py:200-240); fragile positional job-ID parse (:243) |
| kubernetes | PARTIAL | Results via stdout + ast.literal_eval (executor_kubernetes.py:339) — non-literal returns silently degrade to str; any API error returns "completed" (:304-308) |
| local | PARTIAL | Auto-parallel injects _parallel_{var} into kwargs (decorator.py:778) that user functions cannot accept; the resulting TypeError is swallowed (:708-716) |
| pbs | BROKEN | Script runs python execute_function.py (utils.py:1226) — that file is never created anywhere in the package. Also never calls setup_remote_environment (executor_schedulers.py:157-198) yet sources a venv. |
| cloud: lambda | BROKEN | Producer/consumer key mismatch → guaranteed KeyError (utils.py:151 vs executor_cloud.py:390) |
| cloud: aws/azure/gcp/hf | STUB | Dispatch requires create_instance (executor_cloud.py:222); classes define create_ec2_instance/create_vm/create_compute_instance/create_space → NotImplementedError (:234) |
Additionally, three providers return fabricated hostnames from get_cluster_config() error branches — placeholder.gcp.com (cloud_providers/gcp.py:544), placeholder.azure.com (azure.py:669), placeholder.lambdalabs.com (lambda_cloud.py:327,338) — which executor_cloud.py:260 then attempts to SSH into.
4. New capability: HuggingFace Jobs as the integration-test substrate
Verified working today against the contextlab org (academia plan). A real container ran and returned:
CLUSTRIX_HF_OK 3.12.14 x86_64
Available flavors: cpu-basic, cpu-upgrade, cpu-xl, t4-small, t4-medium, l4x1, l4x4, l40sx1, l40sx4, l40sx8, a10g-small, a10g-large, a10g-largex2, a10g-largex4, a100-large, h100, h100x8.
Why this matters. The reason the cloud path stayed broken for months is that verifying it meant provisioning, IAM, billing, and teardown — so nobody did, and a first-line KeyError went unnoticed. HF Jobs collapses that loop to one CLI call and a few cents, with no cluster reservation, no VPN, no institutional SSH credentials. That makes it a far more reliable CI substrate than the Dartmouth hosts the current real-world tests depend on (ndoli.dartmouth.edu, tensor01.dartmouth.edu), whose scheduled workflow has failed on all of its last 60 runs.
Note also that clustrix currently targets the wrong HF primitive. cloud_providers/huggingface_spaces.py targets Spaces — long-lived web apps. Jobs is exactly clustrix's model: hand over a container, run a function, collect a result, exit. An HF Jobs backend is both easier to implement correctly and immediately useful as test infrastructure.
Cost guardrail: anything CI-facing must pin --flavor cpu-basic with an explicit --timeout. GPU flavors stay manual and opt-in. h100x8 on the academia plan is real money.
5. Plan of record
Phases are ordered by dependency. Phase 0 must complete first — until then, running the test suite is unsafe and measuring anything is meaningless.
Phase 0 — Stop the bleeding (safety + ability to measure)
Nothing else can be trusted until these land.
- Contain the billable-test landmine
- Make the dev environment installable and the suite runnable
- Security hardening + token rotation + secret scanning
- Reconcile the 30 unpushed commits
Phase 1 — Establish ground truth
- Make CI actually run the suite
- Fix the 127 failures and 8 collection errors
- Set an honest, measured coverage baseline and a gate that means something
Phase 2 — De-mock
- Remove mock-awareness from shipped code
- Replace assertion-free mock tests with tests that execute
- Land the HF Jobs backend and adopt it as the default integration substrate
Phase 3 — Fix the core
- Cloud execution path and the provider interface
- PBS backend
- Local auto-parallelization
- The pickle trust model and SSH host-key verification
Phase 4 — Pay down architecture debt
- Delete ~5,100 lines of orphaned modules
- Replace silent-failure handling; fix resource leaks
Phase 5 — Make the docs true
- Fix every broken README example; unify versions
- Rewrite
CLAUDE.mdandMIGRATION.mdagainst the real module layout
Phase 6 — Red team
- Adversarial review of serialization, credentials, injection, and the remote trust boundary
Phase 7 — Release
- Cut an honest
v0.2.0; define v1.0 exit criteria
6. Exit criteria for v1.0
A release is production-ready when all of the following hold:
pip install -e ".[dev]"on a clean machine, followed by the documented test command, runs the suite to completion — with zero network calls to billable services.- CI executes >95% of the test suite on every PR, with
continue-on-errorremoved everywhere. - Zero failing tests. Zero collection errors.
unittest.mockappears zero times inclustrix/(shipped code).- Every backend claimed in
README.mdhas a test that actually executed it end to end, with a dated CI run link as evidence. - Coverage is measured, published, and gated at a threshold the project actually meets.
- Every code example in
README.mdis executed by a doc test in CI. - A single source of truth for the version; all four current locations agree.
- Secret scanning + push protection enabled; no plaintext credential ever written world-readable.
- The remote→local trust boundary is documented and either authenticated or explicitly opt-in.
7. Audit provenance
Five parallel audits, 2026-08-17, all read-only:
| Track | Method | Headline |
|---|---|---|
| Source tree | Full AST import graph over clustrix/ |
4 of 8 backends broken/stub; ~5,100 orphaned lines |
| Test suite | Suite actually executed in a purpose-built venv | 127 failed, 1738 passed, 26 skipped, 8 errors in 134.78s |
| Issues | All 29 open issues + comments, cross-checked against code | 9 closable, 12 with stale bodies |
| Docs | Every README claim verified against source | All 4 cloud config examples raise on line 1 |
| Security | Working tree + 555-commit history + gh api |
#107 false positive; pickle RCE is the real issue |
Known limits of this audit. The full-suite coverage figure (56.14%) is from combined worker data at 98% completion — the run reproducibly hangs in a real AWS retry loop (kubernetes/aws_provisioner.py:763,771), so it is a slight underestimate. PyPI publication status was not checked. Issue #87 (AWS EKS Service Control Policy) could not be verified from code as it is an AWS Organizations-level condition.
Sub-issues are linked below. Each carries its own evidence, acceptance criteria, and verification command.
8. Sub-issue roadmap
Phase 0 — Stop the bleeding (do these first; nothing else is trustworthy until they land)
- #109 —
tests/integration/is unmarked, so the documented unit-test command provisions billable AWS EKS clusters ⚠️ - #110 —
pip install -e ".[dev]"produces an environment where pytest cannot start - #111 — Security hardening: rotate local-only HF tokens, enable secret scanning, fix credential file permissions
- #112 — Reconcile the 30 unpushed local commits;
quality_gates.ymlhas never run
Phase 1 — Establish ground truth
- #113 — Make CI actually run the test suite (currently 15 of 2,280)
- #114 — Fix the 127 test failures and 8 collection errors that CI never sees
- #115 — Establish a reproducible coverage baseline (4 conflicting numbers in circulation)
Phase 2 — De-mock
- #116 — Remove mock-awareness from shipped code (production branches on
isinstance(..., Mock)) - #117 — Replace assertion-free mock tests with tests that execute (2,513 mock occurrences)
- #118 — Add HuggingFace Jobs backend and adopt it as the integration-test substrate
Phase 3 — Fix the core
- #119 — Fix the cloud execution path:
KeyError, provider interface mismatch, placeholder hostnames - #120 — Fix broken backends: PBS runs a nonexistent file; local auto-parallelization silently no-ops
- #121 — Security: unauthenticated pickle of remote results (RCE) and disabled SSH host-key verification ⚠️
Phase 4 — Architecture debt
- #122 — Delete ~5,100 lines of orphaned modules (13% of the package has zero importers)
- #123 — Replace silent-failure handling; fix resource leaks and import-time side effects
Phase 5 — Make the docs true
- #124 — Fix all broken README/docs examples; unify the version across 4 locations
- #125 — Rewrite
CLAUDE.mdagainst the real architecture; resolve the 3-way mocking-policy contradiction
Phase 6 — Red team
- #126 — Trust boundary, script injection, serialization fuzzing, cost safety
Phase 7 — Release
- #127 — Release v0.2.0 (honest beta) and define v1.0 exit criteria
9. Suggested order of attack
#109 first, alone. It is the only issue where not acting has an ongoing cost — anyone running the documented test command with AWS credentials is billing EKS right now.
Then #110 → #114 → #113: make the suite runnable, make it green, then turn CI on. Turning CI on before #114 makes master permanently red, and the temptation will be to re-add continue-on-error, which is how this situation arose.
#116 before #117. While production code still branches on isinstance(..., Mock), you can delete a thousand mock assertions and still be exercising the wrong code path.
#122 before #115. Deleting dead code moves the coverage number without writing a test; re-baselining first means measuring a denominator you are about to change.
#118 early — it unblocks honest testing for #119, #120, and #126.
10. Issue triage performed alongside this plan
Of the 29 previously-open issues: 9 closed, 8 updated with corrected bodies.
| Action | Issues |
|---|---|
| Closed — false positive | #107 (scanner matched AWS's own doc placeholder) |
| Closed — completed | #85 (SGE is fully implemented), #82, #72 |
| Closed — superseded / duplicate | #61 (dup of #98), #86 (dup of #99+#102) |
| Closed — archival, no deliverable | #92, #93, #94 |
| Updated — stale body corrected | #66, #68, #88, #89, #90, #91, #95 |
- Lingua principale
- Python
- Stelle
- 10
- Fork
- 4
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di ContextLab/clustrix
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
ContextLab/clustrix#175 ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
ContextLab/clustrix#170 ·
-
enhancement epic
Difficoltà 5/5 Più di una settimana Idoneità per principianti 20/100
ContextLab/clustrix#160 ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
ContextLab/clustrix#155 · 1 commento ·
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 42/100
ContextLab/clustrix#151 · 5 commenti ·
Tutte le issue di ContextLab/clustrix
Issue simili
-
Claiming namespace `jft63`Apertanamespace operations
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
EclipseFdn/open-vsx.org#14043 ·
I maintainer di solito rispondono entro 1 giorno
-
netbox status: needs triage type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
netbox-community/netbox#23376 ·
I maintainer di solito rispondono entro 1 giorno
-
feedback simulation workshop
Difficoltà 2/5 1-3 ore Idoneità per principianti 73/100
githubnext/gh-aw-workshop#4455 ·
I maintainer di solito rispondono entro 1 giorno
-
Triage 🩺
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitApertaneeds-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 77/100
krkn-chaos/krkn#1627 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno