Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

TrainerRank MoE memory admission: backward-recompute OOM and profile calibration failures

Đang mở
#848 28 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 3 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
20/100
Loại issue
Lỗi
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Sôi nổi
Công nghệ
python, pytorch

Hướng nghiên cứu

Bắt đầu với dev/trainer_rank_landing_acceptance.py bằng phase cost-calibrate, sau đó kiểm tra src/art/megatron/lora.py và src/art/megatron/kernels/cute_grouped_lora_quack.py quanh forward path đã được báo cáo. Chạy observer-off control và unchunked 24-view no-grad regression; hoàn tất yêu cầu có admission result được chứng minh mà không có OOM đã báo cáo, đồng thời duy trì tính nhất quán số học.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

September 25, 2026 — merged corrections; broader admission problem remains open.

PR #898 is now merged as a655c051105038960b0dc97be1e74bc98eeed663; PR #888 is merged as e2e8969ff1f64151ae613abf0050e236ee77e35e. Earlier dated statements below that these two changes are held are historical. Separate #931 and #971 remain open; frozen experiment packages are not automatically repinned.

The newly audited two-H200, 19-history instrumented workload still reached an actual expert-combine OOM. A 3,526,709,248-byte FC1 backing allocation was live before the failed 856 MiB request. This identifies a candidate for further memory reduction, not proof it can be released safely: bounded saved-tensor logging ended before FC1, ownership remains unresolved, and the observer affected compilation. The owner ended the stalled attempt through its registered cleanup; native termination and partial evidence are preserved, and both GPUs/process families are closed.

PR #971 addresses live HybridEP communication extent in recompute pricing. Its small native fixture and focused source tests passed; the fixture demonstrated the workspace-floor correction, but total admission remained unchanged. Required hosted GPU CI has not qualified it because scheduling failed before the tests. General memory bounds, the full 19-history workload, packing-dependent numerical differences, and arbitrary backward safety remain unresolved. Independent closed-result evidence: full19-outer-hook02-independent-audit-01/REVIEW-RECEIPT.json (af823635) and hybridep-extent-native-attempt02-audit/REVIEW-RECEIPT.json (891f7bef) under /home/brad/.local/share/schulman/successor-reviews-20260925/.


Earlier dated history is preserved below.

September 23, 2026 — additional workloads qualified; general admission issue remains open.

Current #898 is 4a519f6dcfe4f9dd18d966a9c98d4cc39fc1d125, based on held #888 3b0c6d4e. Both have three source clearances and green quality/two-H200 CI. The changed admission/recovery behavior is still held; it is not on main. Separate #931 is 3bdf3310, with green CI and source reviews, and is also held.

Two more one-H200 workloads passed on combined ART f80dba32: the complete 052 hidden-output probe (256 items, 384 requests, 5,622,779 logical tokens through one API call, 27 planner-selected waves, no caller splitting), and the 053-shaped discriminator/DPO callback (two groups, four alternatives each, one finite joint optimizer update). Original tasks and managed processes returned 0; known owned resources are closed. Independent reviews support completion within the recorded scope. The 14 planner modules match public #898, but the complete package identity differs. These runs exclude #931's later workspace terms.

The 052 diagnostic report capture is incomplete; admission operands and memory peaks were not recovered. A private source-hardlink reader fix passed focused controls/review for future use, without replaying or relabeling the original capture. The 053 workload is a retained-data structural surrogate, not exact historical step32, and retains its two caller cache-release sites. Neither proves numerical parity, estimator accuracy across shapes, or arbitrary backward safety. The separate current-source #888 ON/OFF pair is now complete, with all 12 updates and known cleanup successful in each arm. The unchanged strict comparator is INCONCLUSIVE because packing/call records and starting reserved-cache bytes differed. The separate complete-work observation is +2.756% elapsed time enabled (518.711 versus 504.801 seconds), not an isolated policy-cost estimate or a five-percent guarantee.

Durable completion evidence: /home/brad/.local/share/schulman/art848-resume-20260921-root/052-local-forward-success-20260923/acceptance.json (b9f155f3) and 053-local-structural-success-20260923/acceptance.json (a90b2033). Earlier completed full-eight/update/flat16 evidence and original failures remain intact. Numerical questions #901/#902 and physical-headroom #870 remain distinct.


Earlier dated history follows; its old heads and pending-work statements are superseded above.

September 22, 2026 update — still open.

Current #898 is 804bf4b6c3e05854efd831081ea95ed7a511b18c; exact-head source/CPU and hosted two-H200 CI passed. A composed one-H200 qualification now completed the full retained 049 prefix (six backwards and an optimizer update), then the intended 101,280-token target backward. For the target, the measured reset-to-peak increment was 43.330 GiB against a 60.919 GiB prediction and 62.747 GiB usable budget. This is a conditional measured interval, not a bound on unmeasured pre-reset or external-library allocations; numerical equivalence was not established. All owned resources are closed. This supersedes the earlier statement that the target had not been reached.

Both current baseline and candidate separately completed the unchunked 24-view no-grad workload but failed the original pooled-hidden repeatability assertion; NLL checks passed. They ran on different GPUs, so these outcomes do not attribute a regression. A baseline-only layer trace is now being collected to locate the first observed difference, preserving the original assertion and instrumented scope. #901/#902 remain separate unresolved correctness questions. The current-source full 660-gradient comparison is staged, not run; historical same-source variability and observer-off failures are not relabeled as current-source qualification.

Separate #931 (c683b282) covers additional overlapping eager head/statistics buffers. Focused CPU storage/algebra tests and an exact-head hosted two-H200 run passed. A duplicate cancelled workflow left a failed gate, retained separately from the completed successful run. It remains held; its head-buffer correction is not an explanation of every hidden-output or backward OOM.

Schulman retains ownership. Near-cap calibration, arbitrary backward/library safety and numerical acceptance remain unproved; no issue closure or merge is inferred. Current evidence: art848-resume-20260921-root/898-888-fullprefix-flat16-root-completion-closure.json (400bd6ea), 898-layer-trace-baseline-native-root-release.json (929a05dd) and CURRENT.json. No tolerance change or art.megatron edit.


Earlier dated history is retained below; this update supersedes old head and pending-run statements.

Current status — September 21, 2026, 6:35 AM Mountain: unresolved; focused corrections and renewed qualification in progress.

  • Held PR #898 remains 05e8ebf32dd7edb6d3e7793b552e24f32aecf054, with green exact-head CPU and hosted two-H200 CI. The restored-state one-H200 attempt failed during the original 049 prefix: after four backward returns, admission refused 70,758 packed / 88,576 logical tokens (predicted 37.200 GiB versus usable 33.278 GiB). No optimizer update completed, the intended 101,280-row target backward was not reached, and no CUDA OOM was observed. This is not proof of true infeasibility. All exact owned resources and known host identities are closed.
  • The frozen 049 prefix lacks merged Caladan #411's prepass-local lifetime fix. An exact #411-only overlay is under independent review for a fresh package; no loss/input/order/ART change or measured freed-byte attribution is claimed.
  • PR #929 merged as c6ac48f3698c7eefe7c914533f30d596ecc513a4. Three reviews and CPU/two-H200/main CPU CI passed. Automatic image build remains in progress. Frozen test packages are not automatically repinned.
  • A separate private estimator correction reserves overlapping eager head/statistics and no-gradient logits-copy buffers. Original 12 cases and the final composed 17 focused CPU cases passed; the earlier 180-second timeout is preserved. The correction is not GPU-qualified or a whole-model bound and does not explain Shannon's hidden-state/target-only failures.
  • Retained 052/053 logs locate six no-gradient OOMs in expert FC2 LoRA and three backward failures in checkpoint recomputation. The planner already reads live counters and prices hidden outputs. An original-052 callable over recovered 048 data is being integrated as an explicit surrogate, not an exact recovered failure.

Primary owner: Schulman with six existing lanes. #870 physical headroom and #901/#902 numerical findings remain separate. Durable report: /home/brad/.local/share/schulman/art848-resume-20260921-root/PROGRESS.md.


Earlier status/history below is preserved; current statements above supersede old heads and pending-run descriptions.

Current status — September 18, 2026, late evening Mountain: unresolved. Schulman and six delegated lanes remain active.

Current PR #898 is at 8b3e1dfb0059554223666a84db99bfb86371da63. Three exact-head source reviews and CI, including hosted two-H200 tests, are clear. The new full-model H200 run completed three forward/backward passes without OOM: its predicted incremental peak was 30.471 GiB, while the measured cold/warm increments were 23.451 / 21.679 / 21.679 GiB. These are measured reset-to-peak intervals; the admission-to-first-reset interval was not measured. This is not a general backward-memory bound or proof that one added estimate term caused the improvement.

The native test exited 1 because the unchanged numerical consistency check failed on 30 and 33 of 660 gradient tensors. All tensor relative-L2 comparisons passed the existing 0.03 threshold. Many failing tensor groups overlap failures on main and the previous candidate; causation and training impact remain unresolved. No optimizer updates ran in this diagnostic. Both temporary 40-entry weak-storage banks expired before explicit garbage collection. All four exact provider resources and recorded host processes/groups were independently retired.

Next work: an otherwise matched observer-off control to test the temporary instrumentation, and a separate unchunked 24-view no-grad regression on pristine current source. The earlier long-input native test did pass two caller-chunked 16-view calls of 430,000 logical rows each with three loaded slots, but it explicitly cleared the caller cache before each call. It does not establish that callers can remove those workarounds. The new source proposal contains 538,616 logical rows; physical packed size remains planner-selected, and this is not an exact reconstruction of the lost original Shannon failure.

Evidence: /home/brad/.local/share/schulman/art898-checkpoint-gradient-prototype-20260919-halley/native-fresh-launch-source-v1/root-native-closure.json (035ad2db); scientific assessment /var/tmp/art898-gradient-gpeak-native-assessment-20260919-vnea9__9/manifest.json (bbc6621b); numerical diagnosis /var/tmp/art898-gradient-family-diagnosis-20260919-89082zgh/manifest.json (9888815b). PR remains draft/held. General calibration, representative no-grad admission, arbitrary backward safety and issue #870 remain open.


Earlier status/history, preserved; current statements above supersede old head and pending-run descriptions:

Current status — September 17, 2026: unresolved; Schulman owns the follow-up.

GitHub closed this issue when PR #907 merged. The closing event identifies #907, whose scope was distributed error propagation, not memory estimation. Reopening corrects that status; the general admission issue remains unresolved.

PR #898 is the current accounting correction, at 04863263436b2d6977c171f059e8681636a6218d. It covers additional retained activations/workspaces and the actual selected adapter layout. The latest two fixture regressions pass, and fresh reviews/CI are underway. The PR remains held for the changed admission behavior. Merged #899 checks that the admission budget is still current before execution; merged #900 uses physical free memory and budgeted cache release. These address specific defects, not a general memory bound.

Remaining evidence: a frozen eight-sequence backward workload has an observed cold incremental peak about 8.03 GB above its estimate. Recent runs complete three backwards without OOM but fail the original gradient comparisons; this uses an older frozen runtime and does not qualify the current PR. Active lanes are (1) repeat that workload on current planner code, (2) identify retained tensors around recomputation, and (3) investigate repeated-gradient variation under #902. All completed diagnostic GPU resources are independently cleaned up. General backward safety, near-cap admission and representative performance remain open.

Latest retained receipts: /var/tmp/art848-warm-leaf-image-20260917-root/root-readout-acceptance.json (5b281455), /var/tmp/art898-fixture-contract-result-20260917-physical-wpe6ob97/manifest.json (7712f524). Detailed historical evidence below is preserved; its old pending/adoption wording is superseded by this status.


September 12 adoption and diagnostic update: Caladan #400 is merged at0ad1ffaffc343d080de6d1b6e4db9c3ceba56066, aligning all three ART pins to a3a248a8e457cf5f244b7371802a0c531901e143. Exact reviewed-file/actual-parent-tree parity is verified. Main Prek and automatic service rollout34710747657 passed, including sustained gateway health. The separate GPU image built/prewarmed, but smoke34710438574 failed to acquire an H200 via Sky before model execution. That generic provisioning error is not a demonstrated GPU-capacity or model-code cause; no manual retry was dispatched.

The unpressured allocator diagnostic's reader findings are corrected and independently cleared (reader1ada0d69; reviewf208a7d3). Its fresh exact archive936212ff passed the pinned native image CPU qualification: 35 estimator tests, imported source/history API checks and bounded writer/readback; native/outer0 with all owned resources independently removed. CUDA history/model/actor execution was not exercised by that CPU check. Root receipt: /home/brad/.local/share/schulman/art848-workspace-root-review-20260912/cpu-qualification.json.

Active remaining work: compose and qualify the actual GPU factory/Task/artifact-collection/terminal boundaries, then measure allocator lifetimes around one unpressured cold forward. Keep snapshot materialization memory, final artifacts and original failure joins explicit. General cold/warm admission correctness, profile calibration, physical backward headroom#870 and the unrecovered historical refusal#869 remain open. No general memory bound or native allocator result is claimed.


Earlier observations (preserved; current adoption/reader status above supersedes earlier pending wording):

September 12 update: ART #891 is merged at a3a248a8e457cf5f244b7371802a0c531901e143; its tree equals the reviewed 75d984c tree. It accounts for the identified FC2 grouped-LoRA workspace in the cold/fallback estimate. The exact bounded pressure successor refused before execution, whereas the original admitted case OOMed. This is a specific correction, not a complete general memory bound. PR two-H200/source CI and merged-main Prek passed; the image built/prewarmed but its smoke could not provision an H200. Caladan .art-revision remains 1cefd5c1 at current main15cdc48.

Remaining active work: instrument actual allocator lifetimes around one unpressured cold forward. The private composed diagnostic has CPU evidence, but independent review found incomplete error/event/source joins in its offline reader; correction is required before native use. No new GPU memory attribution is claimed. Evidence: /home/brad/.local/share/schulman/art848-workspace-composed-20260912-memory and art848-composed-review-20260912-LK6qsJ. General admission, calibration, warm pricing, #869 and #870 remain distinct.


Earlier record (preserved; status above is current):

Status reconciliation — September 12, 2026 (Schulman)

Partially addressed; general admission correctness remains open and actively investigated. The merged output-lifetime/admission-evidence, inactive-request, and #889 selected-layout/check-consistency fixes address specific defects. They do not qualify a general memory bound. The private warm pricing and cache-release candidates remain distinct from current shared adoption.

New independently verified cold-admission counterexample: one H200, Qwen3.6-35B-A3B, TP/CP/DP/EP1, a fixed no-gradient auxiliary forward with 45,981 packed / 172,687 logical tokens and 14 hidden-output requests. The test deliberately held 68,451,041,280 bytes (63.75 GiB) of private CUDA ballast. The original candidate admitted 4,092,810,444 required bytes against 4,652,784,333 available bytes. That exact admitted native forward raised CUDA OutOfMemoryError; ART preserved it as the explicit cause of TrainerRankMemoryError. Actor PID/thread/admission, original check, actual forward, final container replay and cleanup are joined. This is an observed OOM after admission, not a counterfactual inferred from an earlier peak.

The failure ran on the retained private _impl89cd2194; independent AST comparison confirms that its cold arithmetic, availability check and OOM wrapper match current ART1cefd5c1 / _impleb6073bd. No warm-profile term or backward cache-release helper was exercised. The selected pricing observation agrees numerically with the original check but remains a separate observation. There were zero backwards/optimizer updates, no completed output and no measured post-failure weight-equality claim. The intentional pressure case does not establish failure of ordinary unpressured execution or a universal reserve coefficient.

The earlier identical-package attempt failed on a model-download HTTP503 before model/Trainer/forward initialization; it remains separately preserved. The one infrastructure retry above is complete. All four owned Kubernetes UIDs and host process groups are independently absent; actual outer OS wait and native container exit were1. No resource remains allocated for this test.

Private durable evidence: /home/brad/.local/share/schulman/art848-pressure-attempt02-evaluation-20260912-root/ (result SHA e7b70f0ec81024a6d21d5816f9cca78a4147976e2e975b117138ca9146ef0ede); independent review /home/brad/.local/share/schulman/art848-pressure-attempt02-review-20260912-memory/, manifest 4f6a7c4c71ee5b151abe24432071a003073325ab0f95c0b4620f8f2f788247bb. Raw captures remain private. Next: trace the actual failed allocation and derive a narrowly justified cold workspace term or pre-execution refusal; avoid tuning a generic coefficient from a single peak. Shared policy changes remain held pending behavior/scope review.

Owner: Schulman and subagents. Keep calibration overestimation, warm underpricing, #869's historical witness and #870's physical backward headroom separate. No art.megatron or API change is proposed by this evidence.


Historical report (preserved):

Found by the expanded cost-model calibration campaign (Qwen3.5-35B-A3B, TP1 × CP1 × EP1, one H200 141 GB, bf16, active LoRA slot, dev/trainer_rank_landing_acceptance.py --phase cost-calibrate). Two related memory-admission failures on this MoE model at CP1; both cells are fine at CP2/CP4.

1. Cold admission lets a forward run that then hard-OOMs. cal-grpo-g16 (106,432 logical tokens, no_sharing layout = 106k packed tokens) was admitted by the cold static estimate and died with a CUDA out-of-memory inside the expert grouped-LoRA GEMM on the first forward:

torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 3.25 GiB. GPU 0 has a total capacity of 139.81 GiB of which 2.31 GiB is free ... 134.11 GiB is allocated by PyTorch
  src/art/megatron/lora.py:1161 _expert_grouped_lora_forward
  src/art/megatron/kernels/cute_grouped_lora_quack.py:326 _varlen_quack_gemm

The contract is a TrainerRankMemoryError refusal before execution, never an OOM. The cold estimate does not account for MoE expert-path working memory (grouped GEMM / LoRA intermediates over top-8 routed rows), which for a 256-expert model with 70 GB of bf16 weights leaves far less headroom than the dense estimate assumes.

2. After one observation, everything is refused. cal-grpo-g4x4 (73,696 logical tokens): the first no_sharing warm-up ran (37 s, peak 129.3 GB), and every later forward in the process was refused with "forward is predicted to exceed available memory; unable to find a feasible split: every rung of the bounded ladder ...", including uniform_depth_2 at 21,487 packed tokens (less than a third of the observed forward). The retained-memory profile learned from the single 129 GB observation makes the predictor refuse layouts that plainly fit, and the best-effort splitting ladder (#831) finds no rung feasible either, so the cell produced no measurable rows.

Evidence: scratch/trainer_rank_cost_calibration/lattice35-ep1-synth-20260904-0007/{evidence.jsonl,tp1-cp1-ep1-etp1-cal-grpo-g16-0-g0.log,tp1-cp1-ep1-etp1-cal-grpo-g4x4-0-g0.log} on the (auto-downed) cluster; local copies with the calibration artifacts. Peaks observed on the same model at CP1: cal-grpo-g8 (53k tokens) 113 GB unshared / 74 GB shared, so the admissible envelope on one H200 ends somewhere between 53k and 74k unshared tokens.

Relates to the "planner-driven head chunking / memory margins" follow-up and the cold static estimate ignoring sharding; MoE adds an expert-path term the estimate lacks. For the calibration campaign the two CP1 cells are recorded as unmeasurable on this class (the CP2/CP4/EP shapes cover the same workloads).

Ngôn ngữ chính
Python
Star
10.8k
Fork
989
Merge trung bình
11 giờ 38 phút
Pull request đã merge (30 ngày)
104

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của OpenPipe/ART

Tất cả issue của OpenPipe/ART

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.