Determine whether a large dynamics RL batch refusal is necessary or avoidable by execution layout
I maintainer di solito rispondono entro 3 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- machine-learning, performance
Direzione di ricerca
Inizia da /home/brad/.local/share/schulman/memory-planning-audit-20260909/REPORT.md e dal file adiacente comment-869-progress.md, quindi esamina il report R0e e separate-lora-phases-scope/REPORT.md. Recupera, se possibile, il fixture, il profilo, il budget e lo stato esatti; il lavoro è completo quando documenta chiaramente se il rifiuto è necessario o se è adatto un layout di esecuzione equivalente, senza modificare il carico di lavoro né dichiarare una ripetizione non verificata.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Status reconciliation — September 11, 2026 (Schulman)
Open feasibility investigation; no incorrect refusal has been established. The exact historical replay remains blocked by missing fixture/profile state. Existing phased-execution comparisons do not prove a memory improvement or matching gradient execution.
Planned lane: Schulman and subagents, lower priority than reproducible #848/#870 boundaries. Recover an exact witness from retained evidence if possible; otherwise state that limit and use a separately identified representative case, without claiming reconstruction. Do not spend new GPU work trying to reproduce an unbound historical state.
Historical report (preserved):
Type: feasibility investigation, not a demonstrated incorrect refusal.
Owner: Schulman. Exact-witness replay is blocked by missing historical state; general planner work continues under #848.
Shannon's retail49 R0e run stopped on a TrainerRankMemoryError refusal. Preserved run-state v300 proves successful updates through batch 53; batches 54–56 contain rollout-only records and next_number is 57. Thus 57 describes driver progress, not a proven failed optimizer step. The guard raised before this attempted execution; this is distinct from a hard CUDA OOM.
The log reports 208288 packed / 568147 logical tokens, predicted 114.938 GiB against usable 63.101 GiB on one H200. Under the inspected source, token geometry describes the minimum wave's unsplit full-sharing plan, while the budget belongs to the final rejected split rung. At DP1, the matching 048 code makes one pair a top-level item, so this would be one indivisible pair, not the full eight-pair update. Historical ART/runtime-byte identity remains unconfirmed.
Preserve the exact failing batch/checkpoint and source/ART identities. Reconstruct top-level pair boundaries, per-history logical/packed lengths, request mix, sharing depth, cold/profiled state and original admitted/refused budget. Determine whether an indivisible input/request group forces the refusal, whether retained state is necessary, and whether a semantically identical execution schedule fits.
A reduced-history cap, changed pair weight or substitution of prepass values changes the experimental workload/objective; do not call such a change an equivalent execution fix. An auxiliary-first/generator-second backward schedule with one optimizer step is a candidate, not a qualified solution. If refusal is necessary, document the measured boundary and supported alternatives.
Evidence: /home/brad/shannon-logs/report-2026-09-09.md (R0e), associated run logs, and /home/brad/.local/share/schulman/retail49-program-20260909/analysis/separate-lora-phases-scope/REPORT.md. Cross-reference #848's broader cold/warm admission work. No GPU execution was performed to file this issue.
Current evidence gap: exact generated histories/tensors, selected profile/budget state and the pre-update G/aux state were not recovered. Queue indices and cal.log_trajectories metrics do not establish preserved raw inputs. No equivalent replay or necessary-refusal conclusion is claimed. Existing reconstruction: /home/brad/.local/share/schulman/memory-planning-audit-20260909/REPORT.md (SHA256 4c65c3c1e861eaf90286a19ecbdf43a53cdc138621184f2c3e8407f352b5fa95); concise evidence in adjacent comment-869-progress.md.
- Lingua principale
- Python
- Stelle
- 10.8k
- Fork
- 989
- Merge medio
- 11h 38m
- PR unite (30g)
- 104
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di OpenPipe/ART
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 54/100
OpenPipe/ART#961 · 3 commenti ·
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
OpenPipe/ART#949 · 5 commenti ·
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 10/100
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 42/100
I maintainer di solito rispondono entro 3 giorni
Tutte le issue di OpenPipe/ART
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
PedestrianDynamics/pyFDS-Evac#199 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
521xueweihan/HelloGitHub#3790 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
sandialabs/atlas-ui-3#978 ·
I maintainer di solito rispondono entro 1 giorno
-
area: tests perceived difficulty: 2
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
Nitjsefnie-Harness-Commons/daedalus#1255 ·
I maintainer di solito rispondono entro 1 giorno
-
hf-audiolm-qwen: `generate_until` hardcodes `.to("cuda")` and aborts on non-CUDA acceleratorsAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
EleutherAI/lm-evaluation-harness#4256 ·
I maintainer di solito rispondono entro 1 giorno