Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Determine whether a large dynamics RL batch refusal is necessary or avoidable by execution layout

Aperta
#869 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 3 giorni

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
python

Direzione di ricerca

Inizia da /home/brad/.local/share/schulman/memory-planning-audit-20260909/REPORT.md e dal file adiacente comment-869-progress.md, quindi esamina il report R0e e separate-lora-phases-scope/REPORT.md. Recupera, se possibile, il fixture, il profilo, il budget e lo stato esatti; il lavoro è completo quando documenta chiaramente se il rifiuto è necessario o se è adatto un layout di esecuzione equivalente, senza modificare il carico di lavoro né dichiarare una ripetizione non verificata.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Status reconciliation — September 11, 2026 (Schulman)

Open feasibility investigation; no incorrect refusal has been established. The exact historical replay remains blocked by missing fixture/profile state. Existing phased-execution comparisons do not prove a memory improvement or matching gradient execution.

Planned lane: Schulman and subagents, lower priority than reproducible #848/#870 boundaries. Recover an exact witness from retained evidence if possible; otherwise state that limit and use a separately identified representative case, without claiming reconstruction. Do not spend new GPU work trying to reproduce an unbound historical state.


Historical report (preserved):

Type: feasibility investigation, not a demonstrated incorrect refusal.
Owner: Schulman. Exact-witness replay is blocked by missing historical state; general planner work continues under #848.

Shannon's retail49 R0e run stopped on a TrainerRankMemoryError refusal. Preserved run-state v300 proves successful updates through batch 53; batches 54–56 contain rollout-only records and next_number is 57. Thus 57 describes driver progress, not a proven failed optimizer step. The guard raised before this attempted execution; this is distinct from a hard CUDA OOM.

The log reports 208288 packed / 568147 logical tokens, predicted 114.938 GiB against usable 63.101 GiB on one H200. Under the inspected source, token geometry describes the minimum wave's unsplit full-sharing plan, while the budget belongs to the final rejected split rung. At DP1, the matching 048 code makes one pair a top-level item, so this would be one indivisible pair, not the full eight-pair update. Historical ART/runtime-byte identity remains unconfirmed.

Preserve the exact failing batch/checkpoint and source/ART identities. Reconstruct top-level pair boundaries, per-history logical/packed lengths, request mix, sharing depth, cold/profiled state and original admitted/refused budget. Determine whether an indivisible input/request group forces the refusal, whether retained state is necessary, and whether a semantically identical execution schedule fits.

A reduced-history cap, changed pair weight or substitution of prepass values changes the experimental workload/objective; do not call such a change an equivalent execution fix. An auxiliary-first/generator-second backward schedule with one optimizer step is a candidate, not a qualified solution. If refusal is necessary, document the measured boundary and supported alternatives.

Evidence: /home/brad/shannon-logs/report-2026-09-09.md (R0e), associated run logs, and /home/brad/.local/share/schulman/retail49-program-20260909/analysis/separate-lora-phases-scope/REPORT.md. Cross-reference #848's broader cold/warm admission work. No GPU execution was performed to file this issue.

Current evidence gap: exact generated histories/tensors, selected profile/budget state and the pre-update G/aux state were not recovered. Queue indices and cal.log_trajectories metrics do not establish preserved raw inputs. No equivalent replay or necessary-refusal conclusion is claimed. Existing reconstruction: /home/brad/.local/share/schulman/memory-planning-audit-20260909/REPORT.md (SHA256 4c65c3c1e861eaf90286a19ecbdf43a53cdc138621184f2c3e8407f352b5fa95); concise evidence in adjacent comment-869-progress.md.

Lingua principale
Python
Stelle
10.8k
Fork
989
Merge medio
11h 38m
PR unite (30g)
104

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di OpenPipe/ART

Tutte le issue di OpenPipe/ART

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.