fix: isolate provider-rate tests from mutable model catalogs
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 55/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Tranquilla
- Stack tecnologico
- python
- Ambito
- testing-qa
Direzione di ricerca
Start by running the focused command for tests/test_agent_loop_usage_sink.py and tests/test_provider_rates.py, then inspect the provider catalog fixtures and cache setup described in the issue. Trace which catalog tier and process-wide state answer before the mocked tier. Done means isolated and full-suite runs pass, including deterministic fetch counters, timestamps, and per-test disk-cache assertions.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
The Python unit job is red on main and on pull requests that do not touch Python because provider-rate tests depend on mutable LiteLLM and OpenRouter catalog state. The exact main base and PR #355 fail with the same 19 assertions while all TUI checks pass.
The tests mock an OpenRouter response containing deepseek/deepseek-v4-pro with a 163,840-token context window and fixed prices. In CI, another catalog tier or process state answers first with a 1,048,576-token window and different prices. The mocked fetch is therefore never called, and disk-cache assertions observe no write.
Steps to reproduce
-
Check out
mainat0e544ecb3493cfd51fdc910dedd037f71188ccd0. -
Run:
uv run pytest tests/test_agent_loop_usage_sink.py tests/test_provider_rates.py -
Compare the full CI runs:
The CI run reports 19 failures and 6,700 passes. A current local targeted run also remains catalog-sensitive: 65 tests pass and two MiniMax catalog assertions fail.
Expected behavior
Provider-rate unit tests should be deterministic and isolated from mutable LiteLLM and OpenRouter metadata, test order, process-wide caches, and local disk state. A TUI-only pull request should not inherit unrelated Python failures from the same base.
Actual behavior
The test fixtures expect mocked values such as context window 163840, but CI resolves 1048576 and different prices before reaching the mock tier. Fetch counters stay at zero, cache files are not written, and allow_fetch=False assertions receive catalog data instead of None.
The exact base commit fails with the same 19 tests as PR #355, so rerunning or rebasing the pull request onto the unchanged main does not address the cause.
Environment
- CI: GitHub-hosted
ubuntu-24.04 - Python: 3.12
- Dependency command:
uv run --frozen - Raven:
main@0e544ecb3493cfd51fdc910dedd037f71188ccd0 - PR evidence:
#355@7a7dc1a34c6769550f12eccf4c80f2cb2be67f61
Logs or screenshots
Representative failures:
assert 1048576 == 163840
assert 1048576 is None
assert 0.0033 == 0.00125
assert counter["calls"] == 1 # actual: 0
FileNotFoundError: .../model-catalog.json
19 failed, 6700 passed, 33 skipped
Acceptance criteria
- Provider catalog tests patch every upstream tier that can answer before the tier under test.
- Process-wide caches and timestamps are restored between tests.
- Disk-cache tests use only their per-test path.
- The focused files pass in isolation and in full-suite order.
- The full Python unit job passes on
main.
- Lingua principale
- Python
- Stelle
- 4.1k
- Fork
- 94
- Merge medio
- 10h 2m
- PR unite (30g)
- 376
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di EverMind-AI/Raven
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
EverMind-AI/Raven#798 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
EverMind-AI/Raven#797 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
EverMind-AI/Raven#640 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
EverMind-AI/Raven#479 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 75/100
EverMind-AI/Raven#474 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di EverMind-AI/Raven
Issue simili
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
letsencrypt/cp-cps#353 ·
-
Marble Madness II is missingAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
DOI-USGS/pywatershed#421 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
python-pillow/Pillow#10087 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno