VideoFloatScaleScorer reports a score count in the rationale as a frame count
I maintainer di solito rispondono entro 2 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 88/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- backend, testing-qa
Direzione di ricerca
Inizia da video_float_scale_scorer.py e segui _score_frames_async in video_scorer.py, quindi leggi tests/unit/score/test_video_scorer.py:214-225. Aggiorna la motivazione in modo che riporti i frame estratti invece degli score restituiti e rafforza il test con un mock che restituisca più score per frame. Correggi anche il commento obsoleto sull’aggregatore in quel test.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
VideoFloatScaleScorer writes a frame count into the persisted rationale that is the number of scores, not the number of frames. With the class's own documented defaults they differ by a factor of four.
What a user sees. VideoFloatScaleScorer(image_capable_scorer=AzureContentFilterScorer()) with no other arguments, on a 5-frame sample (the default, video_scorer.py:28), produces a rationale that reads:
Video scored by analyzing 20 frames
Five frames were extracted and five were sent to the Azure endpoint. The persisted number is 20.
Why. video_float_scale_scorer.py:187:
rationale_prefix = f"Video scored by analyzing {len(frame_scores)} frames"
frame_scores comes from VideoScorer._score_frames_async (video_scorer.py:165-169), which returns whatever image_scorer._score_batch_nested_async(...) produced — one Score per (frame × category), not one per frame. MessageFloatScaleScorer deliberately does not aggregate per frame; contrast MessageTrueFalseScorer._score_async (true_false_scorer.py:172-216), which folds the pieces into one.
So the count is only right when the image scorer happens to emit one score per frame. The class docstring says the default FloatScaleScorerByCategory.MAX is chosen precisely because it suits "scorers like AzureContentFilterScorer that return multiple scores per frame" (video_float_scale_scorer.py:28-30, repeated at :62-68) — the default is documented for the case where the count is wrong. AzureContentFilterScorer defaults to all four TextCategory values (azure_content_filter_scorer.py:126-129, pinned by test_azure_content_filter.py:114).
Why the true/false sibling is right. video_true_false_scorer.py:160 has the same len(frame_scores), but there image_capable_scorer is typed MessageTrueFalseScorer, which must return exactly one score per piece — TrueFalseScorer.validate_return_scores raises otherwise (true_false_scorer.py:73-74). The float sibling has no such invariant, so the identical expression is only accidentally correct there.
The existing test cannot see it. tests/unit/score/test_video_scorer.py:214-225 uses MockFloatScaleScorer(return_value=0.8), which returns one score per frame, so the number comes out right by luck — and it asserts only assert "Video scored by analyzing" in scores[0].score_rationale. The digits are never checked. (Its comment "With MAX aggregator (default)" is also stale: the default is FloatScaleScorerByCategory.MAX, not FloatScaleScoreAggregator.MAX — see video_float_scale_scorer.py:48.)
What I would change. Count the frames actually extracted rather than the scores returned — len(frames) is already in scope in _score_frames_async's caller — and assert the number in the existing test using a mock that emits more than one score per frame. I would also fix the stale aggregator comment in the same PR.
Two smaller things in the same area, offered separately
Separate root causes, so I am not folding them in here. Say if you would like either as its own issue.
AzureContentFilterScorer._score_piece_asyncdocumentsRaises: ValueError: If converted_value_data_type is not "text" or "image_path"(:273-274) but has noelseand no raise, so an unsupported type yields a completed0.0score through_empty_resultinstead. The validator one level up means the pipeline does not reach it today, so it is a latent contract gap rather than a live misreport.- With an
audio_scorerconfigured, the final score's rationale isfinal_result.rationale(video_true_false_scorer.py:183), which drops theFrames (N): …line entirely, so the video's frame detail disappears from the persisted output.
What I have not verified. I have not run this against a real video and the real Azure endpoint; the count above is derived from the two defaults and the call chain. I have verified the code path, both defaults, and that the existing test cannot see the problem.
- Lingua principale
- Python
- Stelle
- 4.5k
- Fork
- 896
- Merge medio
- 2g 22h
- PR unite (30g)
- 220
Preparare l'ambiente
Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/PyRIT
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 2 giorni
-
Bug: triage GUI help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
microsoft/PyRIT#2868 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 76/100
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di microsoft/PyRIT
Issue simili
-
bug server
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
sportsdataverse/sportsdataverse-py#641 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
googleapis/google-cloud-python#18532 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno