VideoFloatScaleScorer reports a score count in the rationale as a frame count
Los mantenedores suelen responder en 2 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 2/5
- Tiempo estimado
- 1-3 horas
- Aptitud para principiantes
- 88/100
- Tipo de issue
- Error
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Stack tecnológico
- python
- Área
- backend, testing-qa
Línea de trabajo
Empieza en video_float_scale_scorer.py y sigue _score_frames_async en video_scorer.py; después, lee tests/unit/score/test_video_scorer.py:214-225. Actualiza la justificación para que informe de los frames extraídos en lugar de los scores devueltos, y refuerza el test con un mock que devuelva varios scores por frame. Corrige también el comentario obsoleto sobre el agregador en ese test.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
VideoFloatScaleScorer writes a frame count into the persisted rationale that is the number of scores, not the number of frames. With the class's own documented defaults they differ by a factor of four.
What a user sees. VideoFloatScaleScorer(image_capable_scorer=AzureContentFilterScorer()) with no other arguments, on a 5-frame sample (the default, video_scorer.py:28), produces a rationale that reads:
Video scored by analyzing 20 frames
Five frames were extracted and five were sent to the Azure endpoint. The persisted number is 20.
Why. video_float_scale_scorer.py:187:
rationale_prefix = f"Video scored by analyzing {len(frame_scores)} frames"
frame_scores comes from VideoScorer._score_frames_async (video_scorer.py:165-169), which returns whatever image_scorer._score_batch_nested_async(...) produced — one Score per (frame × category), not one per frame. MessageFloatScaleScorer deliberately does not aggregate per frame; contrast MessageTrueFalseScorer._score_async (true_false_scorer.py:172-216), which folds the pieces into one.
So the count is only right when the image scorer happens to emit one score per frame. The class docstring says the default FloatScaleScorerByCategory.MAX is chosen precisely because it suits "scorers like AzureContentFilterScorer that return multiple scores per frame" (video_float_scale_scorer.py:28-30, repeated at :62-68) — the default is documented for the case where the count is wrong. AzureContentFilterScorer defaults to all four TextCategory values (azure_content_filter_scorer.py:126-129, pinned by test_azure_content_filter.py:114).
Why the true/false sibling is right. video_true_false_scorer.py:160 has the same len(frame_scores), but there image_capable_scorer is typed MessageTrueFalseScorer, which must return exactly one score per piece — TrueFalseScorer.validate_return_scores raises otherwise (true_false_scorer.py:73-74). The float sibling has no such invariant, so the identical expression is only accidentally correct there.
The existing test cannot see it. tests/unit/score/test_video_scorer.py:214-225 uses MockFloatScaleScorer(return_value=0.8), which returns one score per frame, so the number comes out right by luck — and it asserts only assert "Video scored by analyzing" in scores[0].score_rationale. The digits are never checked. (Its comment "With MAX aggregator (default)" is also stale: the default is FloatScaleScorerByCategory.MAX, not FloatScaleScoreAggregator.MAX — see video_float_scale_scorer.py:48.)
What I would change. Count the frames actually extracted rather than the scores returned — len(frames) is already in scope in _score_frames_async's caller — and assert the number in the existing test using a mock that emits more than one score per frame. I would also fix the stale aggregator comment in the same PR.
Two smaller things in the same area, offered separately
Separate root causes, so I am not folding them in here. Say if you would like either as its own issue.
AzureContentFilterScorer._score_piece_asyncdocumentsRaises: ValueError: If converted_value_data_type is not "text" or "image_path"(:273-274) but has noelseand no raise, so an unsupported type yields a completed0.0score through_empty_resultinstead. The validator one level up means the pipeline does not reach it today, so it is a latent contract gap rather than a live misreport.- With an
audio_scorerconfigured, the final score's rationale isfinal_result.rationale(video_true_false_scorer.py:183), which drops theFrames (N): …line entirely, so the video's frame detail disappears from the persisted output.
What I have not verified. I have not run this against a real video and the real Azure endpoint; the count above is derived from the two defaults and the call chain. I have verified the code path, both defaults, and that the existing test cannot see the problem.
- Lenguaje dominante
- Python
- Estrellas
- 4.6k
- Forks
- 924
- Merge medio
- 3 d 4 h
- PR fusionados (30 d)
- 264
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Sin guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/PyRIT
-
BUG PuzzledConverter cannot select words carrying non-ASCII letters, so the mask falls on articles insteadPosiblemente ocupada @adimalkar la tomó hace 1 día. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
microsoft/PyRIT#3022 · 2 comentarios ·
Los mantenedores suelen responder en 2 días
-
BUG: PlagiarismScorer accepts invalid n-gram size and blank reference textPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Los mantenedores suelen responder en 2 días
-
PackageHallucinationScorer (Python) misses `from pkg.sub import x` and indented importsPosiblemente ocupada @barry166 la tomó hace 8 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
microsoft/PyRIT#2948 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
BUG Configuration keeps runtime-status errors after polling recoversPosiblemente ocupada @rupayon123 la tomó hace 14 días. AbiertoBug: triage GUI help wanted
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
microsoft/PyRIT#2868 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
BUG PlagiarismScorer tokenizer strips combining marks and skips normalization, so a verbatim copy can score 0.0 and different words can score 1.0Posiblemente ocupada @inchang-ing la tomó hace 1 día. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 25/100
Los mantenedores suelen responder en 2 días
Todos los issues de microsoft/PyRIT
Issues similares
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
awslabs/visual-asset-management-system#414 ·
Los mantenedores suelen responder en 1 día
-
bug v1 v2
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
modelcontextprotocol/python-sdk#3670 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
aicell-lab/bioengine#232 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
modelscope/evalscope#1836 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100