Questions about HPSv3 scorer correctness, Figure 2 sample quality, and DiffusionNFT baseline
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- python
- Área
- machine-learning, testing-qa
Línea de trabajo
Start with src/hpsv3_scorer.py and compare its preprocessing and scores against the official HPSv3RewardInferencer.reward(...) on identical image/prompt pairs. Then trace the Figure 2 generation configuration and the released DiffusionNFT baseline entry points. Done means documenting reproducible prompts, seeds, checkpoints, sampling settings, and representative baseline results, or confirming the reported implementation discrepancy.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Hi, thanks for releasing the code. I have a few questions regarding the implementation and the reported qualitative results.
1. HPSv3 scorer implementation appears to miss image normalization
I think there may be a serious correctness issue in the current HPSv3 scorer.
In src/hpsv3_scorer.py, the differentiable preprocessing path calls:
self.ip._preprocess(images01[i:i + 1], do_rescale=False)
However, in the official HPSv3 differentiable image processor, _preprocess() does not resolve do_normalize=None, image_mean=None, and image_std=None to the processor defaults. That resolution only happens in the public preprocess() / preprocess_tensor() path.
As a result, although disabling rescaling is appropriate for an input already in [0, 1], the image does not appear to receive the required CLIP mean/std normalization before being fed into the Qwen2-VL vision encoder.
This seems potentially quite significant: the resulting HPSv3 scores, and especially the image-space reward gradients used for optimization, may not correspond to the official HPSv3 scorer.
Have you checked numerical parity between this implementation and the official
HPSv3RewardInferencer.reward(...)
on exactly the same image/prompt pairs?
It would be helpful if you could provide a simple parity test comparing the raw HPSv3 scores from the released scorer against the official implementation.
2. Figure 2 / teaser image quality
I also have a question about the qualitative results in the paper.
The samples shown in Figure 2 / the main qualitative figure appear substantially higher quality than what I obtain from the released implementation and than some of the other reported qualitative results.
Could you clarify exactly how these images were generated?
In particular, were they generated using exactly the same released checkpoint and inference configuration? It would be useful to provide the corresponding:
- prompts,
- random seeds,
- checkpoints,
- sampling steps,
- CFG/guidance settings,
- resolution, and
- any sample-selection or curation procedure.
This would make the qualitative comparison much easier to reproduce.
3. DiffusionNFT baseline outputs are consistently blurry
Finally, I am having difficulty reproducing a reasonable DiffusionNFT baseline using the released implementation.
The images generated by the provided DiffusionNFT baseline are consistently very blurry / low quality in my runs. This seems unusual enough that I am concerned there may be an implementation or inference-configuration issue with the baseline.
Could you clarify whether you verified this implementation against the original DiffusionNFT implementation?
In particular, could you provide the exact DiffusionNFT training and inference configuration used for the paper, as well as some representative baseline generations? It would also be useful to confirm that DiffusionNFT and DiffusionOPSD are evaluated using the same sampling resolution, number of steps, guidance settings, and other inference hyperparameters.
Thanks — I would appreciate any clarification on these points, especially the HPSv3 preprocessing issue, since that may affect both the reported HPSv3 evaluation numbers and optimization results.
- Lenguaje dominante
- Python
- Estrellas
- 566
- Forks
- 8
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de worldbench/DiffusionOPSD
-
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
worldbench/DiffusionOPSD#6 · 1 comentario ·
-
如何解决reward hackingAbierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
worldbench/DiffusionOPSD#4 · 1 comentario ·
-
【抄袭】Abierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 10/100
worldbench/DiffusionOPSD#3 · 1 reacción ·
-
Verify evals on Papers with CodeAbierto
Dificultad 3/5 1-2 días Aptitud para principiantes 48/100
worldbench/DiffusionOPSD#2 · 4 comentarios · 15 reacciones ·
Todos los issues de worldbench/DiffusionOPSD
Issues similares
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
Los mantenedores suelen responder en 1 día
-
https://search.utilibre.orgAbiertoinstance instance add
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
searxng/searx-instances#941 · 1 comentario ·
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día