Questions about HPSv3 scorer correctness, Figure 2 sample quality, and DiffusionNFT baseline
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 35/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- machine-learning, testing-qa
Direzione di ricerca
Start with src/hpsv3_scorer.py and compare its preprocessing and scores against the official HPSv3RewardInferencer.reward(...) on identical image/prompt pairs. Then trace the Figure 2 generation configuration and the released DiffusionNFT baseline entry points. Done means documenting reproducible prompts, seeds, checkpoints, sampling settings, and representative baseline results, or confirming the reported implementation discrepancy.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hi, thanks for releasing the code. I have a few questions regarding the implementation and the reported qualitative results.
1. HPSv3 scorer implementation appears to miss image normalization
I think there may be a serious correctness issue in the current HPSv3 scorer.
In src/hpsv3_scorer.py, the differentiable preprocessing path calls:
self.ip._preprocess(images01[i:i + 1], do_rescale=False)
However, in the official HPSv3 differentiable image processor, _preprocess() does not resolve do_normalize=None, image_mean=None, and image_std=None to the processor defaults. That resolution only happens in the public preprocess() / preprocess_tensor() path.
As a result, although disabling rescaling is appropriate for an input already in [0, 1], the image does not appear to receive the required CLIP mean/std normalization before being fed into the Qwen2-VL vision encoder.
This seems potentially quite significant: the resulting HPSv3 scores, and especially the image-space reward gradients used for optimization, may not correspond to the official HPSv3 scorer.
Have you checked numerical parity between this implementation and the official
HPSv3RewardInferencer.reward(...)
on exactly the same image/prompt pairs?
It would be helpful if you could provide a simple parity test comparing the raw HPSv3 scores from the released scorer against the official implementation.
2. Figure 2 / teaser image quality
I also have a question about the qualitative results in the paper.
The samples shown in Figure 2 / the main qualitative figure appear substantially higher quality than what I obtain from the released implementation and than some of the other reported qualitative results.
Could you clarify exactly how these images were generated?
In particular, were they generated using exactly the same released checkpoint and inference configuration? It would be useful to provide the corresponding:
- prompts,
- random seeds,
- checkpoints,
- sampling steps,
- CFG/guidance settings,
- resolution, and
- any sample-selection or curation procedure.
This would make the qualitative comparison much easier to reproduce.
3. DiffusionNFT baseline outputs are consistently blurry
Finally, I am having difficulty reproducing a reasonable DiffusionNFT baseline using the released implementation.
The images generated by the provided DiffusionNFT baseline are consistently very blurry / low quality in my runs. This seems unusual enough that I am concerned there may be an implementation or inference-configuration issue with the baseline.
Could you clarify whether you verified this implementation against the original DiffusionNFT implementation?
In particular, could you provide the exact DiffusionNFT training and inference configuration used for the paper, as well as some representative baseline generations? It would also be useful to confirm that DiffusionNFT and DiffusionOPSD are evaluated using the same sampling resolution, number of steps, guidance settings, and other inference hyperparameters.
Thanks — I would appreciate any clarification on these points, especially the HPSv3 preprocessing issue, since that may affect both the reported HPSv3 evaluation numbers and optimization results.
- Lingua principale
- Python
- Stelle
- 566
- Fork
- 8
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di worldbench/DiffusionOPSD
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
worldbench/DiffusionOPSD#6 · 1 commento ·
-
如何解决reward hackingAperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
worldbench/DiffusionOPSD#4 · 1 commento ·
-
【抄袭】Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 10/100
worldbench/DiffusionOPSD#3 · 1 reazione ·
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 48/100
worldbench/DiffusionOPSD#2 · 4 commenti · 15 reazioni ·
Tutte le issue di worldbench/DiffusionOPSD
Issue simili
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedApertaworkflow
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
I maintainer di solito rispondono entro 1 giorno
-
New Submission: TropWATERApertametadata submission
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
canonical/content-cache-operator#163 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
[submission]Apertasubmission
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 65/100
leanprover/lean-eval-submissions#1852 ·
I maintainer di solito rispondono entro 1 giorno