Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Questions about HPSv3 scorer correctness, Figure 2 sample quality, and DiffusionNFT baseline

Aperta
#7 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
35/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
python

Direzione di ricerca

Start with src/hpsv3_scorer.py and compare its preprocessing and scores against the official HPSv3RewardInferencer.reward(...) on identical image/prompt pairs. Then trace the Figure 2 generation configuration and the released DiffusionNFT baseline entry points. Done means documenting reproducible prompts, seeds, checkpoints, sampling settings, and representative baseline results, or confirming the reported implementation discrepancy.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Hi, thanks for releasing the code. I have a few questions regarding the implementation and the reported qualitative results.

1. HPSv3 scorer implementation appears to miss image normalization

I think there may be a serious correctness issue in the current HPSv3 scorer.

In src/hpsv3_scorer.py, the differentiable preprocessing path calls:

self.ip._preprocess(images01[i:i + 1], do_rescale=False)

However, in the official HPSv3 differentiable image processor, _preprocess() does not resolve do_normalize=None, image_mean=None, and image_std=None to the processor defaults. That resolution only happens in the public preprocess() / preprocess_tensor() path.

As a result, although disabling rescaling is appropriate for an input already in [0, 1], the image does not appear to receive the required CLIP mean/std normalization before being fed into the Qwen2-VL vision encoder.

This seems potentially quite significant: the resulting HPSv3 scores, and especially the image-space reward gradients used for optimization, may not correspond to the official HPSv3 scorer.

Have you checked numerical parity between this implementation and the official

HPSv3RewardInferencer.reward(...)

on exactly the same image/prompt pairs?

It would be helpful if you could provide a simple parity test comparing the raw HPSv3 scores from the released scorer against the official implementation.

2. Figure 2 / teaser image quality

I also have a question about the qualitative results in the paper.

The samples shown in Figure 2 / the main qualitative figure appear substantially higher quality than what I obtain from the released implementation and than some of the other reported qualitative results.

Could you clarify exactly how these images were generated?

In particular, were they generated using exactly the same released checkpoint and inference configuration? It would be useful to provide the corresponding:

  • prompts,
  • random seeds,
  • checkpoints,
  • sampling steps,
  • CFG/guidance settings,
  • resolution, and
  • any sample-selection or curation procedure.

This would make the qualitative comparison much easier to reproduce.

3. DiffusionNFT baseline outputs are consistently blurry

Finally, I am having difficulty reproducing a reasonable DiffusionNFT baseline using the released implementation.

The images generated by the provided DiffusionNFT baseline are consistently very blurry / low quality in my runs. This seems unusual enough that I am concerned there may be an implementation or inference-configuration issue with the baseline.

Could you clarify whether you verified this implementation against the original DiffusionNFT implementation?

In particular, could you provide the exact DiffusionNFT training and inference configuration used for the paper, as well as some representative baseline generations? It would also be useful to confirm that DiffusionNFT and DiffusionOPSD are evaluated using the same sampling resolution, number of steps, guidance settings, and other inference hyperparameters.

Thanks — I would appreciate any clarification on these points, especially the HPSv3 preprocessing issue, since that may affect both the reported HPSv3 evaluation numbers and optimization results.

Lingua principale
Python
Stelle
566
Fork
8
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di worldbench/DiffusionOPSD

Tutte le issue di worldbench/DiffusionOPSD

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.