Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

CRC of objects with references is not comparable across separate databases

Ouverte
#74 2 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Les mainteneurs répondent en général sous 1 jour

Personne n'a encore pris cette issue.

Évaluation

Difficulté
5/5
Temps estimé
Plus d'une semaine
Accessibilité débutants
35/100
Type d'issue
Bug
Clarté
Plutôt claire
Activité
Active
Stack technique
csharp, unity

Piste de recherche

Commencez par lire PPtrAndCrcProcessor.ExtractPPtr et ObjectIdProvider.GetId afin de retracer la manière dont les références résolues entrent dans le CRC, puis examinez sf.ExternalReferences et le workflow view_potential_duplicates. Mettez en place des tests pour vérifier l’égalité du CRC entre bases de données et la détection des doublons entre bundles ; le travail est terminé lorsque les objets référencés sont comparés de manière cohérente sans affaiblir la déduplication.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

Summary

objects.crc32 is meant to be a content fingerprint, and the comparing-builds workflow expects you can analyze two builds into two separate databases and diff CRCs to find which objects changed. That works for leaf assets (Texture2D/Mesh/AudioClip — no references), but it is broken for any object that contains references (Materials, prefabs/GameObjects, MonoBehaviours, etc.): identical content produces different CRCs in two separate analyze runs.

Cause

When PPtrAndCrcProcessor.ExtractPPtr folds a reference into the CRC, it uses the resolved analyzer/database object id returned by the callback, not the PPtr's own identity:

var refId = m_Callback(m_ObjectId, fileId, pathId, ...);   // analyzer db id
m_Crc32 = Crc32Algorithm.Append(m_Crc32, <refId bytes>);

That id comes from ObjectIdProvider.GetId((m_LocalToDbFileId[fileId], pathId)), and both the serialized-file id and the object id are assigned sequentially per analyze run. So the same logical object gets different ids in db1 vs db2 → different CRC for identical content → cross-database comparison reports spurious differences for every object that has references.

Why we can't just hash the raw PPtr (the tradeoff)

The obvious fix is to hash the raw on-disk PPtr (fileId + pathId) instead of the resolved id. But the resolved id is currently what makes within-database duplicate detection (view_potential_duplicates) work across bundles: two copies of the same object in different bundles reference the same target, and resolving through m_LocalToDbFileId (keyed by filename) + pathId normalizes them to the same id → same CRC → detected as duplicates.

fileId is a local index into a serialized file's external-reference list, so two copies of an object in different bundles can have different fileId values for the same target. Hashing the raw PPtr would therefore weaken duplicate detection. Deduplication is an important feature and is probably not well covered by tests yet, so we don't want to risk regressing it.

Options to evaluate

  1. Raw PPtr (fileId + pathId) — simplest; fixes cross-db comparison in the common case; risks weakening view_potential_duplicates (local fileId differs between bundles).
  2. Stable target identity + pathId — resolve fileId to a stable identifier for the target file and hash that + pathId, so it is independent of the local index. This fixes cross-db comparison AND preserves cross-bundle duplicate detection, but the "stable identifier" differs by source:
    • Build output external references carry a path (e.g. archive:/CAB-...), not a GUID.
    • Editor / Library references carry a GUID (the source asset's GUID).
      So the CRC needs to mix in whichever of ExternalReference.Path / ExternalReference.Guid is populated (and a fixed marker for local refs, fileId == 0). Relies on those fields being present and stable.
      More code: thread the external-reference info from sf.ExternalReferences into the CRC.
  3. Status quo — cross-db comparison stays broken for referenced objects.

Prerequisite

Add test coverage for view_potential_duplicates / cross-bundle deduplication before changing the CRC, so a fix can be validated to not regress it.

Context

Discovered while reviewing #73 / #70. Note that this is independent of the CRC changes made there (the ManagedReferenceData size fix, the ComputeCRC chunking fix, and the cah:/ stream hashing) — those also change CRC values vs. older tool versions, so CRCs are not comparable across tool versions regardless.

Related: #44 (refs table).

Langage dominant
C#
Étoiles
825
Forks
72
Merge moyen
20 h 16 min
PR mergées (30 j)
13

Préparer son environnement

Ce projet ne fournit ni conteneur de développement, ni Dockerfile, ni guide de contribution : l'installation est à votre charge. Commencez par son README, et consultez notre guide de la première contribution pour les étapes générales.

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de Unity-Technologies/UnityDataTools

Toutes les issues de Unity-Technologies/UnityDataTools

Issues similaires

Plus d'issues C#

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.