Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Regarding repo-level data processing procedure

Aperta
#8 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
20/100
Tipo di issue
Documentazione
Chiarezza
Da chiarire
Stato di attività
Ferma
Ambito
documentation

Direzione di ricerca

Esamina le sezioni 2.2.1 e 2.2.5 e le cinque domande sulla deduplicazione, il filtraggio della qualità, la concatenazione, i parametri di MinHash e i sottografi. Il lavoro è considerato completato quando viene fornito un chiarimento procedurale autorevole oppure viene aggiornata la documentazione di riferimento; nell’issue non sono identificati file di implementazione né test.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Thank you for your excellent work! I have a few questions regarding your repository-level data processing procedure, specifically concerning sections 2.2.1 and 2.2.5.

Regarding Section 2.2.1 Preprocessing

"... implementing deduplication at both the repository and file levels. For each level, we performed exact-deduplication using SHA256 hashes of contents and near-deduplication via the MinHash algorithm. This two-tier strategy yielded two variants of the code corpus..."

  1. Is this two-tier process in sequence or independently in parallel? Explicitly, did you first concatenate files to create repo-level samples, perform repo-level exact and near-deduplication, and then split into individual files for file-level exact and near-deduplication? Or was the process structured differently?
  2. For the repo-level exact deduplication to work effectively, did you concatenate the files within each repo in a specific order (e.g., lexical order of file paths) beforehand?
  3. Did you use the same MinHash parameters for both the repo-level and file-level near-deduplication? If possible, could you share the threshold?
Regarding Section 2.2.1 Quality Filtering and Section 2.2.5 Long-Context Data for Continued Pretraining

"... This corpus supports 89 programming languages, forged into both the repository-level and file-level code data shown in Figure 2..."
"We selected high-quality repositories based on average file quality scores. For mainstream programming languages (e.g., Python, Java, and C), we implemented topological concatenation based on file dependencies. For HTML, SQL, and Shell, we used random concatenation. "

  1. I'd like to confirm my understanding of this process: For the repo-level deduplicated corpus, did you first apply the file-level quality filter to each individual file within all repositories? Then, did you calculate an average file-quality score for each repository based on its constituent files, and apply the topological/random concatenation only to the subset of selected high-score repositories?

"Each repository was mapped to a single string sequence, with exceptionally large repositories (e.g., PyTorch) being decomposed into multiple independent subgraphs to avoid oversized sequences while preserving logical coherence. "

  1. For repositories that have multiple independent subgraphs but whose concatenated total size < 32k, how did you concat these subgraphs to form the final single string sequence? Just randomly?

Thank you in advance for your time and any clarification you can provide!

Lingua principale
Nessun dato sulla lingua
Stelle
759
Fork
60
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di ByteDance-Seed/Seed-Coder

Tutte le issue di ByteDance-Seed/Seed-Coder

Issue simili

Altre issue su Documentation

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.