Emit a per-file and per-function fact ledger in which every byte is accounted for
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
Línea de trabajo
Start by reading pkg/codemetrics/count_test.go and the tests/conformance/ harness, then review the proposed new pkg/ledger/ledger_test.go and tests/ledger/ structure. This issue is complete only when the versioned schemas, ledger writers, accounting invariants, aggregate reconstruction, fixtures, and documented validation checks are implemented across the listed packages and test harnesses.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
User outcome
One command answers "what, exactly, is in here?" at the finest useful grain. Every inventoried entry gets one row, and every function that a structural analyzer measured gets one row. Every aggregate dircue reports (language percentages, role populations, line counts, component totals, the map summary) is a provable sum over those rows.
The headline number is "explained bytes": the share of bytes that carry a positive, evidence-backed classification. Every remaining byte is attributed to a named reason (unknown format, read limit, unsupported mapping, permission denied, …). Nothing disappears into an unexplained remainder.
Why
- Today the evidence is split across separate reports: languages (the Linguist-compatible path), line metrics (scc through
pkg/codemetrics), roles and formats, function metrics (the BCA worker and hotspots), and map nodes. No artifact joins them per file, so a skeptic can't check that the numbers agree, and a user can't ask ordinary questions. For example:- "Which generated C++ files over 2,000 lines have a function with cyclomatic complexity above 50?"
- "How many bytes on this drive are unexplained?"
- The mission is to marry Linguist, scc and BCA. The natural join point is the file, not three separate summaries.
- A columnar ledger makes dircue a data source, queryable with DuckDB, SQLite, pandas or
jq, without dircue having to predict every question. That's the FFmpeg principle: great defaults, plus raw access for power users. - An accounting identity (sum of the rows = total) is a cheap, powerful correctness invariant that makes a whole class of silent-drop bugs visible. The #94 review found several: dropped workflows, skipped YAML documents, and filename-hint misclassification.
Proposal
dircue files(name to be decided), ormap --ledger FILE, emits the file ledger.- Identity: entry path, entry kind (regular, symlink, special, …), size, and the Git blob ID. The blob ID is computed in both source modes; see the tree-ID issue #101.
- Classification:
- role (source, test, generated, vendored, documentation, data, binary, archive, certificate/key material, …);
- language and the Linguist strategy that decided it;
- format and signature evidence;
- encoding and line-ending facts.
- Line metrics: lines, code, comments, blanks and lexical complexity, from the embedded scc processor. Use the same mapping as
analyze metrics; never a second counter. - Structural summary when a structural analyzer ran (see #98): function count, maximum and total cyclomatic/cognitive complexity, and the maintainability index.
- Relationships: owning component ID(s), plus deployable and interface references from the map.
- Coverage: status and reasons per row (
complete, orpartialbecause of a read limit, unsupported mapping, and so on).
- A function ledger (optional table): one row per measured function or space. Fields: path, qualified name if known, kind, span, and the BCA metrics (cyclomatic, cognitive, Halstead, LOC, MI, …). Keep the metric names exactly as upstream defines them. dircue never invents a quality score (see #27).
- Output formats:
- JSONL (default, streamable);
- CSV;
- Parquet, via a pure-Go writer (no cgo);
- SQLite output is optional behind a build tag only if its binary-size cost is acceptable. DuckDB already reads JSONL, CSV and Parquet directly.
- Documentation: each output format ships with a documented, versioned column schema and a stable sort order (by path).
- Summaries derived from the ledger:
- map aggregates and the legacy Linguist-compatible output become views over the same per-file facts. The legacy output stays byte-identical;
- the map summary prints "explained bytes: 98.7% (unknown format: 1.1%, read limit: 0.2%)".
- Worked queries: a
docs/LEDGER.mdcookbook of ten real questions answered with DuckDB one-liners. This is the most persuasive demonstration material for a public launch.
Tasks
- Define the ledger column schemas (files and functions), with a JSON Schema and a Parquet schema, versioned independently of the map schema.
- Refactor the observers to emit per-entry facts into one row builder (shares work with #105).
- Compute the Git blob ID for directory-source files, reusing the tree-ID work in #101.
- Implement the JSONL, CSV and Parquet writers, with deterministic ordering and floating-point formatting.
- Add the explained-bytes accounting and its reason taxonomy, and surface it in
map --summary. - Rebuild map aggregates as group-bys over the ledger. Prove legacy byte identity on the #75 corpus.
- Write the
docs/LEDGER.mdcookbook, with queries tested in CI against a fixture ledger (DuckDB runs only in the test container, never in dircue).
Test plan
1. Unit tests
New package pkg/ledger/ (new); test file pkg/ledger/ledger_test.go (new).
- Row construction: one sub-test per role (
source,test,generated,vendored,documentation,data,binary,archive) and per entry kind (regular file, symlink, special). Assert every mandatory column is populated and non-zero where applicable. - Accounting identity: for every synthetic multi-file fixture, assert
sum(ledger.size) == scan.TotalBytesandsum(explained) + sum(attributed_unknown_by_reason) == TotalBytes. Use table-driven tests; reuse the fixture builder fromtests/stress/generate.pypatterns. - Path edge cases: invalid UTF-8 paths (encoded as OS bytes), paths > 4096 bytes, paths with control characters or newlines, and NFC/NFD look-alikes. NUL can't occur in POSIX or Git paths, so test that the writers escape or reject NUL defensively in synthetic rows. Assert no panic, row emitted with
coverage: partial, reason: path_encoding. - Float formatting stability: cycle
Counts.complexityandmaintainability_indexthrough JSON and CSV round-trips; assert values are equal within1e-9and string representation is byte-identical across identical inputs. - Writer determinism (JSONL/CSV/Parquet): given the same in-memory rows in shuffled order, assert each writer emits identical bytes in path-sorted stable order. Run 20 permutations.
- Schema validation: round-trip every writer's output through the documented JSON Schema and Parquet schema; assert zero errors.
- Legacy byte identity fence: run the existing scanner on the
tests/conformance/corpus fixtures in both source modes, collect language-percentage aggregates, then recompute the same aggregates from a ledger over the same fixture. Assert byte-identical output in the Linguist-compatible path. Tie to the existingmake conformancetarget. - Existing tests to extend:
pkg/codemetrics/count_test.go— add a case assertingCount()results match the ledger'slines/code/comments/blankscolumns exactly for the same content bytes.
2. Independent oracles
Harness: tests/ledger/ (new), following the Docker pattern in tests/conformance/.
- Linguist oracle: pinned
dircue-linguist:9.7.0image (same astests/conformance/run.py). Rungithub-linguist --breakdown --jsonon each fixture. Compare ledgerlanguagecolumn and byte counts. Acceptable tolerance: exact (language) and ±0 bytes (size). Log mismatches toresults/linguist-differential.json. Tie to #84. - scc oracle: pinned scc binary (same version as
pkg/codemetrics/count.go:EngineVersion), run asscc --by-file --format json. Comparecode/comment/blank/complexityper file. Zero tolerance on same-language, same-scope files; document and commit known deviations intests/ledger/DISCREPANCIES.md(new, modelled ontests/metrics/results/). - BCA worker oracle: run the structural worker from
prototypes/structural/workeron fixture repos with the--network noneDocker flag. Compare function-ledgercyclomatic/cognitivecolumns. Document the BCA upstream version intests/ledger/PROVENANCE.md(new). - Container is always run with
--network none --read-only(matchingtests/conformance/Dockerfilepattern).
3. Hand-labeled ground truth and fixtures
- Add 10 new YAML fixtures under
tests/ledger/fixtures/(new directory; fixtures are small, synthetic, committed to git).- Cover: multi-language repo, all-unknown repo, repo with read-limit truncation, symlink target, binary-only, generated-only, mixed explained/unknown.
- Each fixture has a hand-verified
.expected.jsonl(committed as gzip; per #88 rule on bulky binary data, the Parquet variants are generated by the CI harness from the JSONL, never committed). - Precision/recall gate (#75): 100% precision on
language,roleandexplained_bytescolumns for the committed fixtures; explained bytes and every attributed-unknown reason must equal the hand-labeled expected values exactly for each fixture, including the all-unknown fixture, where explained bytes is 0.
- Do not commit kernel or large corpora; the harness fetches them on demand (network-fetched, not in CI).
4. Metamorphic invariants (#85)
- Worker-count invariant: ledger rows are byte-identical with
--workers 1,--workers 4, and the default. Add as a CI sub-test in thetestjob (it's fast on small fixtures). - Source-mode invariant: for fixtures with identical content in both Git and directory mode (matching blob IDs), ledger rows are byte-identical column-for-column except
blob_id(expected to differ or be absent in directory mode before #101 lands). - Aggregate reconstruction: every aggregate from
dircue mapmust equal the group-by over the corresponding ledger rows (enforced as a property inpkg/ledger/ledger_test.go). - Explained + unknown = total: checked per fixture, across all worker counts and source modes.
5. Fuzz targets (#87)
New file pkg/ledger/fuzz_test.go (new):
FuzzLedgerRowRoundtrip(f *testing.F): seeds with the committed JSONL fixtures; checks that deserializing a serialized row is identity and thatsum(row.size)invariant holds.FuzzWriterDeterminism(f *testing.F): seeds with small row slices; checks JSONL writer emits the same bytes on two calls.
These are on-demand only (never run as campaigns in CI per #87).
6. Mutation-testing focus (#86)
Priority mutation targets:
- The accounting accumulator in the ledger row builder (the
explained + attributed_unknown = totalidentity — a missing branch here silently hides dropped files). - The path-sort comparator for writer determinism.
- The format-decision gate that sets
coverage: partialvscoverage: complete.
Use go-mutesting or equivalent on pkg/ledger/ only; require that all three mutant classes are killed by existing unit tests before merge.
7. Performance and resource checks (#89/#90)
- Deterministic counter gate (PR CI): expose a test-only counter in
pkg/ledger/that counts bytes passed through each writer stage. Assertbytes_written_jsonl == sum(row.size_on_disk)within a tolerance. This is noise-free and belongs in thetestjob. BenchmarkLedgerWrite(new): inpkg/ledger/ledger_bench_test.go(new). Bench the JSONL, CSV and Parquet writers on 10k synthetic rows. Run viamake bench. Assert no allocation regression > 10% in CI withbenchstat(on-demand,workflow_dispatchonly, per #90 tiers).- Streaming check: assert that peak RSS during a 92k-row synthetic run stays below a documented threshold (TBD with #105 work). Add to
tests/resources/(existing cgroup-based harness), gated behindworkflow_dispatch.
8. E2E script
tests/ledger/e2e.py (new). Structured JSONL log fields per invocation:
{"command": [...], "version": "...", "wall_s": 1.23, "rss_mib": 45.6,
"rows": 12345, "explained_bytes": 99.7, "mismatches": 0,
"stage": "linguist|scc|bca|dircue-ledger", "source_mode": "git|directory"}
Steps:
- Run
dircueledger command on the pinned corpus (fetched by harness, not by dircue). - Run
linguist --breakdownin thedircue-linguist:9.7.0container (--network none). - Run
scc --by-file --format jsonat the pinned version. - Cross-check language and byte columns; cross-check line-count columns.
- Emit a human-readable summary table (language, dircue bytes, linguist bytes, delta; line columns, scc vs dircue delta).
- Exit non-zero on any mismatch outside the committed
DISCREPANCIES.mdlist.
Invocation: make ledger-e2e (new target) or workflow_dispatch on the ci.yml workflow. Never run on schedule:.
9. CI placement
| What | Where | Trigger |
|---|---|---|
Unit tests (pkg/ledger/...) |
existing test job in ci.yml |
every PR and push |
| Worker-count + source-mode metamorphic invariants | existing test job |
every PR and push |
| Legacy byte-identity fence (ledger vs map aggregates) | new step in linguist-conformance job |
push + non-draft PR |
Linguist + scc oracle differential (tests/ledger/e2e.py) |
new ledger-conformance job (new) in ci.yml |
push + non-draft PR, same pattern as metrics-conformance |
Benchmark regression (BenchmarkLedgerWrite) |
workflow_dispatch + make bench |
on-demand only |
RSS / streaming stress (tests/resources/) |
workflow_dispatch + make target |
on-demand only |
No schedule: trigger anywhere.
10. Regression fences
- Silent-drop regression (#94 review bugs): add a test in
pkg/ledger/ledger_test.gowith a synthetic fixture containing a YAML file with multiple documents, a workflow YAML, and a file with a misleading extension. Assert all entries appear in the ledger andsum(size) == expected_total. - Unknown-not-absent contract: assert that every row with
coverage: partial, reason: unknown_formatstill has a non-emptypathand non-zerosize; thelanguagefield is""not"unknown"(the sentinel must be the empty string, not the word "unknown", to avoid the "unknown rendered as absent" contract breach).
Risks
- Size. Output grows with the file count: about 92k rows on the kernel is fine, but a 10M-file disk needs streaming. Never buffer the whole ledger.
- Parquet dependency. Parquet adds a dependency; audit its size and determinism (row-group sizing, statistics).
- Contract scope. Column names become a contract. Version them from the start, and mark experimental columns.
Success criteria
- Ledger aggregates reproduce every existing aggregate exactly.
- Explained bytes is reported with a complete reason breakdown.
- Cross-tool parity scripts pass on the corpus.
- The cookbook queries run.
Related: #6, #27, #38, #74, #75, #85, and #12 (scc parity output). Related idea issues: #98, #101, #105, #104, #107, #99.
Sequencing: Foundation for #99, #104, #106 and #107. It can start once #94 merges. It benefits from #105 and #101 but doesn't require them: blob IDs can start in Git mode only.
This proposal came out of an idea-wizard session on 2026-09-23. It isn't a release commitment, and the milestone is deliberately unset so the owner can decide per idea. It must preserve dircue's contracts:
- deterministic, offline output, with no execution of inspected content;
- the legacy Linguist-compatible CLI and output unchanged;
unknownnever rendered as absent;- no provider findings or verdicts imported;
- CI event-driven only (no
schedule:).
- Lenguaje dominante
- Go
- Estrellas
- 0
- Forks
- 0
- Merge medio
- 5 h 8 min
- PR fusionados (30 d)
- 54
Preparar el entorno
- Incluye un Dockerfile o un archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de war-and-code/dircue
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
war-and-code/dircue#200 ·
Los mantenedores suelen responder en 1 día
-
follow-up
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
war-and-code/dircue#179 ·
Los mantenedores suelen responder en 1 día
-
follow-up
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
war-and-code/dircue#178 ·
Los mantenedores suelen responder en 1 día
-
enhancement follow-up
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
war-and-code/dircue#172 ·
Los mantenedores suelen responder en 1 día
-
enhancement follow-up
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
war-and-code/dircue#171 ·
Los mantenedores suelen responder en 1 día
Todos los issues de war-and-code/dircue
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
agent-research agent-review-finding chore
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100
jordansmall/spindrift#4922 ·
Los mantenedores suelen responder en 1 día
-
gcsartifact: deleting a missing version returns an errorPosiblemente ocupada @ktsoator la tomó hoy. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 2 días
-
govulncheck
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
Los mantenedores suelen responder en 1 día
-
Change wording for init command success messagePosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día