Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Emit a per-file and per-function fact ledger in which every byte is accounted for

Abierto
#97 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
25/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
docker, git, go, python

Línea de trabajo

Start by reading pkg/codemetrics/count_test.go and the tests/conformance/ harness, then review the proposed new pkg/ledger/ledger_test.go and tests/ledger/ structure. This issue is complete only when the versioned schemas, ledger writers, accounting invariants, aggregate reconstruction, fixtures, and documented validation checks are implemented across the listed packages and test harnesses.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

enhancement

User outcome

One command answers "what, exactly, is in here?" at the finest useful grain. Every inventoried entry gets one row, and every function that a structural analyzer measured gets one row. Every aggregate dircue reports (language percentages, role populations, line counts, component totals, the map summary) is a provable sum over those rows.

The headline number is "explained bytes": the share of bytes that carry a positive, evidence-backed classification. Every remaining byte is attributed to a named reason (unknown format, read limit, unsupported mapping, permission denied, …). Nothing disappears into an unexplained remainder.

Why

  • Today the evidence is split across separate reports: languages (the Linguist-compatible path), line metrics (scc through pkg/codemetrics), roles and formats, function metrics (the BCA worker and hotspots), and map nodes. No artifact joins them per file, so a skeptic can't check that the numbers agree, and a user can't ask ordinary questions. For example:
    • "Which generated C++ files over 2,000 lines have a function with cyclomatic complexity above 50?"
    • "How many bytes on this drive are unexplained?"
  • The mission is to marry Linguist, scc and BCA. The natural join point is the file, not three separate summaries.
  • A columnar ledger makes dircue a data source, queryable with DuckDB, SQLite, pandas or jq, without dircue having to predict every question. That's the FFmpeg principle: great defaults, plus raw access for power users.
  • An accounting identity (sum of the rows = total) is a cheap, powerful correctness invariant that makes a whole class of silent-drop bugs visible. The #94 review found several: dropped workflows, skipped YAML documents, and filename-hint misclassification.

Proposal

  • dircue files (name to be decided), or map --ledger FILE, emits the file ledger.
    • Identity: entry path, entry kind (regular, symlink, special, …), size, and the Git blob ID. The blob ID is computed in both source modes; see the tree-ID issue #101.
    • Classification:
      • role (source, test, generated, vendored, documentation, data, binary, archive, certificate/key material, …);
      • language and the Linguist strategy that decided it;
      • format and signature evidence;
      • encoding and line-ending facts.
    • Line metrics: lines, code, comments, blanks and lexical complexity, from the embedded scc processor. Use the same mapping as analyze metrics; never a second counter.
    • Structural summary when a structural analyzer ran (see #98): function count, maximum and total cyclomatic/cognitive complexity, and the maintainability index.
    • Relationships: owning component ID(s), plus deployable and interface references from the map.
    • Coverage: status and reasons per row (complete, or partial because of a read limit, unsupported mapping, and so on).
  • A function ledger (optional table): one row per measured function or space. Fields: path, qualified name if known, kind, span, and the BCA metrics (cyclomatic, cognitive, Halstead, LOC, MI, …). Keep the metric names exactly as upstream defines them. dircue never invents a quality score (see #27).
  • Output formats:
    • JSONL (default, streamable);
    • CSV;
    • Parquet, via a pure-Go writer (no cgo);
    • SQLite output is optional behind a build tag only if its binary-size cost is acceptable. DuckDB already reads JSONL, CSV and Parquet directly.
  • Documentation: each output format ships with a documented, versioned column schema and a stable sort order (by path).
  • Summaries derived from the ledger:
    • map aggregates and the legacy Linguist-compatible output become views over the same per-file facts. The legacy output stays byte-identical;
    • the map summary prints "explained bytes: 98.7% (unknown format: 1.1%, read limit: 0.2%)".
  • Worked queries: a docs/LEDGER.md cookbook of ten real questions answered with DuckDB one-liners. This is the most persuasive demonstration material for a public launch.

Tasks

  • Define the ledger column schemas (files and functions), with a JSON Schema and a Parquet schema, versioned independently of the map schema.
  • Refactor the observers to emit per-entry facts into one row builder (shares work with #105).
  • Compute the Git blob ID for directory-source files, reusing the tree-ID work in #101.
  • Implement the JSONL, CSV and Parquet writers, with deterministic ordering and floating-point formatting.
  • Add the explained-bytes accounting and its reason taxonomy, and surface it in map --summary.
  • Rebuild map aggregates as group-bys over the ledger. Prove legacy byte identity on the #75 corpus.
  • Write the docs/LEDGER.md cookbook, with queries tested in CI against a fixture ledger (DuckDB runs only in the test container, never in dircue).

Test plan

1. Unit tests

New package pkg/ledger/ (new); test file pkg/ledger/ledger_test.go (new).

  • Row construction: one sub-test per role (source, test, generated, vendored, documentation, data, binary, archive) and per entry kind (regular file, symlink, special). Assert every mandatory column is populated and non-zero where applicable.
  • Accounting identity: for every synthetic multi-file fixture, assert sum(ledger.size) == scan.TotalBytes and sum(explained) + sum(attributed_unknown_by_reason) == TotalBytes. Use table-driven tests; reuse the fixture builder from tests/stress/generate.py patterns.
  • Path edge cases: invalid UTF-8 paths (encoded as OS bytes), paths > 4096 bytes, paths with control characters or newlines, and NFC/NFD look-alikes. NUL can't occur in POSIX or Git paths, so test that the writers escape or reject NUL defensively in synthetic rows. Assert no panic, row emitted with coverage: partial, reason: path_encoding.
  • Float formatting stability: cycle Counts.complexity and maintainability_index through JSON and CSV round-trips; assert values are equal within 1e-9 and string representation is byte-identical across identical inputs.
  • Writer determinism (JSONL/CSV/Parquet): given the same in-memory rows in shuffled order, assert each writer emits identical bytes in path-sorted stable order. Run 20 permutations.
  • Schema validation: round-trip every writer's output through the documented JSON Schema and Parquet schema; assert zero errors.
  • Legacy byte identity fence: run the existing scanner on the tests/conformance/ corpus fixtures in both source modes, collect language-percentage aggregates, then recompute the same aggregates from a ledger over the same fixture. Assert byte-identical output in the Linguist-compatible path. Tie to the existing make conformance target.
  • Existing tests to extend: pkg/codemetrics/count_test.go — add a case asserting Count() results match the ledger's lines/code/comments/blanks columns exactly for the same content bytes.
2. Independent oracles

Harness: tests/ledger/ (new), following the Docker pattern in tests/conformance/.

  • Linguist oracle: pinned dircue-linguist:9.7.0 image (same as tests/conformance/run.py). Run github-linguist --breakdown --json on each fixture. Compare ledger language column and byte counts. Acceptable tolerance: exact (language) and ±0 bytes (size). Log mismatches to results/linguist-differential.json. Tie to #84.
  • scc oracle: pinned scc binary (same version as pkg/codemetrics/count.go:EngineVersion), run as scc --by-file --format json. Compare code/comment/blank/complexity per file. Zero tolerance on same-language, same-scope files; document and commit known deviations in tests/ledger/DISCREPANCIES.md (new, modelled on tests/metrics/results/).
  • BCA worker oracle: run the structural worker from prototypes/structural/worker on fixture repos with the --network none Docker flag. Compare function-ledger cyclomatic/cognitive columns. Document the BCA upstream version in tests/ledger/PROVENANCE.md (new).
  • Container is always run with --network none --read-only (matching tests/conformance/Dockerfile pattern).
3. Hand-labeled ground truth and fixtures
  • Add 10 new YAML fixtures under tests/ledger/fixtures/ (new directory; fixtures are small, synthetic, committed to git).
    • Cover: multi-language repo, all-unknown repo, repo with read-limit truncation, symlink target, binary-only, generated-only, mixed explained/unknown.
    • Each fixture has a hand-verified .expected.jsonl (committed as gzip; per #88 rule on bulky binary data, the Parquet variants are generated by the CI harness from the JSONL, never committed).
    • Precision/recall gate (#75): 100% precision on language, role and explained_bytes columns for the committed fixtures; explained bytes and every attributed-unknown reason must equal the hand-labeled expected values exactly for each fixture, including the all-unknown fixture, where explained bytes is 0.
  • Do not commit kernel or large corpora; the harness fetches them on demand (network-fetched, not in CI).
4. Metamorphic invariants (#85)
  • Worker-count invariant: ledger rows are byte-identical with --workers 1, --workers 4, and the default. Add as a CI sub-test in the test job (it's fast on small fixtures).
  • Source-mode invariant: for fixtures with identical content in both Git and directory mode (matching blob IDs), ledger rows are byte-identical column-for-column except blob_id (expected to differ or be absent in directory mode before #101 lands).
  • Aggregate reconstruction: every aggregate from dircue map must equal the group-by over the corresponding ledger rows (enforced as a property in pkg/ledger/ledger_test.go).
  • Explained + unknown = total: checked per fixture, across all worker counts and source modes.
5. Fuzz targets (#87)

New file pkg/ledger/fuzz_test.go (new):

  • FuzzLedgerRowRoundtrip(f *testing.F): seeds with the committed JSONL fixtures; checks that deserializing a serialized row is identity and that sum(row.size) invariant holds.
  • FuzzWriterDeterminism(f *testing.F): seeds with small row slices; checks JSONL writer emits the same bytes on two calls.

These are on-demand only (never run as campaigns in CI per #87).

6. Mutation-testing focus (#86)

Priority mutation targets:

  • The accounting accumulator in the ledger row builder (the explained + attributed_unknown = total identity — a missing branch here silently hides dropped files).
  • The path-sort comparator for writer determinism.
  • The format-decision gate that sets coverage: partial vs coverage: complete.

Use go-mutesting or equivalent on pkg/ledger/ only; require that all three mutant classes are killed by existing unit tests before merge.

7. Performance and resource checks (#89/#90)
  • Deterministic counter gate (PR CI): expose a test-only counter in pkg/ledger/ that counts bytes passed through each writer stage. Assert bytes_written_jsonl == sum(row.size_on_disk) within a tolerance. This is noise-free and belongs in the test job.
  • BenchmarkLedgerWrite (new): in pkg/ledger/ledger_bench_test.go (new). Bench the JSONL, CSV and Parquet writers on 10k synthetic rows. Run via make bench. Assert no allocation regression > 10% in CI with benchstat (on-demand, workflow_dispatch only, per #90 tiers).
  • Streaming check: assert that peak RSS during a 92k-row synthetic run stays below a documented threshold (TBD with #105 work). Add to tests/resources/ (existing cgroup-based harness), gated behind workflow_dispatch.
8. E2E script

tests/ledger/e2e.py (new). Structured JSONL log fields per invocation:

{"command": [...], "version": "...", "wall_s": 1.23, "rss_mib": 45.6,
 "rows": 12345, "explained_bytes": 99.7, "mismatches": 0,
 "stage": "linguist|scc|bca|dircue-ledger", "source_mode": "git|directory"}

Steps:

  1. Run dircue ledger command on the pinned corpus (fetched by harness, not by dircue).
  2. Run linguist --breakdown in the dircue-linguist:9.7.0 container (--network none).
  3. Run scc --by-file --format json at the pinned version.
  4. Cross-check language and byte columns; cross-check line-count columns.
  5. Emit a human-readable summary table (language, dircue bytes, linguist bytes, delta; line columns, scc vs dircue delta).
  6. Exit non-zero on any mismatch outside the committed DISCREPANCIES.md list.

Invocation: make ledger-e2e (new target) or workflow_dispatch on the ci.yml workflow. Never run on schedule:.

9. CI placement
What Where Trigger
Unit tests (pkg/ledger/...) existing test job in ci.yml every PR and push
Worker-count + source-mode metamorphic invariants existing test job every PR and push
Legacy byte-identity fence (ledger vs map aggregates) new step in linguist-conformance job push + non-draft PR
Linguist + scc oracle differential (tests/ledger/e2e.py) new ledger-conformance job (new) in ci.yml push + non-draft PR, same pattern as metrics-conformance
Benchmark regression (BenchmarkLedgerWrite) workflow_dispatch + make bench on-demand only
RSS / streaming stress (tests/resources/) workflow_dispatch + make target on-demand only

No schedule: trigger anywhere.

10. Regression fences
  • Silent-drop regression (#94 review bugs): add a test in pkg/ledger/ledger_test.go with a synthetic fixture containing a YAML file with multiple documents, a workflow YAML, and a file with a misleading extension. Assert all entries appear in the ledger and sum(size) == expected_total.
  • Unknown-not-absent contract: assert that every row with coverage: partial, reason: unknown_format still has a non-empty path and non-zero size; the language field is "" not "unknown" (the sentinel must be the empty string, not the word "unknown", to avoid the "unknown rendered as absent" contract breach).

Risks

  • Size. Output grows with the file count: about 92k rows on the kernel is fine, but a 10M-file disk needs streaming. Never buffer the whole ledger.
  • Parquet dependency. Parquet adds a dependency; audit its size and determinism (row-group sizing, statistics).
  • Contract scope. Column names become a contract. Version them from the start, and mark experimental columns.

Success criteria

  • Ledger aggregates reproduce every existing aggregate exactly.
  • Explained bytes is reported with a complete reason breakdown.
  • Cross-tool parity scripts pass on the corpus.
  • The cookbook queries run.

Related: #6, #27, #38, #74, #75, #85, and #12 (scc parity output). Related idea issues: #98, #101, #105, #104, #107, #99.

Sequencing: Foundation for #99, #104, #106 and #107. It can start once #94 merges. It benefits from #105 and #101 but doesn't require them: blob IDs can start in Git mode only.


This proposal came out of an idea-wizard session on 2026-09-23. It isn't a release commitment, and the milestone is deliberately unset so the owner can decide per idea. It must preserve dircue's contracts:

  • deterministic, offline output, with no execution of inspected content;
  • the legacy Linguist-compatible CLI and output unchanged;
  • unknown never rendered as absent;
  • no provider findings or verdicts imported;
  • CI event-driven only (no schedule:).
Lenguaje dominante
Go
Estrellas
0
Forks
0
Merge medio
5 h 8 min
PR fusionados (30 d)
54

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de war-and-code/dircue

Todos los issues de war-and-code/dircue

Issues similares

Más issues de Go

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.