Feature request: opt-in crawl source receipt for RAG ingestion
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- python
- Domain
- backend-api-design
Research direction
Start with crawl4ai/models.py, especially the CrawlResult fields, then trace cache metadata and Docker API serialization. Define the receipt's URL, fetch-time, output-kind, digest, configuration fingerprint, cache, and outcome semantics without exposing secrets. Done means fresh, redirected, cache-hit, raw/fit, changed-configuration, serialization, secret-exclusion, and failed/partial-crawl tests satisfy the acceptance criteria while the default response remains compatible.
Written by the indexing model from the issue text.
Description
Problem / use case
Crawl4AI is positioned as a web-to-Markdown input for RAG and agent data pipelines. A typical ingestion job needs to answer three questions for every stored document: which source and crawl produced it, did the usable text actually change, and can the same processed artifact be traced later? Those answers drive re-embedding cost, stale-document replacement, and source attribution in downstream answers.
Today CrawlResult already exposes useful pieces such as url, redirected_url, response_headers, cached_at, cache_status, and Markdown (current model). The cache also has internal content hashes. But the result does not provide one documented, serialization-friendly record tying the selected output text to a stable digest, its actual fetch time, and its source. Every RAG integrator must assemble this independently, and a cache hit can otherwise be mistaken for a fresh fetch.
Proposed module: optional crawl source receipt
Add an opt-in, small source_receipt attached to each successful CrawlResult (and available through the Docker API serialization). Suggested fields:
- Requested URL and final URL, reusing existing redirect information.
fetched_at: time the underlying page was fetched; on a cache hit preserve the original fetch time. Optionally includeserved_atseparately.- Output kind (
raw_markdownorfit_markdown) and a versioned SHA-256 digest of the exact UTF-8 text returned for that kind. Make the normalization/encoding rule explicit. An unchanged digest lets an ingestion job skip embedding; a changed digest triggers replacement. - Crawl4AI version and a fingerprint of content-affecting, non-secret processing options, so a changed extraction configuration does not look like unchanged content.
- Existing cache status, and a clear distinction between successful extraction and fallback/partial output.
The default path can remain unchanged, with receipt generation behind a config flag. Please exclude cookies, Authorization headers, tokens, proxy credentials, and arbitrary response headers from the receipt/fingerprint. Consumers can retain the receipt next to their vector-store document ID; this does not require Crawl4AI to own a vector store or scheduler.
Acceptance criteria
- Fresh crawl, redirect, and cache-hit examples show accurate URL and fetch-time semantics.
- Identical selected Markdown + processing configuration yields the same digest; a text or relevant configuration change changes it.
- Raw and fit Markdown cannot accidentally share a digest namespace.
- Library and Docker API outputs serialize the receipt consistently; existing responses remain compatible when the option is off.
- Tests cover secret exclusion and a failed/partial crawl.
I checked main and develop result fields, the roadmap, and issue/PR searches for provenance, source lineage, content fingerprints, and crawl manifests. I found a planned incremental embedding index in the roadmap, but not this per-crawl receipt. The receipt could serve that later index while remaining useful to external ingestion pipelines.
- Dominant language
- Python
- Stars
- 84.5k
- Forks
- 8.7k
- Avg merge
- 3d 20h
- Merged PRs (30d)
- 15
Getting set up
- Ships a Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from unclecode/crawl4ai
-
🐞 Bug 🩺 Needs Triage
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
unclecode/crawl4ai#2319 · 2 comments ·
Maintainers usually reply within 1 day
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runOpen
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
unclecode/crawl4ai#2309 · 2 comments ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 84/100
unclecode/crawl4ai#2147 · 3 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
unclecode/crawl4ai#2123 · 1 comment ·
Maintainers usually reply within 1 day
-
🐞 Bug 🩺 Needs Triage
Difficulty 4/5 3-5 days Newbie friendliness 55/100
Maintainers usually reply within 1 day
All issues in unclecode/crawl4ai
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
conda-forge/conda-build-feedstock#289 · 1 comment · 1 reaction ·
-
`pulptest` no longer works in 4.0.0: `ImportError: Start directory is not importable: 'pulp/tests'`Open
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
remove reddit feedsOpen
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
TomCasavant/ohio-sites#224 ·