Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Feature request: opt-in crawl source receipt for RAG ingestion

Open
#2,307 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
python

Research direction

Start with crawl4ai/models.py, especially the CrawlResult fields, then trace cache metadata and Docker API serialization. Define the receipt's URL, fetch-time, output-kind, digest, configuration fingerprint, cache, and outcome semantics without exposing secrets. Done means fresh, redirected, cache-hit, raw/fit, changed-configuration, serialization, secret-exclusion, and failed/partial-crawl tests satisfy the acceptance criteria while the default response remains compatible.

Written by the indexing model from the issue text.

Description

Problem / use case

Crawl4AI is positioned as a web-to-Markdown input for RAG and agent data pipelines. A typical ingestion job needs to answer three questions for every stored document: which source and crawl produced it, did the usable text actually change, and can the same processed artifact be traced later? Those answers drive re-embedding cost, stale-document replacement, and source attribution in downstream answers.

Today CrawlResult already exposes useful pieces such as url, redirected_url, response_headers, cached_at, cache_status, and Markdown (current model). The cache also has internal content hashes. But the result does not provide one documented, serialization-friendly record tying the selected output text to a stable digest, its actual fetch time, and its source. Every RAG integrator must assemble this independently, and a cache hit can otherwise be mistaken for a fresh fetch.

Proposed module: optional crawl source receipt

Add an opt-in, small source_receipt attached to each successful CrawlResult (and available through the Docker API serialization). Suggested fields:

  • Requested URL and final URL, reusing existing redirect information.
  • fetched_at: time the underlying page was fetched; on a cache hit preserve the original fetch time. Optionally include served_at separately.
  • Output kind (raw_markdown or fit_markdown) and a versioned SHA-256 digest of the exact UTF-8 text returned for that kind. Make the normalization/encoding rule explicit. An unchanged digest lets an ingestion job skip embedding; a changed digest triggers replacement.
  • Crawl4AI version and a fingerprint of content-affecting, non-secret processing options, so a changed extraction configuration does not look like unchanged content.
  • Existing cache status, and a clear distinction between successful extraction and fallback/partial output.

The default path can remain unchanged, with receipt generation behind a config flag. Please exclude cookies, Authorization headers, tokens, proxy credentials, and arbitrary response headers from the receipt/fingerprint. Consumers can retain the receipt next to their vector-store document ID; this does not require Crawl4AI to own a vector store or scheduler.

Acceptance criteria

  1. Fresh crawl, redirect, and cache-hit examples show accurate URL and fetch-time semantics.
  2. Identical selected Markdown + processing configuration yields the same digest; a text or relevant configuration change changes it.
  3. Raw and fit Markdown cannot accidentally share a digest namespace.
  4. Library and Docker API outputs serialize the receipt consistently; existing responses remain compatible when the option is off.
  5. Tests cover secret exclusion and a failed/partial crawl.

I checked main and develop result fields, the roadmap, and issue/PR searches for provenance, source lineage, content fingerprints, and crawl manifests. I found a planned incremental embedding index in the roadmap, but not this per-crawl receipt. The receipt could serve that later index while remaining useful to external ingestion pipelines.

Dominant language
Python
Stars
84.5k
Forks
8.7k
Avg merge
3d 20h
Merged PRs (30d)
15

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from unclecode/crawl4ai

All issues in unclecode/crawl4ai

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.