[Bug]: Markdown export loses heading hierarchy and table structure
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
调研方向
Start by reproducing the issue with the documented crwl ... -o markdown and -o html commands on a page containing nested headings and an HTML table, then compare the outputs. Identify the markdown conversion entry point and existing tests, and define done as preserving heading levels and producing valid table structure, or documenting an explicit alternative and configuration behavior.
由索引模型根据 Issue 内容生成。
描述
crawl4ai version
0.8.6
Expected Behavior
When converting a page with clear document structure in HTML, the markdown output should preserve:
- Heading hierarchy —
h1throughh6levels mapped to#through######consistently - Table structure — HTML
<table>elements converted to valid GitHub-flavored markdown tables (or documented alternative)
Alternatively, provide explicit configuration options or separate commands so users can choose between:
- Fast/minimal markdown (current behavior, smaller output)
- Structure-preserving markdown (accurate headings, proper tables)
Suggested naming:
- Config approach:
markdown.mode: "compact" | "semantic"ormarkdown.preserve_structure: true - Separate commands:
md-lite/md-semantic(CLI),crawl4ai_md/crawl4ai_md_semantic(MCP)
Suggested resolution
Option A — Configuration flags:
{
"crawler_config": {
"markdown": {
"mode": "semantic",
"preserve_headings": true,
"preserve_tables": "gfm"
}
}
}
Option B — Separate commands/endpoints:
| Use case | CLI flag | MCP tool name |
|---|---|---|
| Fast, minimal | -o markdown or -o md-lite |
crawl4ai_md |
| Structure-preserving | -o md-semantic |
crawl4ai_md_semantic |
Current Behavior
For the same URL and crawl settings, markdown output loses or degrades structural information that is preserved in HTML:
| Element | HTML Output | Markdown Output |
|---|---|---|
| Headings | Clear h1 > h2 > h3 nesting |
Levels flattened or inconsistent |
| Tables | Valid <table> with rows/columns |
Flattened to lists, paragraphs, or lost entirely |
Users must switch to HTML output and manually extract structure, defeating markdown's purpose as a readable, structured format.
Is this reproducible?
Yes
Inputs Causing the Bug
- **Documentation pages** with hierarchical sections (single `h1`, multiple `h2`/`h3` levels)
- **Data pages** with comparison tables, specification tables, or course lists in `<table>` markup
- **Any public URL** (or local HTML fixture) containing structured content
Steps to Reproduce
1. Identify a page with known heading hierarchy (`h1` → `h2` → `h3`) and at least one HTML table
2. Crawl with **markdown** output:
crwl 'https://example.invalid/test-page' -o markdown
3. Crawl the **same page** with **HTML** output:
crwl 'https://example.invalid/test-page' -o html
4. Compare:
- Count heading levels in markdown vs HTML DOM
- Check if table structure is preserved as `|` delimited markdown
5. Observe: markdown flattens headings and loses table formatting
Code snippets
**CLI comparison:**
# Markdown - structure loss visible here
crwl 'https://example.invalid/test-page' -o markdown > output.md
# HTML - structure preserved
crwl 'https://example.invalid/test-page' -o html > output.html
**HTTP API:**
# Markdown request
curl -sS 'http://localhost:PORT/crawl' \
-H 'Content-Type: application/json' \
-d '{
"urls": ["https://example.invalid/test-page"],
"crawler_config": { "cache_mode": "bypass" }
}' | jq -r '.results[0].markdown.raw_markdown'
# HTML request for comparison
curl -sS 'http://localhost:PORT/crawl' \
-H 'Content-Type: application/json' \
-d '{
"urls": ["https://example.invalid/test-page"],
"crawler_config": { "cache_mode": "bypass" }
}' | jq -r '.results[0].html'
OS
Debian GNU/Linux 12 (bookworm)
Python version
3.12.13
Browser
No response
Browser version
No response
Error logs & Screenshots (if applicable)
No response
- 主要语言
- Python
- 星标
- 84.5k
- 派生
- 8.7k
- 平均合并
- 3 天 9 小时
- 30 天内合并 PR
- 17
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
unclecode/crawl4ai 的其他 Issue
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run未关闭
难度 2/5 1-3 小时 新手友好度 78/100
unclecode/crawl4ai#2309 · 2 条评论 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 84/100
unclecode/crawl4ai#2147 · 3 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
unclecode/crawl4ai#2123 · 1 条评论 ·
维护者通常 1 天内回复
-
🐞 Bug 🩺 Needs Triage
难度 4/5 3-5 天 新手友好度 55/100
维护者通常 1 天内回复
-
难度 5/5 一周以上 新手友好度 35/100
维护者通常 1 天内回复
查看 unclecode/crawl4ai 的全部 Issue
相似的 Issue
-
namespace operations
难度 1/5 1 小时以内 新手友好度 82/100
EclipseFdn/open-vsx.org#13573 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
collective/icalendar#1854 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
rancher/rancher-ai-agent#412 ·
维护者通常 6 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
TUDelftGeodesy/DePSI#134 ·
-
难度 2/5 1-3 小时 新手友好度 88/100
HenriquesLab/rxiv-maker#335 ·