[Bug]: Markdown export loses heading hierarchy and table structure
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 42/100
Hướng nghiên cứu
Start by reproducing the issue with the documented crwl ... -o markdown and -o html commands on a page containing nested headings and an HTML table, then compare the outputs. Identify the markdown conversion entry point and existing tests, and define done as preserving heading levels and producing valid table structure, or documenting an explicit alternative and configuration behavior.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.8.6
Expected Behavior
When converting a page with clear document structure in HTML, the markdown output should preserve:
- Heading hierarchy —
h1throughh6levels mapped to#through######consistently - Table structure — HTML
<table>elements converted to valid GitHub-flavored markdown tables (or documented alternative)
Alternatively, provide explicit configuration options or separate commands so users can choose between:
- Fast/minimal markdown (current behavior, smaller output)
- Structure-preserving markdown (accurate headings, proper tables)
Suggested naming:
- Config approach:
markdown.mode: "compact" | "semantic"ormarkdown.preserve_structure: true - Separate commands:
md-lite/md-semantic(CLI),crawl4ai_md/crawl4ai_md_semantic(MCP)
Suggested resolution
Option A — Configuration flags:
{
"crawler_config": {
"markdown": {
"mode": "semantic",
"preserve_headings": true,
"preserve_tables": "gfm"
}
}
}
Option B — Separate commands/endpoints:
| Use case | CLI flag | MCP tool name |
|---|---|---|
| Fast, minimal | -o markdown or -o md-lite |
crawl4ai_md |
| Structure-preserving | -o md-semantic |
crawl4ai_md_semantic |
Current Behavior
For the same URL and crawl settings, markdown output loses or degrades structural information that is preserved in HTML:
| Element | HTML Output | Markdown Output |
|---|---|---|
| Headings | Clear h1 > h2 > h3 nesting |
Levels flattened or inconsistent |
| Tables | Valid <table> with rows/columns |
Flattened to lists, paragraphs, or lost entirely |
Users must switch to HTML output and manually extract structure, defeating markdown's purpose as a readable, structured format.
Is this reproducible?
Yes
Inputs Causing the Bug
- **Documentation pages** with hierarchical sections (single `h1`, multiple `h2`/`h3` levels)
- **Data pages** with comparison tables, specification tables, or course lists in `<table>` markup
- **Any public URL** (or local HTML fixture) containing structured content
Steps to Reproduce
1. Identify a page with known heading hierarchy (`h1` → `h2` → `h3`) and at least one HTML table
2. Crawl with **markdown** output:
crwl 'https://example.invalid/test-page' -o markdown
3. Crawl the **same page** with **HTML** output:
crwl 'https://example.invalid/test-page' -o html
4. Compare:
- Count heading levels in markdown vs HTML DOM
- Check if table structure is preserved as `|` delimited markdown
5. Observe: markdown flattens headings and loses table formatting
Code snippets
**CLI comparison:**
# Markdown - structure loss visible here
crwl 'https://example.invalid/test-page' -o markdown > output.md
# HTML - structure preserved
crwl 'https://example.invalid/test-page' -o html > output.html
**HTTP API:**
# Markdown request
curl -sS 'http://localhost:PORT/crawl' \
-H 'Content-Type: application/json' \
-d '{
"urls": ["https://example.invalid/test-page"],
"crawler_config": { "cache_mode": "bypass" }
}' | jq -r '.results[0].markdown.raw_markdown'
# HTML request for comparison
curl -sS 'http://localhost:PORT/crawl' \
-H 'Content-Type: application/json' \
-d '{
"urls": ["https://example.invalid/test-page"],
"crawler_config": { "cache_mode": "bypass" }
}' | jq -r '.results[0].html'
OS
Debian GNU/Linux 12 (bookworm)
Python version
3.12.13
Browser
No response
Browser version
No response
Error logs & Screenshots (if applicable)
No response
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
unclecode/crawl4ai#2309 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
customer-reported
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Azure/azure-cli#34150 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
community-request
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 95/100
NVIDIA-NeMo/Curator#2464 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
weblate-discover crashes with an unhandled FileNotFoundError when the directory does not existĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
WeblateOrg/translation-finder#1099 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
trezor/trezor-firmware#7997 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày