[Bug]: cleaned_html drops the page body when it is nested inside an unclosed <noscript>
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 58/100
Hướng nghiên cứu
Start with LXMLWebScrapingStrategy._scrap in the content scraping strategy and run the offline reproduction from the issue. Compare cleaned_html with result.html for unbalanced nested noscript tags; done means the body content remains while noscript fallback blocks are still removed. Check issue #2284 because the report says a fix is proposed there.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.9.4 (develop @ 86e6464); also reproduced on 0.9.0
Expected Behavior
cleaned_html keeps the page's body content. <noscript> fallback blocks are removed, as they are today.
Current Behavior
On pages that contain nested <noscript>, cleaned_html loses part or all of the body, even though result.html contains the full page.
Lazy-load plugins commonly wrap an existing noscript (typically Google Tag Manager's), so the served page contains:
<noscript><iframe ...></iframe><noscript><iframe src="https://www.googletagmanager.com/ns.html?id=..."></iframe></noscript></noscript>
The browser has scripting enabled, so it treats noscript contents as raw text and ends the element at the first </noscript>. The stray second closing tag is dropped when page.content() serializes the DOM, so the capture has more opening tags than closing ones. lxml parses with scripting disabled, so it nests the rest of the document inside the still-open <noscript>, and LXMLWebScrapingStrategy._scrap then deletes that content when it removes noscript elements.
One common source is LiteSpeed Cache's iframe lazy-load (~7M active WordPress installs), which rewrites GTM's noscript iframe this way: WordPress.org support thread.
Visible text in cleaned_html for affected pages, default configs (2026-09-24):
| Page | noscript open/close in capture | Visible chars in cleaned_html |
Visible chars in result.html |
|---|---|---|---|
| https://www.comunio.es/ | 5 / 4 | 0 | ~4,000 |
| https://halalvlees.nl/ | 5 / 4 | 1,834 | 12,480 |
| https://www.dietdoctor.com/ | 63 / 54 | 7,846 | 8,171 |
In a random sample of 2,972 Tranco top-1M homepages (ranks 20k–1M), 6 served nested noscript, and 4 of those lost content this way.
Is this reproducible?
Yes
Inputs Causing the Bug
- URL(s): https://www.comunio.es/ , https://halalvlees.nl/
- Settings used: defaults (BrowserConfig(), CrawlerRunConfig())
- Input data: any HTML with more <noscript> opening tags than closing tags
Steps to Reproduce
1. Crawl https://www.comunio.es/ with default settings
2. Compare result.cleaned_html with result.html
3. cleaned_html has an empty body; result.html has the full page
Code snippets
# Offline repro: the shape page.content() produces for such a page
from crawl4ai.content_scraping_strategy import LXMLWebScrapingStrategy
html = """<html><head><title>Example shop</title></head><body>
<noscript><iframe src="about:blank"></iframe><noscript><iframe src="https://www.googletagmanager.com/ns.html?id=GTM-TEST"></iframe></noscript>
<div><h1>Real heading</h1><p>Real body copy.</p></div>
</body></html>"""
print(LXMLWebScrapingStrategy().scrap("https://example.com/", html).cleaned_html)
# Actual: <html><head><title>Example shop</title></head></html>
# Expected: the <h1> and <p> are kept
# Live repro
import asyncio
from crawl4ai import AsyncWebCrawler
async def main():
async with AsyncWebCrawler() as crawler:
result = await crawler.arun("https://www.comunio.es/")
print(len(result.html), repr(result.cleaned_html[-200:]))
asyncio.run(main())
OS
macOS 26.6 (arm64)
Python version
3.12.11
Browser
Chromium (Playwright headless shell)
Browser version
153.0.8010.12
Error logs & Screenshots (if applicable)
No error is raised; result.success is True and the status code is 200.
A fix is proposed in #2284.
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
Maintainer thường phản hồi trong vòng 1 ngày
-
https://search.utilibre.orgĐang mởinstance instance add
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
searxng/searx-instances#941 · 1 bình luận ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày