Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug]: cleaned_html drops the page body when it is nested inside an unclosed <noscript>

オープン
#2,293 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
58/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python
領域
backend, web-dev

調査の方向性

Start with LXMLWebScrapingStrategy._scrap in the content scraping strategy and run the offline reproduction from the issue. Compare cleaned_html with result.html for unbalanced nested noscript tags; done means the body content remains while noscript fallback blocks are still removed. Check issue #2284 because the report says a fix is proposed there.

索引モデルが issue の本文から書いたものです。

説明

crawl4ai version

0.9.4 (develop @ 86e6464); also reproduced on 0.9.0

Expected Behavior

cleaned_html keeps the page's body content. <noscript> fallback blocks are removed, as they are today.

Current Behavior

On pages that contain nested <noscript>, cleaned_html loses part or all of the body, even though result.html contains the full page.

Lazy-load plugins commonly wrap an existing noscript (typically Google Tag Manager's), so the served page contains:

<noscript><iframe ...></iframe><noscript><iframe src="https://www.googletagmanager.com/ns.html?id=..."></iframe></noscript></noscript>

The browser has scripting enabled, so it treats noscript contents as raw text and ends the element at the first </noscript>. The stray second closing tag is dropped when page.content() serializes the DOM, so the capture has more opening tags than closing ones. lxml parses with scripting disabled, so it nests the rest of the document inside the still-open <noscript>, and LXMLWebScrapingStrategy._scrap then deletes that content when it removes noscript elements.

One common source is LiteSpeed Cache's iframe lazy-load (~7M active WordPress installs), which rewrites GTM's noscript iframe this way: WordPress.org support thread.

Visible text in cleaned_html for affected pages, default configs (2026-09-24):

Page noscript open/close in capture Visible chars in cleaned_html Visible chars in result.html
https://www.comunio.es/ 5 / 4 0 ~4,000
https://halalvlees.nl/ 5 / 4 1,834 12,480
https://www.dietdoctor.com/ 63 / 54 7,846 8,171

In a random sample of 2,972 Tranco top-1M homepages (ranks 20k–1M), 6 served nested noscript, and 4 of those lost content this way.

Is this reproducible?

Yes

Inputs Causing the Bug
- URL(s): https://www.comunio.es/ , https://halalvlees.nl/
- Settings used: defaults (BrowserConfig(), CrawlerRunConfig())
- Input data: any HTML with more <noscript> opening tags than closing tags
Steps to Reproduce
1. Crawl https://www.comunio.es/ with default settings
2. Compare result.cleaned_html with result.html
3. cleaned_html has an empty body; result.html has the full page
Code snippets
# Offline repro: the shape page.content() produces for such a page
from crawl4ai.content_scraping_strategy import LXMLWebScrapingStrategy

html = """<html><head><title>Example shop</title></head><body>
<noscript><iframe src="about:blank"></iframe><noscript><iframe src="https://www.googletagmanager.com/ns.html?id=GTM-TEST"></iframe></noscript>
<div><h1>Real heading</h1><p>Real body copy.</p></div>
</body></html>"""

print(LXMLWebScrapingStrategy().scrap("https://example.com/", html).cleaned_html)
# Actual:   <html><head><title>Example shop</title></head></html>
# Expected: the <h1> and <p> are kept
# Live repro
import asyncio
from crawl4ai import AsyncWebCrawler

async def main():
    async with AsyncWebCrawler() as crawler:
        result = await crawler.arun("https://www.comunio.es/")
        print(len(result.html), repr(result.cleaned_html[-200:]))

asyncio.run(main())
OS

macOS 26.6 (arm64)

Python version

3.12.11

Browser

Chromium (Playwright headless shell)

Browser version

153.0.8010.12

Error logs & Screenshots (if applicable)

No error is raised; result.success is True and the status code is 200.

A fix is proposed in #2284.

主要言語
Python
スター
84.5k
フォーク
8.7k
平均マージ
3日 9時間
マージ済み PR(30日)
17

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

unclecode/crawl4ai のほかの issue

unclecode/crawl4ai の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。