Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug]: preserve_tags/preserve_classes are a no-op for excluded tags (aside, nav, footer, header, form)

オープン
#2,125 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
78/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
静か
技術スタック
python
領域
backend

調査の方向性

Start in content_filter_strategy.py, especially filter_content(), _remove_unwanted_tags(), and _is_preserved(). Run the provided HTML snippet without a browser or network, then trace how excluded tags are removed before preservation is checked. Done means preserve_tags and preserve_classes retain matching aside content while the default still excludes it.

索引モデルが issue の本文から書いたものです。

説明

🐞 Bug 🩺 Needs Triage
crawl4ai version

0.9.2

Expected Behavior

PruningContentFilter(preserve_tags=["aside"]) should keep <aside> content in fit_markdown.
#1900 added these options as an escape hatch for content the fit pipeline treats as boilerplate but the user knows is real content.

Same for preserve_classes=["related"] on an <aside class="related">.

Current Behavior

Both options are silently ignored for every tag in excluded_tags (content_filter_strategy.py:101-110): nav, footer, header, aside, script, style, form, iframe, noscript.

filter_content() removes those tags before pruning ever runs:

self._remove_unwanted_tags(soup)   # content_filter_strategy.py:664, decomposes aside
body = soup.find("body")
self._prune_tree(body)             # :668, the only place _is_preserved() is consulted

_remove_unwanted_tags decomposes by tag name and never consults the whitelist:

def _remove_unwanted_tags(self, soup):   # :685
    """Removes unwanted tags"""
    for tag in self.excluded_tags:
        for element in soup.find_all(tag):
            element.decompose()

By the time _is_preserved() (:691) is reached, the node is gone from the tree.
The filter raises nothing and logs nothing.

excluded_tags is hardcoded in RelevantContentFilter.__init__, not a constructor argument, so the only workaround is mutating the instance attribute (f.excluded_tags.discard("aside")).

Is this reproducible?

Yes

Inputs Causing the Bug
- URL(s): none needed, filter_content() takes an HTML string directly
- Settings used: PruningContentFilter(preserve_tags=["aside"])
- Input data: any HTML with an <aside> holding real content
Steps to Reproduce
1. Run the snippet below (no browser, no network).
2. All three cases print "aside kept: False".
3. Expected: True for the preserve_tags and preserve_classes cases.
Code snippets
from crawl4ai.content_filter_strategy import PruningContentFilter

html = (
    "<html><body>"
    "<article><p>" + "Main article body with plenty of words to pass the pruning threshold easily. " * 8 + "</p></article>"
    '<aside class="related"><h2>Related reading</h2><p>'
    + "This sidebar note is real content the author wrote and wants kept. " * 8
    + "</p></aside>"
    "</body></html>"
)

for kwargs in ({}, {"preserve_tags": ["aside"]}, {"preserve_classes": ["related"]}):
    out = " ".join(PruningContentFilter(**kwargs).filter_content(html))
    print(kwargs, "-> aside kept:", "Related reading" in out)

# 0.9.2 / develop @ 2d8f673:
# {}                                -> aside kept: False   # correct, that's the default
# {'preserve_tags': ['aside']}      -> aside kept: False   # BUG
# {'preserve_classes': ['related']} -> aside kept: False   # BUG
OS

macOS

Python version

3.12

Browser

N/A

Browser version

No response

Error logs & Screenshots (if applicable)

No response

主要言語
Python
スター
84.5k
フォーク
8.7k
平均マージ
3日 9時間
マージ済み PR(30日)
17

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

unclecode/crawl4ai のほかの issue

unclecode/crawl4ai の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。