Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Bug]: preserve_tags/preserve_classes are a no-op for excluded tags (aside, nav, footer, header, form)

Đang mở
#2,125 2 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
78/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
python
Lĩnh vực
backend

Hướng nghiên cứu

Start in content_filter_strategy.py, especially filter_content(), _remove_unwanted_tags(), and _is_preserved(). Run the provided HTML snippet without a browser or network, then trace how excluded tags are removed before preservation is checked. Done means preserve_tags and preserve_classes retain matching aside content while the default still excludes it.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

🐞 Bug 🩺 Needs Triage
crawl4ai version

0.9.2

Expected Behavior

PruningContentFilter(preserve_tags=["aside"]) should keep <aside> content in fit_markdown.
#1900 added these options as an escape hatch for content the fit pipeline treats as boilerplate but the user knows is real content.

Same for preserve_classes=["related"] on an <aside class="related">.

Current Behavior

Both options are silently ignored for every tag in excluded_tags (content_filter_strategy.py:101-110): nav, footer, header, aside, script, style, form, iframe, noscript.

filter_content() removes those tags before pruning ever runs:

self._remove_unwanted_tags(soup)   # content_filter_strategy.py:664, decomposes aside
body = soup.find("body")
self._prune_tree(body)             # :668, the only place _is_preserved() is consulted

_remove_unwanted_tags decomposes by tag name and never consults the whitelist:

def _remove_unwanted_tags(self, soup):   # :685
    """Removes unwanted tags"""
    for tag in self.excluded_tags:
        for element in soup.find_all(tag):
            element.decompose()

By the time _is_preserved() (:691) is reached, the node is gone from the tree.
The filter raises nothing and logs nothing.

excluded_tags is hardcoded in RelevantContentFilter.__init__, not a constructor argument, so the only workaround is mutating the instance attribute (f.excluded_tags.discard("aside")).

Is this reproducible?

Yes

Inputs Causing the Bug
- URL(s): none needed, filter_content() takes an HTML string directly
- Settings used: PruningContentFilter(preserve_tags=["aside"])
- Input data: any HTML with an <aside> holding real content
Steps to Reproduce
1. Run the snippet below (no browser, no network).
2. All three cases print "aside kept: False".
3. Expected: True for the preserve_tags and preserve_classes cases.
Code snippets
from crawl4ai.content_filter_strategy import PruningContentFilter

html = (
    "<html><body>"
    "<article><p>" + "Main article body with plenty of words to pass the pruning threshold easily. " * 8 + "</p></article>"
    '<aside class="related"><h2>Related reading</h2><p>'
    + "This sidebar note is real content the author wrote and wants kept. " * 8
    + "</p></aside>"
    "</body></html>"
)

for kwargs in ({}, {"preserve_tags": ["aside"]}, {"preserve_classes": ["related"]}):
    out = " ".join(PruningContentFilter(**kwargs).filter_content(html))
    print(kwargs, "-> aside kept:", "Related reading" in out)

# 0.9.2 / develop @ 2d8f673:
# {}                                -> aside kept: False   # correct, that's the default
# {'preserve_tags': ['aside']}      -> aside kept: False   # BUG
# {'preserve_classes': ['related']} -> aside kept: False   # BUG
OS

macOS

Python version

3.12

Browser

N/A

Browser version

No response

Error logs & Screenshots (if applicable)

No response

Ngôn ngữ chính
Python
Star
84.5k
Fork
8.7k
Merge trung bình
3 ngày 9 giờ
Pull request đã merge (30 ngày)
17

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của unclecode/crawl4ai

Tất cả issue của unclecode/crawl4ai

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.