[Bug]: preserve_tags/preserve_classes are a no-op for excluded tags (aside, nav, footer, header, form)
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 78/100
Hướng nghiên cứu
Start in content_filter_strategy.py, especially filter_content(), _remove_unwanted_tags(), and _is_preserved(). Run the provided HTML snippet without a browser or network, then trace how excluded tags are removed before preservation is checked. Done means preserve_tags and preserve_classes retain matching aside content while the default still excludes it.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.9.2
Expected Behavior
PruningContentFilter(preserve_tags=["aside"]) should keep <aside> content in fit_markdown.
#1900 added these options as an escape hatch for content the fit pipeline treats as boilerplate but the user knows is real content.
Same for preserve_classes=["related"] on an <aside class="related">.
Current Behavior
Both options are silently ignored for every tag in excluded_tags (content_filter_strategy.py:101-110): nav, footer, header, aside, script, style, form, iframe, noscript.
filter_content() removes those tags before pruning ever runs:
self._remove_unwanted_tags(soup) # content_filter_strategy.py:664, decomposes aside
body = soup.find("body")
self._prune_tree(body) # :668, the only place _is_preserved() is consulted
_remove_unwanted_tags decomposes by tag name and never consults the whitelist:
def _remove_unwanted_tags(self, soup): # :685
"""Removes unwanted tags"""
for tag in self.excluded_tags:
for element in soup.find_all(tag):
element.decompose()
By the time _is_preserved() (:691) is reached, the node is gone from the tree.
The filter raises nothing and logs nothing.
excluded_tags is hardcoded in RelevantContentFilter.__init__, not a constructor argument, so the only workaround is mutating the instance attribute (f.excluded_tags.discard("aside")).
Is this reproducible?
Yes
Inputs Causing the Bug
- URL(s): none needed, filter_content() takes an HTML string directly
- Settings used: PruningContentFilter(preserve_tags=["aside"])
- Input data: any HTML with an <aside> holding real content
Steps to Reproduce
1. Run the snippet below (no browser, no network).
2. All three cases print "aside kept: False".
3. Expected: True for the preserve_tags and preserve_classes cases.
Code snippets
from crawl4ai.content_filter_strategy import PruningContentFilter
html = (
"<html><body>"
"<article><p>" + "Main article body with plenty of words to pass the pruning threshold easily. " * 8 + "</p></article>"
'<aside class="related"><h2>Related reading</h2><p>'
+ "This sidebar note is real content the author wrote and wants kept. " * 8
+ "</p></aside>"
"</body></html>"
)
for kwargs in ({}, {"preserve_tags": ["aside"]}, {"preserve_classes": ["related"]}):
out = " ".join(PruningContentFilter(**kwargs).filter_content(html))
print(kwargs, "-> aside kept:", "Related reading" in out)
# 0.9.2 / develop @ 2d8f673:
# {} -> aside kept: False # correct, that's the default
# {'preserve_tags': ['aside']} -> aside kept: False # BUG
# {'preserve_classes': ['related']} -> aside kept: False # BUG
OS
macOS
Python version
3.12
Browser
N/A
Browser version
No response
Error logs & Screenshots (if applicable)
No response
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
Maintainer thường phản hồi trong vòng 1 ngày
-
https://search.utilibre.orgĐang mởinstance instance add
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
searxng/searx-instances#941 · 1 bình luận ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày