[Bug]: LLMExtractionStrategy not applied when CrawlerRunConfig.cache_mode=ENABLED
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 52/100
Hướng nghiên cứu
Start with the supplied AsyncWebCrawler.arun example and compare CacheMode.ENABLED with the working BYPASS case. Trace how cached results are handled alongside LLMExtractionStrategy, then verify that a cached URL still produces populated result.extracted_content and that the documented example works.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.7.4
Expected Behavior
When using cache_mode=CacheMode.ENABLED on CrawlerRunConfig and a previously crawled and cached URL the result.extracted_content is filled with freshly generated output of LLM.
Current Behavior
When using cache_mode=CacheMode.ENABLED on CrawlerRunConfig and a previously crawled and cached URL the result.extracted_content is empty.
Is this reproducible?
Yes
Code snippets
import os
import asyncio
import json
from pydantic import BaseModel, Field
from typing import List
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode, LLMConfig
from crawl4ai import LLMExtractionStrategy
class Product(BaseModel):
name: str
price: str
async def main():
# 1. Define the LLM extraction strategy
llm_strategy = LLMExtractionStrategy(
llm_config = LLMConfig(provider="openai/gpt-4o-mini", api_token=os.getenv('OPENAI_API_KEY')),
schema=Product.schema_json(), # Or use model_json_schema()
extraction_type="schema",
instruction="Extract all product objects with 'name' and 'price' from the content.",
chunk_token_threshold=1000,
overlap_rate=0.0,
apply_chunking=True,
input_format="markdown", # or "html", "fit_markdown"
extra_args={"temperature": 0.0, "max_tokens": 800}
)
# 2. Build the crawler config
crawl_config = CrawlerRunConfig(
extraction_strategy=llm_strategy,
cache_mode=CacheMode.ENABLED # (!!) BYPASS is working
)
# 3. Create a browser config if needed
browser_cfg = BrowserConfig(headless=True)
async with AsyncWebCrawler(config=browser_cfg) as crawler:
# 4. Let's say we want to crawl a single page
result = await crawler.arun(
url="https://example.com/products",
config=crawl_config
)
if result.success:
# 5. The extracted content is presumably JSON
data = json.loads(result.extracted_content)
print("Extracted items:", data)
# 6. Show usage stats
llm_strategy.show_usage() # prints token usage
else:
print("Error:", result.error_message)
The Code Example is taken from the LLM Strategies Documentation
OS
macOS
Python version
3.11.4
Browser
No response
Browser version
No response
Error logs & Screenshots (if applicable)
No response
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 4 giờ
- Pull request đã merge (30 ngày)
- 14
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
unclecode/crawl4ai#2326 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2319 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
unclecode/crawl4ai#2309 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
BasedHardware/omi#20271 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 92/100
openai/openai-cookbook#3153 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
cvss-severity:high devguard l3montree-cybersecurity/devguard/devguard pkg:golang/github.com/l3montree-dev/devguard risk:low state:open
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
l3montree-dev/devguard#3146 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
-
bug confirmed issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
open-webui/open-webui#31849 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày