[Bug]: Browser hangs indefinitely on WSL when system proxy is required
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 45/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- playwright, python
- 领域
- backend, networking
调研方向
Start with the AsyncWebCrawler and BrowserConfig browser-launch path, then compare it with the working Playwright launch using proxy configuration. Reproduce the WSL2 case from the issue and verify that a crawl of https://example.com completes when a proxy is required, including the proxy and extra_args configurations.
由索引模型根据 Issue 内容生成。
描述
crawl4ai version
0.8.6
Expected Behavior
When proxy or proxy_config is set in BrowserConfig, crawl4ai should
successfully launch Chromium with the proxy and complete the crawl.
Current Behavior
On WSL2 where a proxy is required for internet access, AsyncWebCrawler hangs
indefinitely after printing "[INIT].... → Crawl4AI 0.8.6" and never completes.
Setting proxy via BrowserConfig(proxy=...), BrowserConfig(proxy_config=...),
or extra_args=["--proxy-server=..."] all result in the same hang.
Is this reproducible?
Yes
Inputs Causing the Bug
Any URL requiring internet access (e.g. https://example.com) when running on
WSL2 where Chrome cannot connect to the internet directly.
Steps to Reproduce
1. Install crawl4ai on WSL2:
pip install crawl4ai
crawl4ai-setup
2. Verify that Python requests work fine with proxy (direct Chrome cannot reach internet):
HOST_IP=$(ip route | awk '/default/ {print $3}')
http_proxy=http://$HOST_IP:7892 python -c "import requests; print(requests.get('https://www.google.com', timeout=15).status_code)"
# Output: 200
3. Verify that plain Playwright works fine with proxy:
python -c "
from playwright.sync_api import sync_playwright
p = sync_playwright().start()
b = p.chromium.launch()
print('OK')
b.close()
p.stop()
"
# Output: OK
4. Run a basic crawl4ai crawl (no proxy):
python -c "
import asyncio
from crawl4ai import AsyncWebCrawler
async def test():
async with AsyncWebCrawler() as crawler:
result = await crawler.arun('https://example.com')
print(result.success)
asyncio.run(test())
"
# Hangs forever at: [INIT].... → Crawl4AI 0.8.6
5. Try with proxy set in BrowserConfig:
python -c "
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig
async def test():
config = BrowserConfig(proxy='http://172.29.240.1:7892')
async with AsyncWebCrawler(config=config) as crawler:
result = await crawler.arun('https://example.com')
print(result.success)
asyncio.run(test())
"
# Also hangs forever at: [INIT].... → Crawl4AI 0.8.6
6. Confirmed root cause: crawl4ai uses subprocess.Popen to launch Chrome
without passing proxy env vars, so Chrome cannot reach the internet on WSL2.
Plain Playwright with proxy={"server": ...} works correctly.
Code snippets
# Hangs forever
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig
async def main():
config = BrowserConfig(proxy="http://172.29.240.1:7892")
async with AsyncWebCrawler(config=config) as crawler:
result = await crawler.arun("https://example.com")
asyncio.run(main())
# Workaround: launch browser manually via Playwright and connect via CDP
import subprocess
from playwright.async_api import async_playwright
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig
HOST_IP = subprocess.check_output(
"ip route | awk '/default/ {print $3}'", shell=True).decode().strip()
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(
headless=True,
args=["--no-sandbox", "--disable-dev-shm-usage",
"--remote-debugging-port=9222"],
proxy={"server": f"http://{HOST_IP}:7892"}
)
config = BrowserConfig(cdp_url="http://localhost:9222")
async with AsyncWebCrawler(config=config) as crawler:
result = await crawler.arun("https://example.com",
config=CrawlerRunConfig(page_timeout=15000))
print(result.success)
await browser.close()
asyncio.run(main())
OS
OS: Windows 11 + WSL2 (Ubuntu 24.04)
Python version
3.12
Browser
Chromium (crawl4ai managed)
Browser version
Google Chrome for Testing 145.0.7632.6
Error logs & Screenshots (if applicable)
No response
- 主要语言
- Python
- 星标
- 84.5k
- 派生
- 8.7k
- 平均合并
- 3 天 20 小时
- 30 天内合并 PR
- 15
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
unclecode/crawl4ai 的其他 Issue
-
🐞 Bug 🩺 Needs Triage
难度 2/5 1-3 小时 新手友好度 84/100
unclecode/crawl4ai#2319 · 2 条评论 ·
维护者通常 1 天内回复
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run未关闭
难度 2/5 1-3 小时 新手友好度 78/100
unclecode/crawl4ai#2309 · 2 条评论 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 84/100
unclecode/crawl4ai#2147 · 3 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
unclecode/crawl4ai#2123 · 1 条评论 ·
维护者通常 1 天内回复
-
🐞 Bug 🩺 Needs Triage
难度 4/5 3-5 天 新手友好度 55/100
维护者通常 1 天内回复
查看 unclecode/crawl4ai 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
approved correction metadata
难度 1/5 1 小时以内 新手友好度 88/100
acl-org/acl-anthology#10133 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
BasedHardware/omi#20084 ·
维护者通常 1 天内回复
-
bug needs-acceptance wg/evaluation-quality
难度 2/5 1-3 小时 新手友好度 76/100
vllm-project/semantic-router#4424 ·
维护者通常 1 天内回复