Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Bug]: Browser hangs indefinitely on WSL when system proxy is required

未关闭
#1,930 3 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
45/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
playwright, python

调研方向

Start with the AsyncWebCrawler and BrowserConfig browser-launch path, then compare it with the working Playwright launch using proxy configuration. Reproduce the WSL2 case from the issue and verify that a crawl of https://example.com completes when a proxy is required, including the proxy and extra_args configurations.

由索引模型根据 Issue 内容生成。

描述

❓ Question
crawl4ai version

0.8.6

Expected Behavior

When proxy or proxy_config is set in BrowserConfig, crawl4ai should
successfully launch Chromium with the proxy and complete the crawl.

Current Behavior

On WSL2 where a proxy is required for internet access, AsyncWebCrawler hangs
indefinitely after printing "[INIT].... → Crawl4AI 0.8.6" and never completes.
Setting proxy via BrowserConfig(proxy=...), BrowserConfig(proxy_config=...),
or extra_args=["--proxy-server=..."] all result in the same hang.

Is this reproducible?

Yes

Inputs Causing the Bug
Any URL requiring internet access (e.g. https://example.com) when running on 
WSL2 where Chrome cannot connect to the internet directly.
Steps to Reproduce
1. Install crawl4ai on WSL2:
   pip install crawl4ai
   crawl4ai-setup

2. Verify that Python requests work fine with proxy (direct Chrome cannot reach internet):
   HOST_IP=$(ip route | awk '/default/ {print $3}')
   http_proxy=http://$HOST_IP:7892 python -c "import requests; print(requests.get('https://www.google.com', timeout=15).status_code)"
   # Output: 200

3. Verify that plain Playwright works fine with proxy:
   python -c "
   from playwright.sync_api import sync_playwright
   p = sync_playwright().start()
   b = p.chromium.launch()
   print('OK')
   b.close()
   p.stop()
   "
   # Output: OK

4. Run a basic crawl4ai crawl (no proxy):
   python -c "
   import asyncio
   from crawl4ai import AsyncWebCrawler
   async def test():
       async with AsyncWebCrawler() as crawler:
           result = await crawler.arun('https://example.com')
           print(result.success)
   asyncio.run(test())
   "
   # Hangs forever at: [INIT].... → Crawl4AI 0.8.6

5. Try with proxy set in BrowserConfig:
   python -c "
   import asyncio
   from crawl4ai import AsyncWebCrawler, BrowserConfig
   async def test():
       config = BrowserConfig(proxy='http://172.29.240.1:7892')
       async with AsyncWebCrawler(config=config) as crawler:
           result = await crawler.arun('https://example.com')
           print(result.success)
   asyncio.run(test())
   "
   # Also hangs forever at: [INIT].... → Crawl4AI 0.8.6

6. Confirmed root cause: crawl4ai uses subprocess.Popen to launch Chrome 
   without passing proxy env vars, so Chrome cannot reach the internet on WSL2.
   Plain Playwright with proxy={"server": ...} works correctly.
Code snippets
# Hangs forever
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig

async def main():
    config = BrowserConfig(proxy="http://172.29.240.1:7892")
    async with AsyncWebCrawler(config=config) as crawler:
        result = await crawler.arun("https://example.com")

asyncio.run(main())

# Workaround: launch browser manually via Playwright and connect via CDP
import subprocess
from playwright.async_api import async_playwright
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig

HOST_IP = subprocess.check_output(
    "ip route | awk '/default/ {print $3}'", shell=True).decode().strip()

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(
            headless=True,
            args=["--no-sandbox", "--disable-dev-shm-usage",
                  "--remote-debugging-port=9222"],
            proxy={"server": f"http://{HOST_IP}:7892"}
        )
        config = BrowserConfig(cdp_url="http://localhost:9222")
        async with AsyncWebCrawler(config=config) as crawler:
            result = await crawler.arun("https://example.com",
                config=CrawlerRunConfig(page_timeout=15000))
            print(result.success)
        await browser.close()

asyncio.run(main())
OS

OS: Windows 11 + WSL2 (Ubuntu 24.04)

Python version

3.12

Browser

Chromium (crawl4ai managed)

Browser version

Google Chrome for Testing 145.0.7632.6

Error logs & Screenshots (if applicable)

No response

主要语言
Python
星标
84.5k
派生
8.7k
平均合并
3 天 20 小时
30 天内合并 PR
15

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

unclecode/crawl4ai 的其他 Issue

查看 unclecode/crawl4ai 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。