MLSAKIIT/pixly

Implement a more robust web scrapper

オープン

#12 opened on 2025/10/14

 (2 件のコメント) (0 件のリアクション) (1 人の担当者)Python (15 件のフォーク)auto 404
enhancementhacktoberfest

Repository metrics

Stars
 (20 個のスター)
PR merge metrics
 (PR metrics pending)

説明

Web Scrappers in knowledge_manager.py, extract_wiki_content() and extract_forum_content() sometimes gets blocked by certain websites, (namely the ones present in minecraft.csv)

Recommended Solutions :

  • Use an existing tool such as scrappy or playwright.
  • Mask the scrappers to behave to more like human by introducing delays between different searches.
  • Make the scrapper behave like a browser as much as possible
  • Use Proxies to circumvent IP bans.
  • Rotate a list of User Agents and headers.
  • Make it asynchronous using asyncio + httpx

コントリビューターガイド