Is beartype cost acceptable?
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Facilidade para iniciantes
- 35/100
- Tipo de issue
- Refatoração
- Clareza
- Precisa de esclarecimento
- Status de atividade
- Pouca atividade
- Stack de tecnologia
- python
- Domínio
- performance
Direção de pesquisa
Comece com zimscraperlib/init.py e a chamada beartype_this_package(), depois revise o hot path zimscraperlib.rewriting e o comando de benchmark documentado do warc2zim. Reproduza as execuções pareadas, se necessário, e compare o tempo de parede, o tempo de CPU e o RSS máximo. O trabalho estará concluído quando houver acordo e documentação sobre se a cobertura deve ter o escopo limitado, ser configurada com uma estratégia mais barata ou ser desabilitada para execuções selecionadas.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Context
While investigating a memory leak (fixed in #324 — a closure in RxRewriter.rewrite() was being re-decorated by beartype's claw hook on every call, permanently pinning it via beartype's own is_object_blacklisted cache), the fix removed the leak but also raised the question of how much runtime cost beartype_this_package() (used project-wide in zimscraperlib/__init__.py) actually adds. This issue is to share those numbers and discuss whether the tradeoff is the one we want.
Methodology
Ran warc2zim to completion, twice, on the same input, same machine, back to back:
warc2zim --publisher=openZIM --content-header-bytes-length=2048 \
--zim-file=www.physicsclassroom.com_56e21a6e.zim --name=www.physicsclassroom.com_56e21a6e \
--scraper-suffix "zimit 3.1.2" --output output \
--url https://www.physicsclassroom.com/ \
<local WARC, 731MB compressed, ~14.5K ZIM entries> --overwrite
wrapped in /usr/bin/time -v for wall-clock, CPU, and peak RSS.
- Run A:
zimscraperlibas-is (beartype_this_package()active). - Run B: same install, with the single line
beartype_this_package()inzimscraperlib/__init__.pycommented out (everything else identical), bytecode cache cleared between runs.
Both runs post-date the #324 leak fix, so memory is no longer confounded by that bug.
Results
| With beartype (Run A) | Without beartype (Run B) | Delta | |
|---|---|---|---|
| Wall clock time | 53:50 (3230s) | 44:08 (2648s) | −18.0% |
| User CPU time | 3835s | 3251s | −15.2% |
| Max RSS | 924,060 KB | 885,356 KB | −4.2% (noise-level) |
Both runs completed successfully with an identical entry count. Memory is essentially the same between the two (the small gap is within run-to-run noise).
Takeaway
Runtime type-checking via beartype_this_package() costs roughly 15-18% additional wall-clock/CPU time on this workload, with no meaningful memory cost of its own now that #324 is fixed.
Discussion
Is this tradeoff (broad runtime type-safety across zimscraperlib, for ~15-18% slower conversions) the one we want, or should beartype's coverage be scoped down — e.g. excluding hot paths like zimscraperlib.rewriting from beartype_this_package(), or using a cheaper BeartypeStrategy for those modules — while keeping full checking elsewhere, or allowing to fully disable beartype "on-demand" (e.g. enable it only for unit and e2e tests of scraperlib and scrapers)?
- Linguagem predominante
- Python
- Estrelas
- 31
- Forks
- 31
- Merge médio
- 2d 5h
- PRs com merge (30d)
- 3
Preparar o ambiente
Este projeto não oferece contêiner de desenvolvimento, Dockerfile nem guia de contribuição, então a configuração fica por sua conta: comece pelo README e veja nosso guia da primeira contribuição para os passos gerais.
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de openzim/python-scraperlib
-
HTML rewriting: also rewrite `poster` attributeTalvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
openzim/python-scraperlib#339 ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
openzim/python-scraperlib#292 ·
-
Dificuldade 3/5 1-2 dias Facilidade para iniciantes 68/100
openzim/python-scraperlib#346 ·
-
URL normalisation: do not rewrite consecutive slashes `//` as a single slash `/`Talvez já em andamento @anshuman83-40 assumiu há 7 dias. Aberta
Dificuldade 3/5 1-2 dias Facilidade para iniciantes 45/100
openzim/python-scraperlib#340 ·
-
Add fuzzy rule to rewrite URLs of lesbases.anct.gouv.frTalvez livre de novo @benoit74 assumiu há 50 dias e não há nenhum pull request aberto. Aberta
openzim/python-scraperlib#334 · 1 responsável ·
Todas as issues de openzim/python-scraperlib
Issues semelhantes
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
NousResearch/hermes-agent#136483 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
Mantenedores costumam responder em até 1 dia
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last elementTalvez já em andamento @peterdsharpe assumiu hoje. Abertabug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
pytorch/tensordict#2307 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
Mantenedores costumam responder em até 1 dia
-
GrokModel.generate/a_generate pass an OpenAI-style list-of-dicts to xai_sdk.chat.user(), so every call crashes with a protobuf TypeError before any network I/OTalvez já em andamento @Christian-Sidak assumiu hoje. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 70/100
confident-ai/deepeval#3436 · 1 comentário ·
Mantenedores costumam responder em até 1 dia