Do not index irrelevant content of the pages
还没有人认领这个 Issue。
评估
调研方向
未指定文件、测试或入口点。首先定位 scraper 的页面索引路径,并检查是否使用了 libzim 的 HTML 解析器,然后比较各个项目中应排除的页面部分。完成的标准是就一种适合该项目的方式达成一致,以便从索引中省略无关内容。
由索引模型根据 Issue 内容生成。
描述
We shouldn't index some parts of the page which are not relevant:
This is a hard feature since all this comes from online, and is a bit different in every project.
Note that default HTML parser from libzim (I don't recall if we use it or not) ignores everything inside <!-- htdig_noindex -->...<!-- /htdig_noindex -->
- 主要语言
- Python
- 星标
- 9
- 派生
- 3
- PR 合并指标
- 30 天内没有已合并 PR
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
openzim/devdocs 的其他 Issue
-
bug
难度 2/5 1-3 小时 新手友好度 65/100
-
enhancement
难度 4/5 3-5 天 新手友好度 55/100
-
bug
难度 4/5 3-5 天 新手友好度 35/100
-
bug
难度 4/5 3-5 天 新手友好度 35/100
-
bug
难度 3/5 1-2 天 新手友好度 45/100
相似的 Issue
-
agent-ready documentation needs-triage
难度 1/5 1-3 小时 新手友好度 88/100
-
documentation
难度 1/5 1 小时以内 新手友好度 91/100
-
workflow-status page template still says reusable workflows are "triggered only by workflow_call:" 未关闭
难度 1/5 1 小时以内 新手友好度 92/100
-
instance instance add
难度 1/5 1 小时以内 新手友好度 72/100
searxng/searx-instances#939 · 1 条评论 ·
-
area-deployment area-integrations triage:bot-seen
难度 2/5 半天 新手友好度 86/100