[FEA] Add option to report hash collisions
维护者通常 2 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 35/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 停滞
- 技术栈
- cpp
- 领域
- performance
调研方向
从 include/cuco/detail/static_map.inl 中第 252 行附近的 no-CG static_map::find 路径开始,然后跟踪 issue 中提到的其他 insert 和 find device 函数。在 static_map 和 dynamic_map 中定义选择启用的冲突计数行为,确保计数可在 host 上访问,并且默认禁用,同时不影响标准性能。
由索引模型根据 Issue 内容生成。
描述
Is your feature request related to a problem? Please describe.
Hash collisions impact performance of hash map insert and probe. It will be useful to find a way to report the number of collisions for static_map and dynamic_map to help assess performance of insert or probe when developers are evaluating perf on their dataset. It would also allow developers to tune the hash function or occupancy to reduce collisions and find the right balance for their scenario.
Describe the solution you'd like
The map can have an optional template argument that specifies if we need to count collisions (disabled by default), so it's opt-in and doesn't impact perf for the standard case. The number of collisions would be stored in a class variable that's accessible with something like get_num_collisions(). Implementation: allocate memory for device variable uint64_t *d_num_collisions, update all insert and find device code to do atomicAdd(d_num_collisions, 1) to that variable, then copy the contents to the host variable after the kernel. Here is where we can count the collisions for no-CG static_map::find:
https://github.com/NVIDIA/cuCollections/blob/2196040f0562a0280292eebef5295d914f615e63/include/cuco/detail/static_map.inl#L252
The atomic will be guarded by the template argument check, so should only impact perf if we're asked to count collisions. Similarly, would have to update all other insert and find functions.
Describe alternatives you've considered
None.
Additional context
None.
- 主要语言
- Cuda
- 星标
- 671
- 派生
- 122
- 平均合并
- 4 天 19 小时
- 30 天内合并 PR
- 10
环境准备
在浏览器里用你自己的 GitHub 账号启动这个项目的开发容器。
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/cuCollections 的其他 Issue
-
Add cuco::detail::stream_sync(cuda::stream_ref) to centralize CCCL version-specific API naming可能重新可做 @0z5a 于 22 天前认领,目前没有进行中的 PR。 未关闭
难度 2/5 1-3 小时 新手友好度 74/100
NVIDIA/cuCollections#840 · 1 条评论 ·
维护者通常 2 天内回复
-
nvidia-runners
难度 1/5 1-3 小时 新手友好度 25/100
NVIDIA/cuCollections#853 ·
维护者通常 2 天内回复
-
Add byte-oriented sizing and validation utilities for `bloom_filter`可能已有人在做 @yuweih205 于 35 天前认领。 未关闭helps: rapids topic: bloom_filter type: feature request
难度 4/5 3-5 天 新手友好度 55/100
NVIDIA/cuCollections#829 · 2 条评论 ·
维护者通常 2 天内回复
-
topic: performance type: feature request
难度 5/5 一周以上 新手友好度 35/100
NVIDIA/cuCollections#817 · 7 条评论 · 1 个 reaction ·
维护者通常 2 天内回复
-
good first issue P2: Nice to have type: improvement
难度 4/5 3-5 天 新手友好度 38/100
NVIDIA/cuCollections#805 · 4 条评论 ·
维护者通常 2 天内回复
查看 NVIDIA/cuCollections 的全部 Issue
相似的 Issue
-
CLI contributor: external
难度 2/5 1-3 小时 新手友好度 66/100
维护者通常 1 天内回复
-
external feature request text-splitters
难度 2/5 1-3 小时 新手友好度 60/100
langchain-ai/langchain#41191 ·
维护者通常 1 天内回复
-
Bug: Sliders Re-render on Every Resize Even When Thumb-Alignment="Center"可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭bug confirmed perf
难度 2/5 1-3 小时 新手友好度 72/100
videojs/video.js#9400 · 1 条评论 ·
维护者通常 1 天内回复
-
core priority: medium
难度 2/5 1-3 小时 新手友好度 76/100
Kuldeep2822k/cli#332 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
nightscout/AndroidAPS#5245 ·
维护者通常 1 天内回复