Add byte-oriented sizing and validation utilities for `bloom_filter`
Los mantenedores suelen responder en 2 días
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 55/100
Línea de trabajo
Comienza con include/cuco/detail/bloom_filter/parametric_filter_policy.cuh, donde se definen words_per_block y max_filter_blocks, y luego sigue los constructores existentes de cuco::bloom_filter basados en el número de bloques y los puntos de entrada para calcular tamaños. Añade utilidades de construcción y cálculo de tamaños orientadas a bytes, con resultados alineados, positivos y limitados por la policy, así como una validación coherente; el trabajo estará terminado cuando los llamadores ya no necesiten reimplementar estas restricciones.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Is your feature request related to a problem? Please describe.
cuco::bloom_filter is currently sized only by num_blocks (a raw count of filter blocks). Callers that think in terms of a storage budget in bytes, which is the natural unit for memory allocation and for distributed work, have to convert bytes to blocks themselves and re-derive the policy's constraints: a filter size must be a positive multiple of the block size (words_per_block * sizeof(word_type)) and no greater than the policy maximum (max_filter_blocks).
This came up in rapidsai/cudf#23067, where libcudf_streaming's device_bloom_filter wrapper re-implements this sizing logic: a byte-sized constructor, an aligned_size(bytes) helper that rounds a byte count down to the largest valid filter size, and a max_size() accessor for the policy's upper bound. This is general-purpose logic that every cuco bloom filter user needs, not something specific to cudf, so it belongs in cuco with the corresponding validation checks rather than being re-derived by each caller.
Describe the solution you'd like
Expose byte-oriented sizing utilities directly on cuco::bloom_filter / its policy, mirroring the convenience constructors HyperLogLog already provides (cuco::sketch_size_kb, cuco::standard_deviation). Concretely:
- A way to construct or size a filter from a target storage size in bytes, alongside the existing block-count path.
- An
aligned_size-style helper that returns the largest valid filter size not exceeding a requested byte count (a positive multiple of the block size, capped at the policy maximum). - An accessor for the maximum storage size supported by the filter policy.
- The associated validation (positive, multiple of block size, within the policy limit) built into these utilities so callers get consistent error checking.
Describe alternatives you've considered
Keeping the conversion in each downstream project, as libcudf_streaming does today. This duplicates the block-size and maximum-size rules across users and risks them drifting from the policy's actual constraints.
Additional context
Origin discussion: https://github.com/rapidsai/cudf/pull/23067#discussion_r3562252707. The relevant policy constants (words_per_block, max_filter_blocks) live here: https://github.com/NVIDIA/cuCollections/blob/0883368d39296f3bef3a058033141bcc642c5c54/include/cuco/detail/bloom_filter/parametric_filter_policy.cuh#L101-L119
- Lenguaje dominante
- Cuda
- Estrellas
- 671
- Forks
- 122
- Merge medio
- 4 d 21 h
- PR fusionados (30 d)
- 10
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de NVIDIA/cuCollections
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
NVIDIA/cuCollections#858 ·
Los mantenedores suelen responder en 2 días
-
Add cuco::detail::stream_sync(cuda::stream_ref) to centralize CCCL version-specific API namingQuizá libre de nuevo @0z5a la tomó hace 21 días y no hay ningún pull request abierto. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
NVIDIA/cuCollections#840 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
nvidia-runners
Dificultad 1/5 1-3 horas Aptitud para principiantes 25/100
NVIDIA/cuCollections#853 ·
Los mantenedores suelen responder en 2 días
-
topic: performance type: feature request
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
NVIDIA/cuCollections#817 · 7 comentarios · 1 reacción ·
Los mantenedores suelen responder en 2 días
-
good first issue P2: Nice to have type: improvement
Dificultad 4/5 3-5 días Aptitud para principiantes 38/100
NVIDIA/cuCollections#805 · 4 comentarios ·
Los mantenedores suelen responder en 2 días
Todos los issues de NVIDIA/cuCollections
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
NaturalIntelligence/fast-xml-parser#888 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
infertopics leaves new nodes without a topic when untopiced neighbours outnumber topiced onesPosiblemente ocupada @moneebullah25 la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
NVIDIA/earth2studio#1241 ·
Los mantenedores suelen responder en 3 días