Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Add byte-oriented sizing and validation utilities for `bloom_filter`

Abierto
#829 2 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 2 días

@yuweih205 ya está trabajando en esto.

Desde el 5/9/2026.

  • #841 de @yuweih205 — abierto

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
55/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
cpp
Área
data

Línea de trabajo

Comienza con include/cuco/detail/bloom_filter/parametric_filter_policy.cuh, donde se definen words_per_block y max_filter_blocks, y luego sigue los constructores existentes de cuco::bloom_filter basados en el número de bloques y los puntos de entrada para calcular tamaños. Añade utilidades de construcción y cálculo de tamaños orientadas a bytes, con resultados alineados, positivos y limitados por la policy, así como una validación coherente; el trabajo estará terminado cuando los llamadores ya no necesiten reimplementar estas restricciones.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

helps: rapids topic: bloom_filter type: feature request

Is your feature request related to a problem? Please describe.

cuco::bloom_filter is currently sized only by num_blocks (a raw count of filter blocks). Callers that think in terms of a storage budget in bytes, which is the natural unit for memory allocation and for distributed work, have to convert bytes to blocks themselves and re-derive the policy's constraints: a filter size must be a positive multiple of the block size (words_per_block * sizeof(word_type)) and no greater than the policy maximum (max_filter_blocks).

This came up in rapidsai/cudf#23067, where libcudf_streaming's device_bloom_filter wrapper re-implements this sizing logic: a byte-sized constructor, an aligned_size(bytes) helper that rounds a byte count down to the largest valid filter size, and a max_size() accessor for the policy's upper bound. This is general-purpose logic that every cuco bloom filter user needs, not something specific to cudf, so it belongs in cuco with the corresponding validation checks rather than being re-derived by each caller.

Describe the solution you'd like

Expose byte-oriented sizing utilities directly on cuco::bloom_filter / its policy, mirroring the convenience constructors HyperLogLog already provides (cuco::sketch_size_kb, cuco::standard_deviation). Concretely:

  • A way to construct or size a filter from a target storage size in bytes, alongside the existing block-count path.
  • An aligned_size-style helper that returns the largest valid filter size not exceeding a requested byte count (a positive multiple of the block size, capped at the policy maximum).
  • An accessor for the maximum storage size supported by the filter policy.
  • The associated validation (positive, multiple of block size, within the policy limit) built into these utilities so callers get consistent error checking.

Describe alternatives you've considered

Keeping the conversion in each downstream project, as libcudf_streaming does today. This duplicates the block-size and maximum-size rules across users and risks them drifting from the policy's actual constraints.

Additional context

Origin discussion: https://github.com/rapidsai/cudf/pull/23067#discussion_r3562252707. The relevant policy constants (words_per_block, max_filter_blocks) live here: https://github.com/NVIDIA/cuCollections/blob/0883368d39296f3bef3a058033141bcc642c5c54/include/cuco/detail/bloom_filter/parametric_filter_policy.cuh#L101-L119

Lenguaje dominante
Cuda
Estrellas
671
Forks
122
Merge medio
4 d 21 h
PR fusionados (30 d)
10

Preparar el entorno

Abrir en Codespaces

Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de NVIDIA/cuCollections

Todos los issues de NVIDIA/cuCollections

Issues similares

Más issues de Data Engineering

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.