Unexpected degradation of FPR when `pattern_bits=4`
Maintainer thường phản hồi trong vòng 2 ngày
@sleeepyjack đang làm issue này rồi.
Từ ngày 27/2/2025.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
I am encountering unexpected behavior when using cuco::bloom_filter with pattern_bits = 4. The false positive rate (FPR) degrades too dramatically when changing from pattern_bits = 8 with a constant 'load factor' (i.e., the fraction of bits set in the filter). The issue may be related to the bit pattern selection.
The following code demonstrates the issue:
#include <cuco/bloom_filter.cuh>
#include <iostream>
#include <thrust/count.h>
#include <thrust/device_vector.h>
#include <thrust/sequence.h>
// 'Blocked' filter policy with 8B blocks
using policy_t = cuco::default_filter_policy<cuco::xxhash_64<uint32_t>, uint64_t, 1>;
using bf_t =
cuco::bloom_filter<uint32_t, cuco::extent<std::size_t>, cuda::thread_scope_device, policy_t>;
constexpr size_t bits_per_block = 64;
constexpr uint32_t pattern_bits_A = 4;
constexpr uint32_t pattern_bits_B = 8;
constexpr size_t bits_per_key_A = 2 * pattern_bits_A;
constexpr size_t bits_per_key_B = 2 * pattern_bits_B;
int main()
{
// Initialize non-overlapping build and probe key sets
thrust::device_vector<uint32_t> build_keys(1U << 20U);
thrust::device_vector<uint32_t> probe_keys(1U << 25U);
thrust::device_vector<bool> flags_A(1U << 25U, false);
thrust::device_vector<bool> flags_B(1U << 25U, false);
thrust::sequence(build_keys.begin(), build_keys.end(), 0, 2);
thrust::sequence(probe_keys.begin(), probe_keys.end(), 1, 2);
// Specify pattern bits for the policy
policy_t policy_A(pattern_bits_A);
bf_t filter_A(cuda::ceil_div(bits_per_key_A * build_keys.size(), bits_per_block), {}, policy_A);
filter_A.add(build_keys.begin(), build_keys.end());
filter_A.contains(probe_keys.begin(), probe_keys.end(), flags_A.begin());
size_t fps_A = thrust::count(flags_A.begin(), flags_A.end(), true);
double_t fpr_A = 100.0 * fps_A / flags_A.size();
std::cout << "FPR A: " << fpr_A << "\n";
policy_t policy_B(pattern_bits_B);
bf_t filter_B(cuda::ceil_div(bits_per_key_B * build_keys.size(), bits_per_block), {}, policy_B);
filter_B.add(build_keys.begin(), build_keys.end());
filter_B.contains(probe_keys.begin(), probe_keys.end(), flags_B.begin());
size_t fps_B = thrust::count(flags_B.begin(), flags_B.end(), true);
double_t fpr_B = 100.0 * fps_B / flags_B.size();
std::cout << "FPR B: " << fpr_B << "\n";
return 0;
}
Observed Behavior:
FPR A: 16.9311
FPR B: 0.611573
Expected Behavior:
The FPR should increase more smoothly with decreasing pattern_bits / filter size. This configuration of 8B blocks with 4 bits being set per key is common (arrow/acero) and is not expected to produce such a high FPR with a 'load factor' of 0.5.
Environment:
- Cuco version: 0.0.1
- CUDA version: 12.2
- Compiler: gcc 11.4.0
- GPU: L4
- OS: Ubuntu
Would appreciate any insights into what might be causing this! Or, if I'm missing something. Thanks!
- Ngôn ngữ chính
- Cuda
- Star
- 671
- Fork
- 122
- Merge trung bình
- 4 ngày 19 giờ
- Pull request đã merge (30 ngày)
- 10
Chuẩn bị môi trường
Khởi chạy dev container của dự án ngay trên trình duyệt, bằng tài khoản GitHub của bạn.
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/cuCollections
-
Add cuco::detail::stream_sync(cuda::stream_ref) to centralize CCCL version-specific API namingCó thể làm lại được @0z5a đã nhận 23 ngày trước và không có pull request nào đang mở. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
NVIDIA/cuCollections#840 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
nvidia-runners
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 25/100
NVIDIA/cuCollections#853 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
topic: performance type: feature request
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
NVIDIA/cuCollections#817 · 7 bình luận · 1 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày
-
good first issue P2: Nice to have type: improvement
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 38/100
NVIDIA/cuCollections#805 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[FEA] Add MPSC/MPMC concurrent queueCó thể làm lại được @sleeepyjack đã nhận 262 ngày trước và không có pull request nào đang mở. Đang mởtype: feature request
NVIDIA/cuCollections#791 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày