Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Array-of-strings tuple sketch: key hashing is incompatible with Java (UTF-8 vs UTF-16)

Đã đóng
#533 5 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
68/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
cpp, java
Lĩnh vực
data, testing-qa

Hướng nghiên cứu

Start in tuple/include/array_of_strings_sketch_impl.hpp at hash_array_of_strings_key, then inspect the Java stringArrHash behavior and the datasketches-tck AoS snapshots. Add or update AosSketchCrossLanguageTest to compare C++ hashes with the Java snapshots, including aos_unicode, and update the header comment. Done means the cross-language hash sets match and the existing snapshots remain readable.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

@proost, while checking cross-language binary compatibility of the tuple sketches against the snapshots in datasketches-tck, we found that the C++ array-of-strings (AoS) tuple sketch hashes keys differently from Java. For the same keys, C++ and Java produce completely different hashes, so their sketches cannot be meaningfully combined.

Evidence

Comparing the TCK snapshots aos_*_cpp.sk and aos_*_java.sk, generated from the same keys:

Case Retained (Java / C++) Hashes in common
aos_1_n10 10 / 10 0
aos_1_n1000 1000 / 1000 0
aos_multikey_n1000 1000 / 1000 0
aos_unicode 3 / 3 0

Each language reads the other's files without error, so nothing fails visibly. But a union of a Java sketch and a C++ sketch built from the same keys double-counts every key, and an intersection comes out empty.

Cause

The seed (0x7A3CCA71) and the , separator match Java, but the bytes that get hashed don't:

  • Java (tuple/Util.stringArrHash) hashes the joined key as UTF-16 code units:
    final String s = stringConcat(strArray);              // joined with ','
    return hashCharArr(s.toCharArray(), 0, s.length(), PRIME);
    
    XxHash.hashCharArr hashes each char as 2 bytes, little-endian.
  • C++ (hash_array_of_strings_key in tuple/include/array_of_strings_sketch_impl.hpp) hashes the UTF-8 bytes of each string.

Why C++ and Go should change, not Java

All three implementations need to agree. The cost of changing each one differs a lot:

  • Java has shipped its AoS sketch for years. Changing its hashing would make every AoS sketch Java users have already saved incompatible with new ones. It would also break Java's compatibility with itself across versions. Java also hashes UTF-16 deliberately: that is how Java stores strings, so it can hash them without first converting them.
  • C++ has never released its AoS sketch (#476). Changing it now affects no users.
  • Go released its AoS sketch only recently, in v0.2.0. As a pre-1.0 library, it can make breaking changes between minor versions, provided the release notes call them out.

So Java is the reference, and C++ and Go should match it.

Why this issue wasn't caught earlier

Our current cross-language tests only check that the .sk files can be read, not that the hashes match. We discovered this only recently, when we began Cross-Language Binary (CLB) testing.

Proposed fix

  1. In hash_array_of_strings_key, convert each key string from UTF-8 to UTF-16 code units, using surrogate pairs for code points above U+FFFF. Hash the code units as little-endian bytes, with , as 2c 00 between strings. Invalid UTF-8 should throw std::invalid_argument.
  2. Add a test that builds sketches in C++ from the same keys as the Java generators (AosSketchCrossLanguageTest) and requires the hash sets to equal those in the Java .sk files, including aos_unicode.
  3. Update the header comment, which currently says the hashing matches Java.

We have verified this approach: with UTF-8 to UTF-16LE conversion, C++ reproduces the hashes in every Java AoS snapshot exactly. That includes aos_unicode, whose keys contain Korean, Cyrillic and emoji outside the BMP (🔑, 🗝️), and the multikey and 1,000,000-item cases. Apart from the hashing, the format already matches: every multi-entry Java AoS snapshot round-trips through C++ byte for byte.

Optional, while in this code: compact_array_of_strings_tuple_sketch::serialize() has no default serde, unlike deserialize(). Calling it without passing default_array_of_strings_serde<>() compiles but fails at link time with an undefined serde<array<std::string>> symbol.

Timing

The C++ AoS sketch (#476) has not been released yet. We'd like to fix this before 5.3.0, which we plan to release very soon, so that the first release is compatible with Java. Otherwise, fixing it later would invalidate C++ users' saved sketches.

Could you take this on in the next few days? If you're short on time, let us know and we can prepare the PR for your review.

Go has the same issue; see apache/datasketches-go#191.

Ngôn ngữ chính
C++
Star
274
Fork
89
Merge trung bình
1 ngày 11 giờ
Pull request đã merge (30 ngày)
17

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của apache/datasketches-cpp

Tất cả issue của apache/datasketches-cpp

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.