Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Array-of-strings tuple sketch: key hashing is incompatible with Java (UTF-8 vs UTF-16)

クローズ
#533 コメント 5 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
68/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
cpp, java
領域
data, testing-qa

調査の方向性

Start in tuple/include/array_of_strings_sketch_impl.hpp at hash_array_of_strings_key, then inspect the Java stringArrHash behavior and the datasketches-tck AoS snapshots. Add or update AosSketchCrossLanguageTest to compare C++ hashes with the Java snapshots, including aos_unicode, and update the header comment. Done means the cross-language hash sets match and the existing snapshots remain readable.

索引モデルが issue の本文から書いたものです。

説明

@proost, while checking cross-language binary compatibility of the tuple sketches against the snapshots in datasketches-tck, we found that the C++ array-of-strings (AoS) tuple sketch hashes keys differently from Java. For the same keys, C++ and Java produce completely different hashes, so their sketches cannot be meaningfully combined.

Evidence

Comparing the TCK snapshots aos_*_cpp.sk and aos_*_java.sk, generated from the same keys:

Case Retained (Java / C++) Hashes in common
aos_1_n10 10 / 10 0
aos_1_n1000 1000 / 1000 0
aos_multikey_n1000 1000 / 1000 0
aos_unicode 3 / 3 0

Each language reads the other's files without error, so nothing fails visibly. But a union of a Java sketch and a C++ sketch built from the same keys double-counts every key, and an intersection comes out empty.

Cause

The seed (0x7A3CCA71) and the , separator match Java, but the bytes that get hashed don't:

  • Java (tuple/Util.stringArrHash) hashes the joined key as UTF-16 code units:
    final String s = stringConcat(strArray);              // joined with ','
    return hashCharArr(s.toCharArray(), 0, s.length(), PRIME);
    
    XxHash.hashCharArr hashes each char as 2 bytes, little-endian.
  • C++ (hash_array_of_strings_key in tuple/include/array_of_strings_sketch_impl.hpp) hashes the UTF-8 bytes of each string.

Why C++ and Go should change, not Java

All three implementations need to agree. The cost of changing each one differs a lot:

  • Java has shipped its AoS sketch for years. Changing its hashing would make every AoS sketch Java users have already saved incompatible with new ones. It would also break Java's compatibility with itself across versions. Java also hashes UTF-16 deliberately: that is how Java stores strings, so it can hash them without first converting them.
  • C++ has never released its AoS sketch (#476). Changing it now affects no users.
  • Go released its AoS sketch only recently, in v0.2.0. As a pre-1.0 library, it can make breaking changes between minor versions, provided the release notes call them out.

So Java is the reference, and C++ and Go should match it.

Why this issue wasn't caught earlier

Our current cross-language tests only check that the .sk files can be read, not that the hashes match. We discovered this only recently, when we began Cross-Language Binary (CLB) testing.

Proposed fix

  1. In hash_array_of_strings_key, convert each key string from UTF-8 to UTF-16 code units, using surrogate pairs for code points above U+FFFF. Hash the code units as little-endian bytes, with , as 2c 00 between strings. Invalid UTF-8 should throw std::invalid_argument.
  2. Add a test that builds sketches in C++ from the same keys as the Java generators (AosSketchCrossLanguageTest) and requires the hash sets to equal those in the Java .sk files, including aos_unicode.
  3. Update the header comment, which currently says the hashing matches Java.

We have verified this approach: with UTF-8 to UTF-16LE conversion, C++ reproduces the hashes in every Java AoS snapshot exactly. That includes aos_unicode, whose keys contain Korean, Cyrillic and emoji outside the BMP (🔑, 🗝️), and the multikey and 1,000,000-item cases. Apart from the hashing, the format already matches: every multi-entry Java AoS snapshot round-trips through C++ byte for byte.

Optional, while in this code: compact_array_of_strings_tuple_sketch::serialize() has no default serde, unlike deserialize(). Calling it without passing default_array_of_strings_serde<>() compiles but fails at link time with an undefined serde<array<std::string>> symbol.

Timing

The C++ AoS sketch (#476) has not been released yet. We'd like to fix this before 5.3.0, which we plan to release very soon, so that the first release is compatible with Java. Otherwise, fixing it later would invalidate C++ users' saved sketches.

Could you take this on in the next few days? If you're short on time, let us know and we can prepare the PR for your review.

Go has the same issue; see apache/datasketches-go#191.

主要言語
C++
スター
274
フォーク
89
平均マージ
1日 11時間
マージ済み PR(30日)
17

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

apache/datasketches-cpp のほかの issue

apache/datasketches-cpp の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。