Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Array-of-strings tuple sketch: key hashing is incompatible with Java (UTF-8 vs UTF-16)

已关闭
#533 5 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
68/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
cpp, java
领域
data, testing-qa

调研方向

Start in tuple/include/array_of_strings_sketch_impl.hpp at hash_array_of_strings_key, then inspect the Java stringArrHash behavior and the datasketches-tck AoS snapshots. Add or update AosSketchCrossLanguageTest to compare C++ hashes with the Java snapshots, including aos_unicode, and update the header comment. Done means the cross-language hash sets match and the existing snapshots remain readable.

由索引模型根据 Issue 内容生成。

描述

@proost, while checking cross-language binary compatibility of the tuple sketches against the snapshots in datasketches-tck, we found that the C++ array-of-strings (AoS) tuple sketch hashes keys differently from Java. For the same keys, C++ and Java produce completely different hashes, so their sketches cannot be meaningfully combined.

Evidence

Comparing the TCK snapshots aos_*_cpp.sk and aos_*_java.sk, generated from the same keys:

Case Retained (Java / C++) Hashes in common
aos_1_n10 10 / 10 0
aos_1_n1000 1000 / 1000 0
aos_multikey_n1000 1000 / 1000 0
aos_unicode 3 / 3 0

Each language reads the other's files without error, so nothing fails visibly. But a union of a Java sketch and a C++ sketch built from the same keys double-counts every key, and an intersection comes out empty.

Cause

The seed (0x7A3CCA71) and the , separator match Java, but the bytes that get hashed don't:

  • Java (tuple/Util.stringArrHash) hashes the joined key as UTF-16 code units:
    final String s = stringConcat(strArray);              // joined with ','
    return hashCharArr(s.toCharArray(), 0, s.length(), PRIME);
    
    XxHash.hashCharArr hashes each char as 2 bytes, little-endian.
  • C++ (hash_array_of_strings_key in tuple/include/array_of_strings_sketch_impl.hpp) hashes the UTF-8 bytes of each string.

Why C++ and Go should change, not Java

All three implementations need to agree. The cost of changing each one differs a lot:

  • Java has shipped its AoS sketch for years. Changing its hashing would make every AoS sketch Java users have already saved incompatible with new ones. It would also break Java's compatibility with itself across versions. Java also hashes UTF-16 deliberately: that is how Java stores strings, so it can hash them without first converting them.
  • C++ has never released its AoS sketch (#476). Changing it now affects no users.
  • Go released its AoS sketch only recently, in v0.2.0. As a pre-1.0 library, it can make breaking changes between minor versions, provided the release notes call them out.

So Java is the reference, and C++ and Go should match it.

Why this issue wasn't caught earlier

Our current cross-language tests only check that the .sk files can be read, not that the hashes match. We discovered this only recently, when we began Cross-Language Binary (CLB) testing.

Proposed fix

  1. In hash_array_of_strings_key, convert each key string from UTF-8 to UTF-16 code units, using surrogate pairs for code points above U+FFFF. Hash the code units as little-endian bytes, with , as 2c 00 between strings. Invalid UTF-8 should throw std::invalid_argument.
  2. Add a test that builds sketches in C++ from the same keys as the Java generators (AosSketchCrossLanguageTest) and requires the hash sets to equal those in the Java .sk files, including aos_unicode.
  3. Update the header comment, which currently says the hashing matches Java.

We have verified this approach: with UTF-8 to UTF-16LE conversion, C++ reproduces the hashes in every Java AoS snapshot exactly. That includes aos_unicode, whose keys contain Korean, Cyrillic and emoji outside the BMP (🔑, 🗝️), and the multikey and 1,000,000-item cases. Apart from the hashing, the format already matches: every multi-entry Java AoS snapshot round-trips through C++ byte for byte.

Optional, while in this code: compact_array_of_strings_tuple_sketch::serialize() has no default serde, unlike deserialize(). Calling it without passing default_array_of_strings_serde<>() compiles but fails at link time with an undefined serde<array<std::string>> symbol.

Timing

The C++ AoS sketch (#476) has not been released yet. We'd like to fix this before 5.3.0, which we plan to release very soon, so that the first release is compatible with Java. Otherwise, fixing it later would invalidate C++ users' saved sketches.

Could you take this on in the next few days? If you're short on time, let us know and we can prepare the PR for your review.

Go has the same issue; see apache/datasketches-go#191.

主要语言
C++
星标
274
派生
89
平均合并
1 天 11 小时
30 天内合并 PR
17

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/datasketches-cpp 的其他 Issue

查看 apache/datasketches-cpp 的全部 Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。