Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Bug] Wildcard `**` + ALIGN BY DEVICE + GROUP BY + value filter returns incorrect COUNT when data spans multiple TsFiles

未关闭
#17,520 8 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
48/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
java, sql
领域
databases

调研方向

首先,在禁用压缩、使用四个由 FLUSH 分隔的批次、通配符路径、GROUP BY、ALIGN BY DEVICE 和值过滤器的条件下,复现 Query A 和 Query B。跟踪通配符查询遍历受影响的 TsFiles,并将其时间桶与精确设备查询的时间桶进行比较;完成的标准是两个查询对每个时间桶都返回完全相同的计数,包括在第一次 FLUSH 之后写入的批次。

由索引模型根据 Issue 内容生成。

描述

Search before asking
  • I searched in the issues and found nothing similar.
Version
  • OS: Linux (WSL2, kernel 6.6.87)
  • IoTDB Version: 2.0.5 (production), 2.0.8 (reproduced locally with apache/iotdb:latest Docker image)
  • Deployment: Standalone
Describe the bug and provide the minimal reproduce step
Bug Summary

When querying with a wildcard path (**), ALIGN BY DEVICE, GROUP BY time interval, and a value filter (WHERE on non-timestamp columns), some time buckets return COUNT = 0 instead of the actual count.

The bug is only triggered when the device's data is spread across multiple TsFiles (i.e., written in separate batches before compaction merges them). The same query on the exact device path always returns the correct result.

Minimal Reproduce Steps

Step 1 — Disable compaction (to keep TsFiles separate and prevent the bug from being masked):

Add to iotdb-system.properties:

enable_seq_space_compaction=false
enable_unseq_space_compaction=false
enable_cross_space_compaction=false

Step 2 — Write data in 4 separate batches with FLUSH between each:

-- Batch 1 (day 1 data)
INSERT INTO root.test.service.tenant1.device1(TIMESTAMP, method, code, msgId)
  VALUES (1776147444096, 25, null, 'aaa'),
         (1776148288048, 26, 0,    'aaa');
FLUSH;

-- Batch 2 (day 2 data)
INSERT INTO root.test.service.tenant1.device1(TIMESTAMP, method, code, msgId)
  VALUES (1776233509152, 25, null, 'bbb'),
         (1776233509201, 26, 0,    'bbb');
FLUSH;

-- Batch 3 (day 3 data — this batch will be MISSING in wildcard query)
INSERT INTO root.test.service.tenant1.device1(TIMESTAMP, method, code, msgId)
  VALUES (1776300241254, 25, null, 'ccc'),
         (1776300241302, 26, 0,    'ccc');
FLUSH;

-- Batch 4 (day 7 data — this batch will also be MISSING in wildcard query)
INSERT INTO root.test.service.tenant1.device1(TIMESTAMP, method, code, msgId)
  VALUES (1776642467102, 25, null, 'ddd'),
         (1776642467150, 26, 0,    'ddd');
FLUSH;

Confirm multiple TsFiles exist (all ending in -0-0.tsfile, meaning not yet compacted):

sequence/root.test/3/xxxx/...-1-0-0.tsfile
sequence/root.test/3/xxxx/...-2-0-0.tsfile
sequence/root.test/3/yyyy/...-1-0-0.tsfile
sequence/root.test/3/yyyy/...-2-0-0.tsfile

Step 3 — Run Query A (wildcard):

SELECT COUNT(method)
FROM root.test.service.**
WHERE method = 26 AND code = 0
GROUP BY([1773936000000, 1776700800000), 1d)
ALIGN BY DEVICE;

Step 4 — Run Query B (exact path):

SELECT COUNT(method)
FROM root.test.service.tenant1.device1
WHERE method = 26 AND code = 0
GROUP BY([1773936000000, 1776700800000), 1d)
ALIGN BY DEVICE;
What did you expect to see?

Query A and Query B return identical COUNT values for every time bucket.

What did you see instead?

Query A returns COUNT = 0 for the time buckets that correspond to Batch 3 and Batch 4 (data written into separate TsFiles after the first flush). Query B returns the correct non-zero counts for those same buckets.

Results from real production data

Device: root.test.service.39eaf2bd9ca130da3368e3ac9cfcbdff.BaV7gK84WlEZz4RGArXa92knaxTEJU

Time Bucket (UTC) Query A ** (wrong) Query B exact path (correct)
2026-04-13T16:00:00.000Z 17 17
2026-04-14T16:00:00.000Z 6 6
2026-04-15T16:00:00.000Z 0 6
2026-04-16T16:00:00.000Z 0 0
2026-04-17T16:00:00.000Z 0 0
2026-04-18T16:00:00.000Z 0 0
2026-04-19T16:00:00.000Z 0 6

The two incorrect buckets (Apr 16 and Apr 20 CST) correspond exactly to data batches written into TsFiles that were created after the first flush had already occurred.

Anything else?

Why this happens in production but is hard to reproduce in dev:

In a production IoT environment, devices continuously write data across multiple days. Each daily batch triggers a new TsFile via memtable flush. Compaction may not run fast enough to merge them before the next query arrives, so the bug is consistently observable. In a development environment with sparse data (small volume, single batch), all data lands in one TsFile and the bug never triggers.

Compaction masks the bug:

With compaction enabled (default), if the compaction task completes before the query runs, TsFiles are merged and the wildcard query returns correct results. This makes the bug appear intermittent, but it is actually deterministic: it always occurs when data spans multiple uncompacted TsFiles.

Affected versions:

Both 2.0.5 and 2.0.8 (apache/iotdb:latest as of 2026-04-20) were confirmed to reproduce the bug under identical conditions (compaction disabled, data in 4 separate TsFiles, 3+ devices under the wildcard path).

Workaround:

Query specific device paths instead of using **. Enumerate device paths from application-layer metadata and pass them explicitly to the query.

Are you willing to submit a PR?
  • I'm willing to submit a PR!
主要语言
Java
星标
6.4k
派生
1.2k
平均合并
1 天 8 小时
30 天内合并 PR
129

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/iotdb 的其他 Issue

查看 apache/iotdb 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。