[Feature] Cut metadata round trips when planning a scan (known manifest sizes, listing probes, parallel manifest-list reads)
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 58/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
调研方向
从 ObjectsFile::Read、ReadIfFileExist、ReadArrowBatches 和 ReadFileSegment 开始,然后跟踪 ManifestFile、ManifestList、FileStoreScan 和 SnapshotFileCollector 的调用方。还要检查 FileUtils::ListVersionedFileStatus 和 JindoFileSystem::ListDir。完成的标准是:已知大小时避免 metadata status 调用,缺失目录的列表操作仍能正常工作,并且 base/delta manifest 的读取在并发运行时保持顺序。
由索引模型根据 Issue 内容生成。
描述
Search before asking
- I searched in the issues and found nothing similar.
Motivation
Scan planning pays several object-store round trips that are avoidable, either because the answer is already in metadata it has read or because one call can answer what two are asking.
- Manifest reads re-resolve a length the metadata already carries. A manifest list records each manifest's
fileSize, and a snapshot records its base/delta/changelog manifest-list sizes. ButObjectsFile::Read*opened these files with a bareOpen(path), so on a remote store every manifest and manifest list paid agetFileStatus/HeadObjectround trip just to learn a length planning already had. This is the metadata-path counterpart of the data-file fast path inOpen(const FileStatus&). ListVersionedFileStatusprobes existence before listing. It calledExists(dir)thenListDir(dir). Every file system already lists a missing directory as an empty result rather than an error, so the probe only decided whether to make a call that answers the same question — an extra round trip on every schema/snapshot/versioned-file lookup.- Jindo
ListDirasks the store twice. It calledExists(dir)thenGetFileStatus(dir); a singleGetFileStatusanswers both "is it there" and "is it a directory". ScanMode::ALLreads the base and delta manifest lists serially. They are two independent files and neither read depends on the other, so the two metadata round trips are paid one after the other instead of together.
For scans over many manifests, or against a high-latency object store, these round trips are a measurable and entirely avoidable part of planning latency.
Solution
- Thread an optional known length through the metadata read path:
ObjectsFile::Read/ReadIfFileExist/ReadArrowBatches/ReadFileSegmentgain astd::optional<int64_t> file_size(defaultstd::nullopt), and a newOpenForReadhelper opens withOpen(FileStatus(path, size))when the length is known and falls back toOpen(path)when it is not.ManifestFile::ReadBucketEntriesforwards it,ManifestList::ReadBase/Delta/ChangelogManifestspass the sizes recorded on the snapshot, andFileStoreScan/SnapshotFileCollectorpass eachManifestFileMeta::FileSize(). - Drop the
Exists()probe inFileUtils::ListVersionedFileStatusand list directly. - Collapse Jindo
ListDirto a singleGetFileStatus, mapping the SDK not-found to an empty listing (as the other file systems do) and propagating any other error. - In
FileStoreScan::ReadManifestsWithSnapshot, read the base and delta manifest lists concurrently through the existingexecutor_(Via+CollectAll), preserving base-then-delta order.
The size fields stay optional, so a snapshot or manifest list written before they existed keeps reading through the Open(path) fallback — this is an optimization, not a new requirement on the metadata.
Anything else?
- The length handed to
Open(FileStatus)is trusted, not re-validated, so a stale or wrong size would surface as a short or failed read. Manifest and manifest-list files are write-once and never rewritten, and the size is recorded by the same commit that wrote the file, so it cannot go stale underneath the read; a negative length is rejected by the baseOpen(const FileStatus&). - The size fast path benefits stores that override
Open(const FileStatus&)(todayObjectStoreFileSystem;ResolvingFileSystemforwards it).JindoFileSystemcurrently only overridesOpen(path), so for Jindo the manifest-size threading is inert until it gains that override; the Jindo win here is the single-callListDir. - No storage format or protocol change, and no new public API: the changed signatures are internal
src/paimonhelpers and the new parameter is defaulted.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- 主要语言
- C++
- 星标
- 65
- 派生
- 29
- 平均合并
- 2 天 11 小时
- 30 天内合并 PR
- 79
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/paimon-cpp 的其他 Issue
-
bug
apache/paimon-cpp#385 · 已指派 1 人 ·
-
enhancement
apache/paimon-cpp#381 · 已指派 1 人 ·
-
难度 4/5 3-5 天 新手友好度 30/100
apache/paimon-cpp#375 · 已指派 1 人 ·
-
enhancement
难度 5/5 一周以上 新手友好度 45/100
apache/paimon-cpp#361 · 已指派 1 人 ·
-
enhancement
难度 4/5 3-5 天 新手友好度 45/100
apache/paimon-cpp#325 · 已指派 1 人 ·
查看 apache/paimon-cpp 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 70/100
-
难度 2/5 1-3 小时 新手友好度 65/100
duckdb/duckdb-wasm#2258 ·
-
难度 2/5 1-3 小时 新手友好度 75/100
objectionary/eo-graphs#75 ·
-
Coarray integration tests carry no LABELS, so run_tests.py silently skips them under every backend 未关闭coarray
难度 2/5 1-3 小时 新手友好度 70/100
-
难度 2/5 1-3 小时 新手友好度 75/100
FISCO-BCOS/FISCO-BCOS#5642 ·