[Feature] Cut metadata round trips when planning a scan (known manifest sizes, listing probes, parallel manifest-list reads)
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 58/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Domain
- data-engineering
Research direction
Start with ObjectsFile::Read, ReadIfFileExist, ReadArrowBatches, and ReadFileSegment, then trace ManifestFile, ManifestList, FileStoreScan, and SnapshotFileCollector callers. Also inspect FileUtils::ListVersionedFileStatus and JindoFileSystem::ListDir. Done means known sizes avoid metadata status calls, missing-directory listings still work, and base/delta manifest reads preserve order while running concurrently.
Written by the indexing model from the issue text.
Description
Search before asking
- I searched in the issues and found nothing similar.
Motivation
Scan planning pays several object-store round trips that are avoidable, either because the answer is already in metadata it has read or because one call can answer what two are asking.
- Manifest reads re-resolve a length the metadata already carries. A manifest list records each manifest's
fileSize, and a snapshot records its base/delta/changelog manifest-list sizes. ButObjectsFile::Read*opened these files with a bareOpen(path), so on a remote store every manifest and manifest list paid agetFileStatus/HeadObjectround trip just to learn a length planning already had. This is the metadata-path counterpart of the data-file fast path inOpen(const FileStatus&). ListVersionedFileStatusprobes existence before listing. It calledExists(dir)thenListDir(dir). Every file system already lists a missing directory as an empty result rather than an error, so the probe only decided whether to make a call that answers the same question — an extra round trip on every schema/snapshot/versioned-file lookup.- Jindo
ListDirasks the store twice. It calledExists(dir)thenGetFileStatus(dir); a singleGetFileStatusanswers both "is it there" and "is it a directory". ScanMode::ALLreads the base and delta manifest lists serially. They are two independent files and neither read depends on the other, so the two metadata round trips are paid one after the other instead of together.
For scans over many manifests, or against a high-latency object store, these round trips are a measurable and entirely avoidable part of planning latency.
Solution
- Thread an optional known length through the metadata read path:
ObjectsFile::Read/ReadIfFileExist/ReadArrowBatches/ReadFileSegmentgain astd::optional<int64_t> file_size(defaultstd::nullopt), and a newOpenForReadhelper opens withOpen(FileStatus(path, size))when the length is known and falls back toOpen(path)when it is not.ManifestFile::ReadBucketEntriesforwards it,ManifestList::ReadBase/Delta/ChangelogManifestspass the sizes recorded on the snapshot, andFileStoreScan/SnapshotFileCollectorpass eachManifestFileMeta::FileSize(). - Drop the
Exists()probe inFileUtils::ListVersionedFileStatusand list directly. - Collapse Jindo
ListDirto a singleGetFileStatus, mapping the SDK not-found to an empty listing (as the other file systems do) and propagating any other error. - In
FileStoreScan::ReadManifestsWithSnapshot, read the base and delta manifest lists concurrently through the existingexecutor_(Via+CollectAll), preserving base-then-delta order.
The size fields stay optional, so a snapshot or manifest list written before they existed keeps reading through the Open(path) fallback — this is an optimization, not a new requirement on the metadata.
Anything else?
- The length handed to
Open(FileStatus)is trusted, not re-validated, so a stale or wrong size would surface as a short or failed read. Manifest and manifest-list files are write-once and never rewritten, and the size is recorded by the same commit that wrote the file, so it cannot go stale underneath the read; a negative length is rejected by the baseOpen(const FileStatus&). - The size fast path benefits stores that override
Open(const FileStatus&)(todayObjectStoreFileSystem;ResolvingFileSystemforwards it).JindoFileSystemcurrently only overridesOpen(path), so for Jindo the manifest-size threading is inert until it gains that override; the Jindo win here is the single-callListDir. - No storage format or protocol change, and no new public API: the changed signatures are internal
src/paimonhelpers and the new parameter is defaulted.
Are you willing to submit a PR?
- I'm willing to submit a PR!
- Dominant language
- C++
- Stars
- 65
- Forks
- 29
- Avg merge
- 2d 4h
- Merged PRs (30d)
- 78
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/paimon-cpp
-
enhancement
apache/paimon-cpp#381 · 1 assignee ·
-
Difficulty 4/5 3-5 days Newbie friendliness 30/100
apache/paimon-cpp#375 · 1 assignee ·
-
enhancement
Difficulty 5/5 Over a week Newbie friendliness 45/100
apache/paimon-cpp#361 · 1 assignee ·
-
enhancement
Difficulty 4/5 3-5 days Newbie friendliness 45/100
apache/paimon-cpp#325 · 1 assignee ·
-
enhancement
Difficulty 5/5 Over a week Newbie friendliness 35/100
apache/paimon-cpp#319 · 1 reaction · 1 assignee ·
All issues in apache/paimon-cpp
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
objectionary/eo-graphs#74 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 95/100
-
enhancement
Difficulty 1/5 Under an hour Newbie friendliness 88/100
QuantStack/git2cpp#187 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100