Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Feature] Improve scan performance in hot read paths

Open
#240 1 comment 0 reactions 1 assignee View on GitHub

@gripleaf is already working on this.

Since Sep 21, 2026.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Needs clarification
Activity status
Active
Tech stack
cpp
Domain
performance

Research direction

Start by profiling the manifest reader's StructArray::fields() calls and the Avro decoder's ArrayBuilder::type() calls in concurrent scan paths. Identify the appropriate batch, reader, or builder lifetime for immutable Arrow metadata, then verify caches are invalidated when the corresponding Arrow object tree is replaced. Done means reduced shared-pointer synchronization and reference-counting overhead under concurrent scans.

Written by the indexing model from the issue text.

Description

enhancement
Search before asking
  • I searched in the issues and found nothing similar.
Motivation

Recent profiling of highly concurrent scans has revealed several performance bottlenecks caused by
repeated operations on Arrow-returned shared_ptr objects in hot read loops.

Two significant cases have been identified:

  1. Manifest readers repeatedly call StructArray::fields() while processing individual rows. This
    copies shared_ptr<Array> objects and, with GCC 8.3's libstdc++, can introduce substantial lock
    contention through _Sp_locker, pthread_mutex_lock, and futex waits when multiple workers read
    manifests concurrently.

  2. Avro decoding calls ArrayBuilder::type() for every integer and timestamp value. Because this
    method returns std::shared_ptr<DataType> by value, concurrent scans repeatedly modify reference
    counts on shared Arrow primitive data types, causing cache-line contention. Profiling showed
    ArrayBuilder::type() and shared-pointer release operations accounting for a large proportion of
    samples after the manifest bottleneck was removed.

This issue tracks the broader effort to identify and eliminate similar shared-pointer operations
from scan hot paths. The goal is to cache immutable Arrow metadata at an appropriate batch, reader,
or builder lifetime, while ensuring caches are invalidated whenever the corresponding Arrow object
tree is replaced.

The expected outcome is lower synchronization and reference-counting overhead under concurrent
scans, allowing CPU time to return to actual decoding, memory copying, and buffer management.

Solution

No response

Anything else?

No response

Are you willing to submit a PR?
  • I'm willing to submit a PR!
Dominant language
C++
Stars
65
Forks
29
Avg merge
2d 4h
Merged PRs (30d)
78

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/paimon-cpp

All issues in apache/paimon-cpp

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.