[Feature] Improve scan performance in hot read paths
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- cpp
- Lĩnh vực
- performance
Hướng nghiên cứu
Bắt đầu bằng cách lập hồ sơ các lệnh gọi StructArray::fields() của manifest reader và các lệnh gọi ArrayBuilder::type() của Avro decoder trong các đường dẫn quét đồng thời. Xác định thời gian tồn tại phù hợp của batch, reader hoặc builder cho metadata Arrow bất biến, sau đó xác minh rằng các cache được vô hiệu hóa khi cây đối tượng Arrow tương ứng được thay thế. Hoàn thành có nghĩa là giảm được chi phí đồng bộ hóa shared-pointer và đếm tham chiếu trong các lần quét đồng thời.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Search before asking
- I searched in the issues and found nothing similar.
Motivation
Recent profiling of highly concurrent scans has revealed several performance bottlenecks caused by
repeated operations on Arrow-returned shared_ptr objects in hot read loops.
Two significant cases have been identified:
-
Manifest readers repeatedly call
StructArray::fields()while processing individual rows. This
copiesshared_ptr<Array>objects and, with GCC 8.3's libstdc++, can introduce substantial lock
contention through_Sp_locker,pthread_mutex_lock, and futex waits when multiple workers read
manifests concurrently. -
Avro decoding calls
ArrayBuilder::type()for every integer and timestamp value. Because this
method returnsstd::shared_ptr<DataType>by value, concurrent scans repeatedly modify reference
counts on shared Arrow primitive data types, causing cache-line contention. Profiling showed
ArrayBuilder::type()and shared-pointer release operations accounting for a large proportion of
samples after the manifest bottleneck was removed.
This issue tracks the broader effort to identify and eliminate similar shared-pointer operations
from scan hot paths. The goal is to cache immutable Arrow metadata at an appropriate batch, reader,
or builder lifetime, while ensuring caches are invalidated whenever the corresponding Arrow object
tree is replaced.
The expected outcome is lower synchronization and reference-counting overhead under concurrent
scans, allowing CPU time to return to actual decoding, memory copying, and buffer management.
Solution
No response
Anything else?
No response
Are you willing to submit a PR?
- I'm willing to submit a PR!
- Ngôn ngữ chính
- C++
- Star
- 65
- Fork
- 29
- Merge trung bình
- 2 ngày 30 phút
- Pull request đã merge (30 ngày)
- 77
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/paimon-cpp
-
enhancement
apache/paimon-cpp#381 · 1 người được giao ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 30/100
apache/paimon-cpp#375 · 1 người được giao ·
-
enhancement
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 30/100
apache/paimon-cpp#369 · 1 người được giao ·
-
enhancement
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 45/100
apache/paimon-cpp#361 · 1 người được giao ·
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 68/100
apache/paimon-cpp#347 · 1 người được giao ·
Tất cả issue của apache/paimon-cpp
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
AXERA-TECH/ax-llm#77 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
games-on-whales/wolf#509 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
-
bug-unconfirmed
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
NVIDIA/cuda-samples#453 ·