Native blocks without a JVM input publish SQL metrics on every batch, ignoring spark.comet.metrics.updateInterval
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 78/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- rust
- Lĩnh vực
- backend, performance
Hướng nghiên cứu
Bắt đầu trong native/core/src/execution/jni_api.rs tại nhánh batch_receiver và so sánh đường dẫn xuất bản metric của nó với đường dẫn đầu vào JVM. Tái hiện bằng spark.comet.metrics.updateInterval=-1 và spark.comet.batchSize=1000, sau đó xác minh rằng các block thuần native chỉ xuất bản metric theo khoảng thời gian đã cấu hình và releasePlan xuất bản các giá trị cuối cùng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
When a native block has no JVM input, meaning every leaf is a native scan, executePlan publishes the whole metric tree to the JVM after every output batch (the batch_receiver branch in native/core/src/execution/jni_api.rs). Blocks with a JVM input publish only after spark.comet.metrics.updateInterval has passed. So pure-native blocks ignore the setting, including a negative value, which the config doc says means "metrics will be updated upon task completion".
This came in with #3553 (0.14.0), which moved pure-native blocks onto a channel. Before that, both kinds of block checked the interval.
Each publish walks the plan, runs aggregate_by_name over every node's metrics, encodes a protobuf, and calls into the JVM, which decodes it and sets each SQLMetric. A native Parquet scan registers about 45 metrics for each file it opens, so the cost grows with the number of files a task has read.
I timed update_metrics in place on a release build (Apple M3 Max, local[1], 16M rows in 2,048 output batches from one task):
| Plan | Cost per publish | Per task |
|---|---|---|
| Scan only, 1 file | 19 µs | 40 ms |
| Scan + filter + project | 30 µs | 62 ms |
| Scan only, 64 files in one task (2,905 raw metrics) | 38 µs | 78 ms |
Scan + filter + project, noop write |
26 µs | 54 ms |
Switching that branch to the interval check gave these medians of 7 runs:
| Plan | Wall time | Process CPU time |
|---|---|---|
| Scan only, 1 file | 174 → 143 ms | 343 → 209 ms |
| Scan + filter + project | 714 → 716 ms | 882 → 834 ms |
| Scan only, 64 files in one task | 210 → 167 ms | 313 → 220 ms |
Scan + filter + project, noop write |
917 → 848 ms | 1608 → 1505 ms |
The scan + filter + project wall time doesn't move because the producer is the bottleneck there, and the publish runs on the consumer thread. It still spends CPU that other tasks on the executor could use.
This affects pure-native blocks whose output reaches the JVM batch by batch: a scan feeding a Spark write or a collect, a broadcast build side, JVM shuffle, or a fallback operator. A block that ends in a native shuffle write hands back only one batch, so it isn't affected.
Steps to reproduce
Set spark.comet.metrics.updateInterval=-1 and spark.comet.batchSize=1000, and read a 10,000-row Parquet file in one task. Inside the task, read the native scan's output_rows metric after the first batch. It is already non-zero, when it should stay 0 until the iterator closes.
Expected behavior
Pure-native blocks publish on the configured interval, as blocks with a JVM input do, and releasePlan publishes the final values.
Additional context
Found while looking at #1381.
- Ngôn ngữ chính
- Scala
- Star
- 1.3k
- Fork
- 377
- Merge trung bình
- 2 ngày 5 giờ
- Pull request đã merge (30 ngày)
- 337
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/datafusion-comet
-
arrays_zip with two same-named inputs fails with "ArrowArray struct has 2 children (expected 1)"Có thể đã có người làm @mohitgurav20 đã nhận 2 ngày trước. Đang mởarea:expressions area:ffi bug priority:high
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/datafusion-comet#6251 · 4 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area:ci bug priority:low
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
apache/datafusion-comet#6060 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
requires-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
apache/datafusion-comet#5661 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area:scan enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
apache/datafusion-comet#5319 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area:ci priority:low
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 66/100
apache/datafusion-comet#4586 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của apache/datafusion-comet
Issue tương tự
-
Code Cleanup dependencies
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 68/100
ProjectSidewalk/SidewalkWebpage#5609 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
apache/texera#8775 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
go 🏃 testing 🧪
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
valkey-io/valkey-glide#7239 ·
Maintainer thường phản hồi trong vòng 3 ngày
-
itype:bug stat:needs triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày