Python images can't run default Python UDFs since Spark 4.2 (PyArrow and pandas not installed)
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 1/5
- Thời gian dự kiến
- Dưới một giờ
- Mức phù hợp với người mới
- 85/100
Hướng nghiên cứu
Chỉnh sửa r-python.template để thêm dòng pip install cho pyarrow và pandas trong khối HAVE_PY. Build image đã sửa đổi và xác nhận rằng một @udf đơn giản không còn gây ra ModuleNotFoundError.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Since Spark 4.2.0, regular Python UDFs are Arrow-optimized by default (SPARK-54555; spark.sql.execution.pythonUDF.arrow.enabled now defaults to true). The python3 images only install python3 and python3-pip (r-python.template), so the Python worker has neither PyArrow nor pandas, and a plain @udf with no options fails on the unmodified image:
from pyspark.sql import SparkSession, functions as F
spark = SparkSession.builder.remote("sc://localhost:15002").getOrCreate()
@F.udf("int")
def plus_one(x): return x + 1
spark.range(3).select(plus_one("id")).collect()
File "/opt/spark/python/lib/pyspark.zip/pyspark/worker.py", line 3004, in read_udfs
import pyarrow as pa
ModuleNotFoundError: No module named 'pyarrow'
Server: apache/spark:4.2.0-scala2.13-java21-python3-ubuntu running start-connect-server.sh --wait; client: pyspark-client==4.2.0 on Python 3.10. Results:
| UDF | apache/spark:4.2.0-scala2.13-java21-python3-ubuntu |
same image + pyarrow and pandas |
|---|---|---|
@udf (default) |
ModuleNotFoundError: No module named 'pyarrow' |
OK |
@udf(useArrow=False) |
OK | OK |
@pandas_udf |
ModuleNotFoundError: No module named 'pandas' |
OK |
Before 4.2 a plain @udf didn't need either library, so this only started breaking with the new default. (SPARK-37554 added them to the release-build image, not these.)
Proposal: install both in the HAVE_PY block of r-python.template so 4.2.1 and 4.3.0 pick them up:
PIP_BREAK_SYSTEM_PACKAGES=1 python3 -m pip install --no-cache-dir "pyarrow>=18.0.0" "pandas>=2.2.0,<3"; \
The floors are PySpark 4.2's own requirements. PIP_BREAK_SYSTEM_PACKAGES=1 is there because the Ubuntu 26.04 base in #133 refuses a plain system pip install, and <3 because PySpark warns about pandas 3.x; tested on both jammy and resolute. It adds roughly 320 MB to the image.
That said, it's possible my understanding is off, or that there's a better way to handle this (e.g. documenting it instead of growing the image). Open to either. Thank you!
- Ngôn ngữ chính
- Dockerfile
- Star
- 174
- Fork
- 62
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/spark-docker
-
Upgrade base images to Ubuntu 24.04Có thể đã có người làm @RobbertDM đã nhận 26 ngày trước. Đang mở
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 68/100
apache/spark-docker#132 · 7 bình luận ·
Tất cả issue của apache/spark-docker
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Effect-TS/effect#8881 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
mail processing verified
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 7 ngày
-
feature
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 66/100
-
L: github:actions L: php:composer
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
dependabot/dependabot-core#16493 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
api-platform/core#8649 ·
Maintainer thường phản hồi trong vòng 1 ngày