Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Python images can't run default Python UDFs since Spark 4.2 (PyArrow and pandas not installed)

Đang mở Phù hợp với người mới
#137 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
1/5
Thời gian dự kiến
Dưới một giờ
Mức phù hợp với người mới
85/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
dockerfile, python
Lĩnh vực
backend

Hướng nghiên cứu

Chỉnh sửa r-python.template để thêm dòng pip install cho pyarrow và pandas trong khối HAVE_PY. Build image đã sửa đổi và xác nhận rằng một @udf đơn giản không còn gây ra ModuleNotFoundError.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Since Spark 4.2.0, regular Python UDFs are Arrow-optimized by default (SPARK-54555; spark.sql.execution.pythonUDF.arrow.enabled now defaults to true). The python3 images only install python3 and python3-pip (r-python.template), so the Python worker has neither PyArrow nor pandas, and a plain @udf with no options fails on the unmodified image:

from pyspark.sql import SparkSession, functions as F
spark = SparkSession.builder.remote("sc://localhost:15002").getOrCreate()

@F.udf("int")
def plus_one(x): return x + 1

spark.range(3).select(plus_one("id")).collect()
  File "/opt/spark/python/lib/pyspark.zip/pyspark/worker.py", line 3004, in read_udfs
    import pyarrow as pa
ModuleNotFoundError: No module named 'pyarrow'

Server: apache/spark:4.2.0-scala2.13-java21-python3-ubuntu running start-connect-server.sh --wait; client: pyspark-client==4.2.0 on Python 3.10. Results:

UDF apache/spark:4.2.0-scala2.13-java21-python3-ubuntu same image + pyarrow and pandas
@udf (default) ModuleNotFoundError: No module named 'pyarrow' OK
@udf(useArrow=False) OK OK
@pandas_udf ModuleNotFoundError: No module named 'pandas' OK

Before 4.2 a plain @udf didn't need either library, so this only started breaking with the new default. (SPARK-37554 added them to the release-build image, not these.)

Proposal: install both in the HAVE_PY block of r-python.template so 4.2.1 and 4.3.0 pick them up:

PIP_BREAK_SYSTEM_PACKAGES=1 python3 -m pip install --no-cache-dir "pyarrow>=18.0.0" "pandas>=2.2.0,<3"; \

The floors are PySpark 4.2's own requirements. PIP_BREAK_SYSTEM_PACKAGES=1 is there because the Ubuntu 26.04 base in #133 refuses a plain system pip install, and <3 because PySpark warns about pandas 3.x; tested on both jammy and resolute. It adds roughly 320 MB to the image.

That said, it's possible my understanding is off, or that there's a better way to handle this (e.g. documenting it instead of growing the image). Open to either. Thank you!

Ngôn ngữ chính
Dockerfile
Star
174
Fork
62
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của apache/spark-docker

Tất cả issue của apache/spark-docker

Issue tương tự

Thêm issue về Backend & API Design

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.