Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Python images can't run default Python UDFs since Spark 4.2 (PyArrow and pandas not installed)

オープン 初心者向け
#137 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
1/5
見積もり時間
1時間未満
初心者へのやさしさ
85/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
dockerfile, python
領域
backend

調査の方向性

r-python.templateを編集し、HAVE_PYブロックにpyarrowとpandasのpip install行を追加してください。変更したイメージをビルドし、単純な@udfでModuleNotFoundErrorが発生しなくなったことを確認してください。

索引モデルが issue の本文から書いたものです。

説明

Since Spark 4.2.0, regular Python UDFs are Arrow-optimized by default (SPARK-54555; spark.sql.execution.pythonUDF.arrow.enabled now defaults to true). The python3 images only install python3 and python3-pip (r-python.template), so the Python worker has neither PyArrow nor pandas, and a plain @udf with no options fails on the unmodified image:

from pyspark.sql import SparkSession, functions as F
spark = SparkSession.builder.remote("sc://localhost:15002").getOrCreate()

@F.udf("int")
def plus_one(x): return x + 1

spark.range(3).select(plus_one("id")).collect()
  File "/opt/spark/python/lib/pyspark.zip/pyspark/worker.py", line 3004, in read_udfs
    import pyarrow as pa
ModuleNotFoundError: No module named 'pyarrow'

Server: apache/spark:4.2.0-scala2.13-java21-python3-ubuntu running start-connect-server.sh --wait; client: pyspark-client==4.2.0 on Python 3.10. Results:

UDF apache/spark:4.2.0-scala2.13-java21-python3-ubuntu same image + pyarrow and pandas
@udf (default) ModuleNotFoundError: No module named 'pyarrow' OK
@udf(useArrow=False) OK OK
@pandas_udf ModuleNotFoundError: No module named 'pandas' OK

Before 4.2 a plain @udf didn't need either library, so this only started breaking with the new default. (SPARK-37554 added them to the release-build image, not these.)

Proposal: install both in the HAVE_PY block of r-python.template so 4.2.1 and 4.3.0 pick them up:

PIP_BREAK_SYSTEM_PACKAGES=1 python3 -m pip install --no-cache-dir "pyarrow>=18.0.0" "pandas>=2.2.0,<3"; \

The floors are PySpark 4.2's own requirements. PIP_BREAK_SYSTEM_PACKAGES=1 is there because the Ubuntu 26.04 base in #133 refuses a plain system pip install, and <3 because PySpark warns about pandas 3.x; tested on both jammy and resolute. It adds roughly 320 MB to the image.

That said, it's possible my understanding is off, or that there's a better way to handle this (e.g. documenting it instead of growing the image). Open to either. Thank you!

主要言語
Dockerfile
スター
174
フォーク
62
PR マージ指標
30日以内にマージされた PR はありません

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

apache/spark-docker のほかの issue

apache/spark-docker の issue をすべて見る

似ている issue

Backend & API Design の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。