Python images can't run default Python UDFs since Spark 4.2 (PyArrow and pandas not installed)
まだ誰も着手していません。
評価
- 難易度
- 1/5
- 見積もり時間
- 1時間未満
- 初心者へのやさしさ
- 85/100
調査の方向性
r-python.templateを編集し、HAVE_PYブロックにpyarrowとpandasのpip install行を追加してください。変更したイメージをビルドし、単純な@udfでModuleNotFoundErrorが発生しなくなったことを確認してください。
索引モデルが issue の本文から書いたものです。
説明
Since Spark 4.2.0, regular Python UDFs are Arrow-optimized by default (SPARK-54555; spark.sql.execution.pythonUDF.arrow.enabled now defaults to true). The python3 images only install python3 and python3-pip (r-python.template), so the Python worker has neither PyArrow nor pandas, and a plain @udf with no options fails on the unmodified image:
from pyspark.sql import SparkSession, functions as F
spark = SparkSession.builder.remote("sc://localhost:15002").getOrCreate()
@F.udf("int")
def plus_one(x): return x + 1
spark.range(3).select(plus_one("id")).collect()
File "/opt/spark/python/lib/pyspark.zip/pyspark/worker.py", line 3004, in read_udfs
import pyarrow as pa
ModuleNotFoundError: No module named 'pyarrow'
Server: apache/spark:4.2.0-scala2.13-java21-python3-ubuntu running start-connect-server.sh --wait; client: pyspark-client==4.2.0 on Python 3.10. Results:
| UDF | apache/spark:4.2.0-scala2.13-java21-python3-ubuntu |
same image + pyarrow and pandas |
|---|---|---|
@udf (default) |
ModuleNotFoundError: No module named 'pyarrow' |
OK |
@udf(useArrow=False) |
OK | OK |
@pandas_udf |
ModuleNotFoundError: No module named 'pandas' |
OK |
Before 4.2 a plain @udf didn't need either library, so this only started breaking with the new default. (SPARK-37554 added them to the release-build image, not these.)
Proposal: install both in the HAVE_PY block of r-python.template so 4.2.1 and 4.3.0 pick them up:
PIP_BREAK_SYSTEM_PACKAGES=1 python3 -m pip install --no-cache-dir "pyarrow>=18.0.0" "pandas>=2.2.0,<3"; \
The floors are PySpark 4.2's own requirements. PIP_BREAK_SYSTEM_PACKAGES=1 is there because the Ubuntu 26.04 base in #133 refuses a plain system pip install, and <3 because PySpark warns about pandas 3.x; tested on both jammy and resolute. It adds roughly 320 MB to the image.
That said, it's possible my understanding is off, or that there's a better way to handle this (e.g. documenting it instead of growing the image). Open to either. Thank you!
- 主要言語
- Dockerfile
- スター
- 174
- フォーク
- 62
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
apache/spark-docker のほかの issue
-
Upgrade base images to Ubuntu 24.04対応中かも @RobbertDM が 26 日前に担当しました。 オープン
難易度 3/5 1〜2日 初心者へのやさしさ 68/100
apache/spark-docker#132 · コメント 7 件 ·
apache/spark-docker の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
glpi-project/glpi#25842 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 84/100
vllm-project/recipes#1081 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
CommunityToolkit/Aspire#2231 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
volcengine/OpenViking#5711 ·
メンテナーはふだん 1 日以内に返信