Python images can't run default Python UDFs since Spark 4.2 (PyArrow and pandas not installed)
Nadie ha tomado este issue todavía.
- #138 de @lisancao — cerrado sin fusionar
Evaluación
- Dificultad
- 1/5
- Tiempo estimado
- Menos de una hora
- Aptitud para principiantes
- 85/100
Línea de trabajo
Edita r-python.template para añadir la línea de instalación de pip para pyarrow y pandas en el bloque HAVE_PY. Construye la imagen modificada y confirma que un @udf simple ya no genera ModuleNotFoundError.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Since Spark 4.2.0, regular Python UDFs are Arrow-optimized by default (SPARK-54555; spark.sql.execution.pythonUDF.arrow.enabled now defaults to true). The python3 images only install python3 and python3-pip (r-python.template), so the Python worker has neither PyArrow nor pandas, and a plain @udf with no options fails on the unmodified image:
from pyspark.sql import SparkSession, functions as F
spark = SparkSession.builder.remote("sc://localhost:15002").getOrCreate()
@F.udf("int")
def plus_one(x): return x + 1
spark.range(3).select(plus_one("id")).collect()
File "/opt/spark/python/lib/pyspark.zip/pyspark/worker.py", line 3004, in read_udfs
import pyarrow as pa
ModuleNotFoundError: No module named 'pyarrow'
Server: apache/spark:4.2.0-scala2.13-java21-python3-ubuntu running start-connect-server.sh --wait; client: pyspark-client==4.2.0 on Python 3.10. Results:
| UDF | apache/spark:4.2.0-scala2.13-java21-python3-ubuntu |
same image + pyarrow and pandas |
|---|---|---|
@udf (default) |
ModuleNotFoundError: No module named 'pyarrow' |
OK |
@udf(useArrow=False) |
OK | OK |
@pandas_udf |
ModuleNotFoundError: No module named 'pandas' |
OK |
Before 4.2 a plain @udf didn't need either library, so this only started breaking with the new default. (SPARK-37554 added them to the release-build image, not these.)
Proposal: install both in the HAVE_PY block of r-python.template so 4.2.1 and 4.3.0 pick them up:
PIP_BREAK_SYSTEM_PACKAGES=1 python3 -m pip install --no-cache-dir "pyarrow>=18.0.0" "pandas>=2.2.0,<3"; \
The floors are PySpark 4.2's own requirements. PIP_BREAK_SYSTEM_PACKAGES=1 is there because the Ubuntu 26.04 base in #133 refuses a plain system pip install, and <3 because PySpark warns about pandas 3.x; tested on both jammy and resolute. It adds roughly 320 MB to the image.
That said, it's possible my understanding is off, or that there's a better way to handle this (e.g. documenting it instead of growing the image). Open to either. Thank you!
- Lenguaje dominante
- Dockerfile
- Estrellas
- 176
- Forks
- 64
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de apache/spark-docker
-
Upgrade base images to Ubuntu 24.04Posiblemente ocupada @RobbertDM la tomó hace 28 días. Abierto
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
apache/spark-docker#132 · 7 comentarios ·
Todos los issues de apache/spark-docker
Issues similares
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
Code Quality
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Automattic/safe-publish#708 ·
Los mantenedores suelen responder en 1 día
-
gcsartifact: deleting a missing version returns an errorPosiblemente ocupada @ktsoator la tomó hoy. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 2 días
-
Dificultad 1/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
-
enhancement pkg:sdk
Dificultad 2/5 1-3 horas Aptitud para principiantes 80/100
aws/aws-durable-execution-sdk-go#144 ·
Los mantenedores suelen responder en 1 día