Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Python images can't run default Python UDFs since Spark 4.2 (PyArrow and pandas not installed)

Abierto Apto para principiantes
#137 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

  • #138 de @lisancao — cerrado sin fusionar

Evaluación

Dificultad
1/5
Tiempo estimado
Menos de una hora
Aptitud para principiantes
85/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
dockerfile, python
Área
backend

Línea de trabajo

Edita r-python.template para añadir la línea de instalación de pip para pyarrow y pandas en el bloque HAVE_PY. Construye la imagen modificada y confirma que un @udf simple ya no genera ModuleNotFoundError.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Since Spark 4.2.0, regular Python UDFs are Arrow-optimized by default (SPARK-54555; spark.sql.execution.pythonUDF.arrow.enabled now defaults to true). The python3 images only install python3 and python3-pip (r-python.template), so the Python worker has neither PyArrow nor pandas, and a plain @udf with no options fails on the unmodified image:

from pyspark.sql import SparkSession, functions as F
spark = SparkSession.builder.remote("sc://localhost:15002").getOrCreate()

@F.udf("int")
def plus_one(x): return x + 1

spark.range(3).select(plus_one("id")).collect()
  File "/opt/spark/python/lib/pyspark.zip/pyspark/worker.py", line 3004, in read_udfs
    import pyarrow as pa
ModuleNotFoundError: No module named 'pyarrow'

Server: apache/spark:4.2.0-scala2.13-java21-python3-ubuntu running start-connect-server.sh --wait; client: pyspark-client==4.2.0 on Python 3.10. Results:

UDF apache/spark:4.2.0-scala2.13-java21-python3-ubuntu same image + pyarrow and pandas
@udf (default) ModuleNotFoundError: No module named 'pyarrow' OK
@udf(useArrow=False) OK OK
@pandas_udf ModuleNotFoundError: No module named 'pandas' OK

Before 4.2 a plain @udf didn't need either library, so this only started breaking with the new default. (SPARK-37554 added them to the release-build image, not these.)

Proposal: install both in the HAVE_PY block of r-python.template so 4.2.1 and 4.3.0 pick them up:

PIP_BREAK_SYSTEM_PACKAGES=1 python3 -m pip install --no-cache-dir "pyarrow>=18.0.0" "pandas>=2.2.0,<3"; \

The floors are PySpark 4.2's own requirements. PIP_BREAK_SYSTEM_PACKAGES=1 is there because the Ubuntu 26.04 base in #133 refuses a plain system pip install, and <3 because PySpark warns about pandas 3.x; tested on both jammy and resolute. It adds roughly 320 MB to the image.

That said, it's possible my understanding is off, or that there's a better way to handle this (e.g. documenting it instead of growing the image). Open to either. Thank you!

Lenguaje dominante
Dockerfile
Estrellas
176
Forks
64
Métricas de merge de PR
Sin PR fusionados en 30 d

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de apache/spark-docker

Todos los issues de apache/spark-docker

Issues similares

Más issues de Backend & API Design

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.