PyHive, Presto connector returning wrong resultset
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 35/100
Línea de trabajo
Comience ejecutando los scripts proporcionados de PyHive y JDBC Python oficial contra la misma tabla de Presto y compare sus recuentos de filas. Rastree el comportamiento del conector de PyHive al obtener los resultados; el trabajo estará terminado cuando identifique por qué faltan filas y confirme que la consulta devuelve el recuento esperado.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
I'm using Presto Cluster for processing large amount of data.
To visualize the data I use the connector provided and suggested by the official Superset documentation, which is PyHive from the SQLAlchemy library and I'm using the default settings for the connection.
When using the provided pyhive presto connector and executing a very simple query - "SELECT * FROM test_table", the returned number of rows by the resultset is incorrect compared with the same query executed in the presto-cli app, the official connector provided by the Presto documentation.
I created two simple python scripts to test Presto connection using PyHive and the official jdbc.jar driver.
The PyHive connector returned wrong number of rows in the resultset about 817000 rows, exactly the same number of rows that was returned by the Superset chart. The connector with the official jdbc driver returned the correct amount of data - 875000 rows.
It looks like the issue is caused by the PyHive connector. Is it possible to change the connection method from PyHive to the official JDBC driver?
I'm attaching the two python scripts that I used to reproduce the issue.
#This Python script is using PyHive
from pyhive import presto
def execute_presto_query(host, port, user, catalog, schema, table, max_rows):
connection = presto.connect(host=host, port=port, username=user, catalog=catalog, schema=schema, session_props={'query_max_output_size': '1TB'})
try:
cursor = connection.cursor()
query = f"""SELECT * FROM test_table"""
cursor.execute(query)
total_rows = 0
while True:
rows = cursor.fetchmany(max_rows)
if not rows:
break
for row in rows:
total_rows += 1
print(row)
except Exception as e:
print("Error executing the query:", e)
finally:
print(total_rows)
cursor.close()
connection.close()
if __name__ == "__main__":
host = "localhost"
port = 30000
user = "testUser"
catalog = "pinot"
schema = "default"
table = "test_table"
max_rows = 1000000
execute_presto_query(host, port, user, catalog, schema, table, max_rows)
#This Python script is using the official JDBC driver
import jaydebeapi
import jpype
def execute_presto_query(host, port, user, catalog, schema, table, max_rows):
jar_file = '/home/admin1/Downloads/presto-jdbc-0.282.jar'
jpype.startJVM(jpype.getDefaultJVMPath(), "-Djava.class.path=" + jar_file)
connection_url = f'jdbc:presto://{host}:{port}/{catalog}/{schema}'
conn = jaydebeapi.connect(
'com.facebook.presto.jdbc.PrestoDriver',
connection_url,
{'user': user},
jar_file
)
try:
cursor = conn.cursor()
query = f"SELECT * FROM test_table"
cursor.execute(query)
rows = cursor.fetchall()
for row in rows:
print(row)
print(f"Total rows returned: {len(rows)}")
except Exception as e:
print("Error executing the query:", e)
finally:
cursor.close()
conn.close()
jpype.shutdownJVM()
if __name__ == "__main__":
host = "localhost"
port = 30000
user = "testUsername"
catalog = "pinot"
schema = "default"
table = "test_table"
max_rows = 1000000
execute_presto_query(host, port, user, catalog, schema, table, max_rows)
- Lenguaje dominante
- Python
- Estrellas
- 1.7k
- Forks
- 545
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de dropbox/PyHive
-
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
-
Dificultad 3/5 1-2 días Aptitud para principiantes 35/100
-
Dificultad 3/5 1-2 días Aptitud para principiantes 35/100
-
pyHive mTLS for NGINX proxyAbierto
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 20/100
Todos los issues de dropbox/PyHive
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
EvaluationSuite.run fails with default args_for_task and mutates supplied kwargsPosiblemente ocupada @ktz03 la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
huggingface/evaluate#825 ·
Los mantenedores suelen responder en 1 día
-
Add `django-upgrade` to the CIAbiertodependencies feature github_actions good first issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
wemake-services/wemake-django-template#3149 ·
Los mantenedores suelen responder en 1 día
-
[request] vsg/1.1.16Abiertoupstream update
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
conan-io/conan-center-index#31142 ·
Los mantenedores suelen responder en 1 día
-
area:core bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día