PyHive, Presto connector returning wrong resultset
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 35/100
Direzione di ricerca
Iniziate eseguendo gli script PyHive e JDBC Python ufficiali forniti sulla stessa tabella Presto e confrontandone il numero di righe. Tracciate il comportamento del connettore PyHive durante il recupero dei risultati; il lavoro è completato quando avete identificato il motivo per cui mancano delle righe e confermato che la query restituisce il numero previsto.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
I'm using Presto Cluster for processing large amount of data.
To visualize the data I use the connector provided and suggested by the official Superset documentation, which is PyHive from the SQLAlchemy library and I'm using the default settings for the connection.
When using the provided pyhive presto connector and executing a very simple query - "SELECT * FROM test_table", the returned number of rows by the resultset is incorrect compared with the same query executed in the presto-cli app, the official connector provided by the Presto documentation.
I created two simple python scripts to test Presto connection using PyHive and the official jdbc.jar driver.
The PyHive connector returned wrong number of rows in the resultset about 817000 rows, exactly the same number of rows that was returned by the Superset chart. The connector with the official jdbc driver returned the correct amount of data - 875000 rows.
It looks like the issue is caused by the PyHive connector. Is it possible to change the connection method from PyHive to the official JDBC driver?
I'm attaching the two python scripts that I used to reproduce the issue.
#This Python script is using PyHive
from pyhive import presto
def execute_presto_query(host, port, user, catalog, schema, table, max_rows):
connection = presto.connect(host=host, port=port, username=user, catalog=catalog, schema=schema, session_props={'query_max_output_size': '1TB'})
try:
cursor = connection.cursor()
query = f"""SELECT * FROM test_table"""
cursor.execute(query)
total_rows = 0
while True:
rows = cursor.fetchmany(max_rows)
if not rows:
break
for row in rows:
total_rows += 1
print(row)
except Exception as e:
print("Error executing the query:", e)
finally:
print(total_rows)
cursor.close()
connection.close()
if __name__ == "__main__":
host = "localhost"
port = 30000
user = "testUser"
catalog = "pinot"
schema = "default"
table = "test_table"
max_rows = 1000000
execute_presto_query(host, port, user, catalog, schema, table, max_rows)
#This Python script is using the official JDBC driver
import jaydebeapi
import jpype
def execute_presto_query(host, port, user, catalog, schema, table, max_rows):
jar_file = '/home/admin1/Downloads/presto-jdbc-0.282.jar'
jpype.startJVM(jpype.getDefaultJVMPath(), "-Djava.class.path=" + jar_file)
connection_url = f'jdbc:presto://{host}:{port}/{catalog}/{schema}'
conn = jaydebeapi.connect(
'com.facebook.presto.jdbc.PrestoDriver',
connection_url,
{'user': user},
jar_file
)
try:
cursor = conn.cursor()
query = f"SELECT * FROM test_table"
cursor.execute(query)
rows = cursor.fetchall()
for row in rows:
print(row)
print(f"Total rows returned: {len(rows)}")
except Exception as e:
print("Error executing the query:", e)
finally:
cursor.close()
conn.close()
jpype.shutdownJVM()
if __name__ == "__main__":
host = "localhost"
port = 30000
user = "testUsername"
catalog = "pinot"
schema = "default"
table = "test_table"
max_rows = 1000000
execute_presto_query(host, port, user, catalog, schema, table, max_rows)
- Lingua principale
- Python
- Stelle
- 1.7k
- Fork
- 545
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di dropbox/PyHive
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 25/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 35/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 35/100
-
pyHive mTLS for NGINX proxyAperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 20/100
Tutte le issue di dropbox/PyHive
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
NousResearch/hermes-agent#136483 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last elementForse già presa @peterdsharpe l’ha presa oggi. Apertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
pytorch/tensordict#2307 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
GrokModel.generate/a_generate pass an OpenAI-style list-of-dicts to xai_sdk.chat.user(), so every call crashes with a protobuf TypeError before any network I/OForse già presa @Christian-Sidak l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
confident-ai/deepeval#3436 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno