Slow engine creation - only when using FastAPI
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 25/100
Direzione di ricerca
Inizia con la riproduzione FastAPI in main.py e test_route.py, quindi confronta il relativo percorso create_engine con l’esempio di connessione autonomo di databricks.sql. Esegui il profiling dell’importazione e dell’esecuzione di databricks/sql/thrift_api/TCLIService/ttypes.py con le versioni fissate in requirements.txt. Il lavoro sarà completato quando sarà stata stabilita la causa del ritardo esclusivo di FastAPI e sarà stato dimostrato un rimedio supportato senza modificare manualmente il file generato.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hello,
I am currently working on a FastAPI application which calls Databricks using databricks-sql-connector, however the code appears to slow down massively on the engine creation stage.
I've narrowed the issue down to .venv\Lib\site-packages\databricks\sql\thrift_api\TCLIService\ttypes.py, where when running in debug mode the script is loaded incredibly slowly (around 5-10 minutes). However when the exact same code is run outside of a FastAPI function (ie, in a standalone script, notebook, etc) it runs almost instantly and the engine is created in under 1 second.
The details of my test application are:
main.py:
from fastapi import FastAPI
from fastapi.responses import RedirectResponse
import uvicorn
from routes import test_route
app = FastAPI(
title="TestAPI",
version="1.0",
description="Test API",
)
@app.get("/", include_in_schema=False)
async def docs_redirect():
return RedirectResponse(url='/docs')
app.include_router(test_route.router)
if __name__ == '__main__':
uvicorn.run('main:app', host='0.0.0.0', port=8000)
test_route.py:
import fastapi
from fastapi import APIRouter
from sqlalchemy import create_engine
from sqlalchemy import text
router = APIRouter(tags=["Test"])
@router.post("/test_route", description="Test")
async def test():
import time
print("Connecting to engine....")
startT = time.time()
engine = create_engine(
url = f"databricks://token:<<TOKEN>>@<<HOST>>?http_path=/sql/1.0/warehouses/<<WAREHOUSE>>&catalog=<<CATALOG>>&schema=<<SCHEMA>>"
)
print("Engine Connected!")
print(f"Time taken: {time.time() - startT} seconds.")
with engine.connect() as connection:
result = connection.execute(text("SELECT * from range(10)"))
for row in result:
print(row)
print(f"Execution time: {time.time() - startT} seconds.")
pass
return "Ok."
requirements.txt:
fastapi==0.111.1
uvicorn==0.30.3
databricks-sql-connector[SQLAlchemy]==3.3.0
The same issue occurs when using:
from databricks import sql
connection = sql.connect(
server_hostname = <HOST>,
http_path = <HTTP>
access_token = <TOKEN>)
cursor = connection.cursor()
cursor.execute("SELECT * from range(10)")
print(cursor.fetchall())
cursor.close()
connection.close()
but given it all goes to create_engine under the hood I was trying to simplify my test case.
As mentioned, the call stack points to .venv\Lib\site-packages\databricks\sql\thrift_api\TCLIService\ttypes.py taking all of this extra time. This can be fixed through editing ttypes.py directly and removing all rows which begin None, #. However this directly violates the warning given at the top of ttypes.py about not editing. Removing these rows reduces the file from ~100,000 lines to ~10,000, and completely solves the time delay issue when using FastAPI.
So my main questions are:
- Why does this delay in engine creation only occur in a FastAPI app (is it something to do with it being asynchronous?).
- How can this delay be remedied?
- Can
ttypes.pybe safely edited to remove all "None" rows or is this not a viable solution?
Thanks!
- Lingua principale
- Python
- Stelle
- 233
- Fork
- 152
- Merge medio
- 21h 5m
- PR unite (30g)
- 10
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di databricks/databricks-sql-python
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
Tutte le issue di databricks/databricks-sql-python
Issue simili
-
essnmx good first issue
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 95/100
-
[Feature] 奇物选择添加优先级 Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
syfoud/Simulated_Scepter#174 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
Giskard-AI/giskard-oss#2840 · 1 commento ·
-
A claim comment carrying the issue number is silently declined while the workflow reports success Apertaarea: repo bug perceived difficulty: 2
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
yeti-platform/yeti#1380 ·