Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

get_columns() missing @reflection.cache causes a warehouse round-trip on every reflection call

Aperta Adatta ai principianti
#75 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

@TangoEnSkai ci sta già lavorando.

Dal 23/8/2026.

  • #74 di @TangoEnSkai — aperta

Valutazione

Difficoltà
1/5
Tempo stimato
Meno di un'ora
Idoneità per principianti
92/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python, sqlalchemy
Ambito
databases

Direzione di ricerca

Iniziate in src/databricks/sqlalchemy/base.py, in get_columns(), quindi confrontate la relativa gestione della reflection con get_pk_constraint() e i dialect upstream menzionati nell’issue. Eseguite chiamate ripetute con lo stesso info_cache e verificate che la cache venga popolata e che il round-trip verso il warehouse avvenga una sola volta.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Summary

DatabricksDialect.get_columns() is missing the @reflection.cache decorator, so it never participates in SQLAlchemy's reflection cache. Every call issues a fresh GetColumns round-trip to the warehouse, however many times the same table is reflected through the same Inspector.

This is the sibling of #72 (get_foreign_keys()), which is addressed in #74. get_columns() is a separate occurrence of the same omission and is not covered by that PR.

Expected behaviour

Like get_pk_constraint(), has_table(), get_table_names(), get_view_names(), get_materialized_view_names(), get_temp_view_names(), get_schema_names() and get_table_comment() — and like get_columns() in SQLAlchemy's own SQLite, PostgreSQL and MySQL dialects, all three of which are decorated — repeated reflection of the same table through one Inspector should cost one round-trip, not one per call.

Inspector.get_columns() explicitly threads the cache through to the dialect (sqlalchemy/engine/reflection.py):

with self._operation_context() as conn:
    col_defs = self.dialect.get_columns(
        conn, table_name, schema, info_cache=self.info_cache, **kw
    )

The info_cache kwarg arrives, lands in **kwargs, and is discarded.

Actual behaviour

Reproduced against main (SQLAlchemy 2.0.52), stubbing out the transport so the call count is directly observable:

info_cache = {}
for _ in range(3):
    dialect.get_columns(conn, "t", None, info_cache=info_cache)
method calls server round-trips info_cache keys
get_columns() 3 3 [] — never populated
get_pk_constraint() 3 1 [('get_pk_constraint', ('t',), (('schema', None),))]

Note the cost is not always a single statement: when cur.columns() returns an empty list, get_columns() follows up with DESCRIBE TABLE EXTENDED to distinguish a genuinely column-less table from a missing one (base.py:156). For such tables an uncached call is two round-trips, repeated every time.

Callers that reuse one Inspector across many lookups feel this directly — Alembic's autogenerate is the common case, as is any long-lived application that reflects per request.

Root cause

src/databricks/sqlalchemy/base.py:139 — get_columns() lacks @reflection.cache. Compare get_pk_constraint() at line 212, which is decorated.

Fix

Add @reflection.cache to get_columns().

Two things worth noting for whoever picks this up, both checked:

  • Caching is safe with respect to Inspector._instantiate_types(), which mutates the returned column dicts in place. It is guarded by if not isinstance(coltype, TypeEngine), so re-running it over an already-instantiated cached list is a no-op. This is the same situation the upstream dialects are in.
  • get_indexes() is also undecorated but returns the EMPTY_INDEX constant without touching the server, so it needs no cache. get_columns() is the only remaining method where the omission costs a round-trip.

I'm happy to open a PR for this if it's welcome.

Lingua principale
Python
Stelle
24
Fork
20
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di databricks/databricks-sqlalchemy

Tutte le issue di databricks/databricks-sqlalchemy

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.