FsspecFileIO: `_adls` mutates shared properties, so a second storage account gets the first account's filesystem

Aperta Adatta ai principianti
#3,885 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
78/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
backend, cloud

Direzione di ricerca

Inizia in pyiceberg/io/fsspec.py su _adls e sul suo chiamante intorno alla riga 515, quindi esegui la riproduzione in due posizioni con il AzureBlobFileSystem simulato. Verifica che ogni hostname produca il proprio account e che io.properties rimanga invariato dopo entrambe le letture.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Apache Iceberg version

main (development)

Please describe the bug 🐞

_adls writes the account name it infers back into the properties dict it is handed, and the call site passes the FileIO's own self.properties. So the first ADLS location a FsspecFileIO touches pins adls.account-name for the life of that FileIO, and every later location is served a filesystem built for the first account.

In pyiceberg/io/fsspec.py:

def _adls(properties: Properties, hostname: str | None = None) -> AbstractFileSystem:
    ...
    # Fallback: extract account_name from URI hostname
    if hostname and ADLS_ACCOUNT_NAME not in properties:
        properties[ADLS_ACCOUNT_NAME] = hostname.split(".")[0]

and the caller at fsspec.py:515:

if scheme in _ADLS_SCHEMES:
    return _adls(self.properties, hostname)

The lru_cache on (scheme, hostname) just above it is not the problem. It does correctly build a second filesystem for a second hostname. But by then ADLS_ACCOUNT_NAME is already present in the shared properties, so the inference is skipped and the second filesystem gets the first account.

Steps to reproduce

Two locations in two different storage accounts, starting from empty properties:

from unittest.mock import patch
from pyiceberg.io.fsspec import FsspecFileIO

loc_a = "abfss://data@accountone.dfs.core.windows.net/wh/t/a.parquet"
loc_b = "abfss://data@accounttwo.dfs.core.windows.net/wh/t/b.parquet"

captured = []
class FakeFS:
    def __init__(self, **kw): captured.append(kw.get("account_name"))

io = FsspecFileIO(properties={})
with patch("adlfs.AzureBlobFileSystem", FakeFS):
    io.new_input(loc_a)
    print(dict(io.properties))
    io.new_input(loc_b)
    print(dict(io.properties))

print(captured)

Output:

{'adls.account-name': 'accountone'}
{'adls.account-name': 'accountone'}
['accountone', 'accountone']

Expected ['accountone', 'accounttwo'].

What I expected

A FileIO given no adls.account-name should infer the account per location, not once. The inference is already per hostname at the cache layer, only the write into the shared dict breaks it.

The same mutation happens a few lines above in the SAS token loop, which sets ADLS_ACCOUNT_NAME from the token key, so that path has the same effect.

Impact

A catalog whose tables span two storage accounts silently reads from the wrong account. Depending on whether a same named container exists there, this either fails with a confusing not found or resolves to the wrong data. It also means a FileIO's properties change as a side effect of reading, which is surprising for anything that inspects or reuses them.

Suggested fix

Do not write into properties. Resolve the account into a local value inside _adls and pass that to AzureBlobFileSystem, leaving the caller's dict untouched. Happy to send a PR.

Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
Lingua principale
Python
Stelle
1.1k
Fork
589
Merge medio
2g 4h
PR unite (30g)
72

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di apache/iceberg-python

Tutte le issue di apache/iceberg-python

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.