FsspecFileIO: `_adls` mutates shared properties, so a second storage account gets the first account's filesystem

Open Beginner friendly
#3,885 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
backend, cloud

Research direction

Start in pyiceberg/io/fsspec.py at _adls and its caller around line 515, then run the two-location reproduction with the mocked AzureBlobFileSystem. Verify that each hostname produces its own account and that io.properties remains unchanged after both reads.

Written by the indexing model from the issue text.

Description

Apache Iceberg version

main (development)

Please describe the bug 🐞

_adls writes the account name it infers back into the properties dict it is handed, and the call site passes the FileIO's own self.properties. So the first ADLS location a FsspecFileIO touches pins adls.account-name for the life of that FileIO, and every later location is served a filesystem built for the first account.

In pyiceberg/io/fsspec.py:

def _adls(properties: Properties, hostname: str | None = None) -> AbstractFileSystem:
    ...
    # Fallback: extract account_name from URI hostname
    if hostname and ADLS_ACCOUNT_NAME not in properties:
        properties[ADLS_ACCOUNT_NAME] = hostname.split(".")[0]

and the caller at fsspec.py:515:

if scheme in _ADLS_SCHEMES:
    return _adls(self.properties, hostname)

The lru_cache on (scheme, hostname) just above it is not the problem. It does correctly build a second filesystem for a second hostname. But by then ADLS_ACCOUNT_NAME is already present in the shared properties, so the inference is skipped and the second filesystem gets the first account.

Steps to reproduce

Two locations in two different storage accounts, starting from empty properties:

from unittest.mock import patch
from pyiceberg.io.fsspec import FsspecFileIO

loc_a = "abfss://data@accountone.dfs.core.windows.net/wh/t/a.parquet"
loc_b = "abfss://data@accounttwo.dfs.core.windows.net/wh/t/b.parquet"

captured = []
class FakeFS:
    def __init__(self, **kw): captured.append(kw.get("account_name"))

io = FsspecFileIO(properties={})
with patch("adlfs.AzureBlobFileSystem", FakeFS):
    io.new_input(loc_a)
    print(dict(io.properties))
    io.new_input(loc_b)
    print(dict(io.properties))

print(captured)

Output:

{'adls.account-name': 'accountone'}
{'adls.account-name': 'accountone'}
['accountone', 'accountone']

Expected ['accountone', 'accounttwo'].

What I expected

A FileIO given no adls.account-name should infer the account per location, not once. The inference is already per hostname at the cache layer, only the write into the shared dict breaks it.

The same mutation happens a few lines above in the SAS token loop, which sets ADLS_ACCOUNT_NAME from the token key, so that path has the same effect.

Impact

A catalog whose tables span two storage accounts silently reads from the wrong account. Depending on whether a same named container exists there, this either fails with a confusing not found or resolves to the wrong data. It also means a FileIO's properties change as a side effect of reading, which is surprising for anything that inspects or reuses them.

Suggested fix

Do not write into properties. Resolve the account into a local value inside _adls and pass that to AzureBlobFileSystem, leaving the caller's dict untouched. Happy to send a PR.

Willingness to contribute
  • I can contribute a fix for this bug independently
  • I would be willing to contribute a fix for this bug with guidance from the Iceberg community
  • I cannot contribute a fix for this bug at this time
Dominant language
Python
Stars
1.1k
Forks
589
Avg merge
2d 4h
Merged PRs (30d)
72

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/iceberg-python

All issues in apache/iceberg-python

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.