PyArrowFile class is not compatible with ABFS uri syntax
Les mainteneurs répondent en général sous 1 jour
@krishnakaanchan-png y travaille déjà.
Depuis le 11/9/2026.
- #3936 par @krishnakaanchan-png — ouverte
Évaluation
- Difficulté
- 3/5
- Temps estimé
- 1-2 jours
- Accessibilité débutants
- 62/100
Piste de recherche
Commencez dans pyiceberg/io/pyarrow.py, au niveau de PyArrowFile.init, create et exists(), et suivez la manière dont l’emplacement ABFSS est transmis à PyArrow. Reproduisez l’échec avec une URI ABFS contenant le nom du compte, puis vérifiez que les opérations de lecture et d’écriture de PyArrow fonctionnent sans le segment d’URI non valide.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Apache Iceberg version
0.10.0 (latest release)
Please describe the bug 🐞
Starting from version 20, Pyarrow has support for Azure filesystems.
ABFS URIs have this format: abfs[s]://<file_system>@<account_name>.dfs.core.windows.net//<file_name>
But Pyarrow library expects the following path format for Azure: abfs[s]://<file_system>//<file_name>.
As you see, the part "@<account_name>.<dfs|blob>.core.windows.net" prevents users to use pyarrow file io in Azure environment. This issue CAN be fixed in Pyiceberg by removing account_name part.
The proposed fix is just to start a conversation around the issue. I am not 100% sure how and where this should be fixed.
We know similar issues do not occur with Fsspec file io.
Examples
We have a very basic setup with RestCatalog:
def create_iceberg_catalog():
CATALOG_URI = "https://lakehouse.../catalog"
catalog_config = {
"uri": CATALOG_URI,
PY_IO_IMPL: "pyiceberg.io.pyarrow.PyArrowFileIO",
ADLS_ACCOUNT_NAME: "lakehouseaccount",
}
return RestCatalog("lakehouse", **catalog_config)
When we create a table "testns.testtable", it is assigned a following location : abfss://[email protected]/testns/testtable
Then, when we try to append data to the table:
data = pa.table(
{
"id": pa.array(range(5), type=pa.int32()), # Ensure 'id' is int32 to match Iceberg schema
"value": [random.choice(["Heads", "Tails"]) for _ in range(5)],
}
)
table.append(data)
it throws the following exception:
OSError: ListBlobsByHierarchy failed for prefix='aip_test[/test_table-xxx/metadata/snap-xxx.avro](https://xxx/test_table-xxx.avro)'. GetFileInfo is unable to determine whether the path exists. Azure Error: [InvalidResourceName] 400 The specified resource name contains invalid characters.
This is because exists() method is called:
File [~/.official-venvs/amd64.ipykernel-default.master/lib/python3.12/site-packages/pyiceberg/io/pyarrow.py:368](https://xxx/user/nikita-matckevich/.official-venvs/amd64.ipykernel-default.master/lib/python3.12/site-packages/pyiceberg/io/pyarrow.py#line=367), in PyArrowFile.create(self, overwrite)
366 if not overwrite and self.exists() is True:
And it expects the uri without "@akehouseaccount.dfs.core.windows.net". When we monkey-patch the PyArrowFile.init everything works fine:
PyArrowFile.old_init = PyArrowFile.__init__
def patched_init(self, location: str, path: str, fs: FileSystem, buffer_size: int = ONE_MEGABYTE):
# Call the original __init__ method
self.old_init(location, path, fs, buffer_size)
self._path = remove_section_between_at_and_slash(path)
print("Logging: PyArrowFile initialized")
PyArrowFile.__init__ = patched_init
It does not matter how and with which engine the table was created and written before: all pyarrow methods are not working, even those that are on read path, so it will be impossible to scan a non-empty table as well. We tested it by creating a table with fsspec file io and reading it with pyarrow file io.
It is hard to test this behavior with Azurite, because Azurite uris are different and do not contain "@<account_name>" part.
Willingness to contribute
- I can contribute a fix for this bug independently
- I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- I cannot contribute a fix for this bug at this time
- Langage dominant
- Python
- Étoiles
- 1.2k
- Forks
- 618
- Merge moyen
- 1 j 18 h
- PR mergées (30 j)
- 70
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Propose un modèle de pull request
- Aucun guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de apache/iceberg-python
-
PyArrowFileIO: every small S3 write is a 3-request multipart upload; expose allow_delayed_openOuverte
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
apache/iceberg-python#4093 ·
Les mainteneurs répondent en général sous 1 jour
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's responsePeut-être pris @Soumo-git-hub l’a pris il y a 3 jours. Ouvertekind:bug
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100
apache/iceberg-python#4073 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
apache/iceberg-python#4010 · 3 commentaires · 1 réaction ·
Les mainteneurs répondent en général sous 1 jour
-
to_bytes silently rescales a Decimal with a negative scalePeut-être pris @Rodrigo-Palma l’a pris il y a 23 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3996 ·
Les mainteneurs répondent en général sous 1 jour
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationPeut-être pris @ghoshp83 l’a pris il y a 23 jours. Ouvertebug
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
apache/iceberg-python#3979 ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de apache/iceberg-python
Issues similaires
-
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
NousResearch/hermes-agent#136483 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 1/5 Moins d'une heure Accessibilité débutants 88/100
Les mainteneurs répondent en général sous 1 jour
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last elementPeut-être pris @peterdsharpe l’a pris aujourd’hui. Ouvertebug
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
pytorch/tensordict#2307 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
Les mainteneurs répondent en général sous 1 jour
-
GrokModel.generate/a_generate pass an OpenAI-style list-of-dicts to xai_sdk.chat.user(), so every call crashes with a protobuf TypeError before any network I/OPeut-être pris @Christian-Sidak l’a pris aujourd’hui. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
confident-ai/deepeval#3436 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour