[Feature Request] Add pyiceberg.catalog.hadoop.HadoopCatalog (filesystem-only catalog)
Les mainteneurs répondent en général sous 1 jour
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Accessibilité débutants
- 42/100
- Type d'issue
- Fonctionnalité
- Clarté
- Plutôt claire
- Activité
- Active
- Domaine
- data-engineering, databases
Piste de recherche
Commencez par lire les implémentations de catalogue existantes et MetastoreCatalog, puis examinez PyArrowFileIO._initialize_fs et l’utilisation en aval de daft/catalog/__gravitino/__catalog.py mentionnée dans l’issue. Comparez le comportement demandé avec les références HadoopCatalog et HadoopTables de Java Iceberg. Le travail est terminé lorsque les opérations sur les namespaces et les tables s’effectuent uniquement via le système de fichiers, que la résolution des métadonnées est compatible avec Java et qu’un chemin documenté ou public existe pour les schémas de systèmes de fichiers personnalisés.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Is your feature request related to a problem? Please describe.
Java Iceberg ships a filesystem-only HadoopCatalog / HadoopTables, where table metadata lives under <warehouse>/<db>.db/<table>/metadata/ with no external metastore. PyIceberg currently has no equivalent — the available catalog types are rest / hive / glue / dynamodb / sql / in-memory / bigquery.
This gap matters in two ways:
-
Interop with Java-side HadoopCatalog tables. Tables created by Java
HadoopCatalog(common in lightweight deployments without a metastore) cannot be opened through any supported PyIceberg catalog. Users must fall back toStaticTable.from_metadataand resolve the latestmetadata.jsonthemselves, which loses catalog semantics (no namespace listing, no create/commit). -
Downstream projects already assume the module exists. Daft's Gravitino integration imports
from pyiceberg.catalog.hadoop import HadoopCatalogand callsHadoopCatalog("gravitino_reader", props).load_table(table_dir)to open a table from a storage location (daft/catalog/__gravitino/_catalog.py). Against PyIceberg 0.11.x this raisesTypeError: HadoopCatalog.__init__() takes 2 positional arguments but 3 were given, and against versions without the module it fails at import time.
Describe the solution you'd like
A pyiceberg.catalog.hadoop.HadoopCatalog (subclassing MetastoreCatalog) implementing Java HadoopCatalog semantics:
warehouseproperty as the root location- table dir =
<warehouse>/<namespace>/<table> - metadata at
<table_dir>/metadata/v{n}.metadata.jsonplus aversion-hint.textholding the current version - latest-version resolution: read
version-hint.text, fall back to scanningmetadata/for the maxv{n}(matching Java behavior) - namespace/table create/list/commit driven purely by the warehouse filesystem (no metastore calls)
Additional context / pitfalls observed while prototyping
Happy to contribute a PR if this is in scope. A few notes from an internal prototype:
-
Metadata file naming. Java HadoopCatalog uses
v{n}.metadata.json+version-hint.text, but tables created by JDBC/REST catalogs use00000-<uuid>.metadata.jsonwith no version-hint. To open those as well, the scan fallback should accept both patterns (v(\d+)\.metadata\.jsonand\d{5}-.*\.metadata\.json), or at least document the limitation. -
Filesystem abstraction.
__init__should derive the filesystem from the catalog's FileIO (PyArrowFileIO) instead of hardcodingpyarrow.fs.HadoopFileSystem.from_uri(warehouse). The JVM-backedHadoopFileSystemonly supportshdfs://and fails for object-store schemes (s3://, and custom schemes), so routing through FileIO keeps it scheme-agnostic. -
Atomicity.
create_table/commit_tableuse create-if-absent onv{n}.metadata.jsonfor optimistic concurrency — safe on HDFS but not atomic on plain object stores (S3 has no create-if-absent guarantee). Java has the same caveat; worth documenting or using a conditional-write primitive where available.
Adapting a custom storage scheme (Tencent Cloud TBDSFS as a concrete case)
A related gap surfaced while prototyping against Tencent Cloud TBDS's distributed filesystem scheme tbdsfs://<cluster>/<path> (exposed by a Python client, plus a JVM fs.tbdsfs.impl):
-
PyArrowFileIO._initialize_fs(scheme, netloc)only understands a fixed set of schemes (hdfs / s3 / gs / file / abfs / ...), so anytbdsfs://...location raisesValueError: Unrecognized filesystem type in URI: tbdsfs. -
There is no public, documented way to plug in a custom filesystem. The only workaround today is monkey-patching a private method:
- Implement a
pyarrow.fs.FileSystemHandlersubclass wrapping the TBDSFS Python client, wrap it inpyarrow.fs.PyFileSystem, then patchPyArrowFileIO._initialize_fsto return that filesystem whenscheme == "tbdsfs". - pyarrow 21's
PyFileSystemcallback also has non-obvious contracts any custom handler must satisfy: single paths arrive as one-element lists;get_file_infomust return a one-element list; and for scheme'd URIs the netloc is prepended into the path (e.g.internal/usr/..., without a leading slash).
- Implement a
This works, but it depends on patching a private API (_initialize_fs), which is brittle across PyIceberg releases.
Suggested improvement: a documented, public extension point for registering an arbitrary pyarrow.fs.FileSystem (or a custom FileIO) per scheme — e.g. a register_file_system(scheme, factory) helper, or a scheme → FileSystem mapping read from FileIO/catalog properties — so non-standard object stores and filesystems can be integrated without touching internals. This would also naturally address the HadoopCatalog.__init__ filesystem-abstraction point above.
References
- Java:
org.apache.iceberg.hadoop.HadoopCatalog/HadoopTables - Downstream usage that currently breaks:
daft/catalog/__gravitino/_catalog.py→_open_iceberg_table
- Langage dominant
- Python
- Étoiles
- 1.1k
- Forks
- 606
- Merge moyen
- 1 j 18 h
- PR mergées (30 j)
- 87
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Propose un modèle de pull request
- Aucun guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de apache/iceberg-python
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's responseOuvertekind:bug
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100
apache/iceberg-python#4073 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
apache/iceberg-python#4010 · 3 commentaires · 1 réaction ·
Les mainteneurs répondent en général sous 1 jour
-
to_bytes silently rescales a Decimal with a negative scalePeut-être pris @Rodrigo-Palma l’a pris il y a 16 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3996 ·
Les mainteneurs répondent en général sous 1 jour
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationPeut-être pris @ghoshp83 l’a pris il y a 17 jours. Ouvertebug
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
apache/iceberg-python#3979 ·
Les mainteneurs répondent en général sous 1 jour
-
FsspecFileIO: `_adls` mutates shared properties, so a second storage account gets the first account's filesystemPeut-être pris @krishnakaanchan-png l’a pris il y a 34 jours. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
apache/iceberg-python#3885 ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de apache/iceberg-python
Issues similaires
-
Action calls retired claude-3-5-haiku-20241022, generating failing API requests for every userOuverte
Difficulté 1/5 Moins d'une heure Accessibilité débutants 91/100
-
Difficulté 1/5 1-3 heures Accessibilité débutants 92/100
-
enhancement P2
Difficulté 2/5 1-3 heures Accessibilité débutants 78/100
Toloka/tolokaforge#1776 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 88/100
TencentCloud/Octop#1622 ·
Les mainteneurs répondent en général sous 1 jour
-
[arch] gate_sync.sh: stale remote tmp files on install failure; overlapping cron runs unguardedOuvertearch area:fleet priority:p3 severity:low track:hosted-product
Difficulté 2/5 1-3 heures Accessibilité débutants 76/100
Les mainteneurs répondent en général sous 1 jour