[Feature Request] Add pyiceberg.catalog.hadoop.HadoopCatalog (filesystem-only catalog)
Maintainer antworten meist innerhalb von 1 Tag
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Anfängerfreundlichkeit
- 42/100
- Issue-Typ
- Feature
- Klarheit
- Größtenteils klar
- Aktivitätsstatus
- Aktiv
- Bereich
- data-engineering, databases
Rechercherichtung
Beginne damit, die vorhandenen Katalogimplementierungen und MetastoreCatalog zu lesen, und untersuche anschließend PyArrowFileIO._initialize_fs sowie die im Issue genannte nachgelagerte Verwendung von daft/catalog/__gravitino/__catalog.py. Vergleiche das angeforderte Verhalten mit den Referenzen HadoopCatalog und HadoopTables von Java Iceberg. Erledigt ist die Aufgabe, wenn Namespace- und Tabellenoperationen ausschließlich über das Dateisystem möglich sind, die Metadatenauflösung mit Java kompatibel ist und ein dokumentierter oder öffentlicher Pfad für benutzerdefinierte Dateisystemschemata vorhanden ist.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Is your feature request related to a problem? Please describe.
Java Iceberg ships a filesystem-only HadoopCatalog / HadoopTables, where table metadata lives under <warehouse>/<db>.db/<table>/metadata/ with no external metastore. PyIceberg currently has no equivalent — the available catalog types are rest / hive / glue / dynamodb / sql / in-memory / bigquery.
This gap matters in two ways:
-
Interop with Java-side HadoopCatalog tables. Tables created by Java
HadoopCatalog(common in lightweight deployments without a metastore) cannot be opened through any supported PyIceberg catalog. Users must fall back toStaticTable.from_metadataand resolve the latestmetadata.jsonthemselves, which loses catalog semantics (no namespace listing, no create/commit). -
Downstream projects already assume the module exists. Daft's Gravitino integration imports
from pyiceberg.catalog.hadoop import HadoopCatalogand callsHadoopCatalog("gravitino_reader", props).load_table(table_dir)to open a table from a storage location (daft/catalog/__gravitino/_catalog.py). Against PyIceberg 0.11.x this raisesTypeError: HadoopCatalog.__init__() takes 2 positional arguments but 3 were given, and against versions without the module it fails at import time.
Describe the solution you'd like
A pyiceberg.catalog.hadoop.HadoopCatalog (subclassing MetastoreCatalog) implementing Java HadoopCatalog semantics:
warehouseproperty as the root location- table dir =
<warehouse>/<namespace>/<table> - metadata at
<table_dir>/metadata/v{n}.metadata.jsonplus aversion-hint.textholding the current version - latest-version resolution: read
version-hint.text, fall back to scanningmetadata/for the maxv{n}(matching Java behavior) - namespace/table create/list/commit driven purely by the warehouse filesystem (no metastore calls)
Additional context / pitfalls observed while prototyping
Happy to contribute a PR if this is in scope. A few notes from an internal prototype:
-
Metadata file naming. Java HadoopCatalog uses
v{n}.metadata.json+version-hint.text, but tables created by JDBC/REST catalogs use00000-<uuid>.metadata.jsonwith no version-hint. To open those as well, the scan fallback should accept both patterns (v(\d+)\.metadata\.jsonand\d{5}-.*\.metadata\.json), or at least document the limitation. -
Filesystem abstraction.
__init__should derive the filesystem from the catalog's FileIO (PyArrowFileIO) instead of hardcodingpyarrow.fs.HadoopFileSystem.from_uri(warehouse). The JVM-backedHadoopFileSystemonly supportshdfs://and fails for object-store schemes (s3://, and custom schemes), so routing through FileIO keeps it scheme-agnostic. -
Atomicity.
create_table/commit_tableuse create-if-absent onv{n}.metadata.jsonfor optimistic concurrency — safe on HDFS but not atomic on plain object stores (S3 has no create-if-absent guarantee). Java has the same caveat; worth documenting or using a conditional-write primitive where available.
Adapting a custom storage scheme (Tencent Cloud TBDSFS as a concrete case)
A related gap surfaced while prototyping against Tencent Cloud TBDS's distributed filesystem scheme tbdsfs://<cluster>/<path> (exposed by a Python client, plus a JVM fs.tbdsfs.impl):
-
PyArrowFileIO._initialize_fs(scheme, netloc)only understands a fixed set of schemes (hdfs / s3 / gs / file / abfs / ...), so anytbdsfs://...location raisesValueError: Unrecognized filesystem type in URI: tbdsfs. -
There is no public, documented way to plug in a custom filesystem. The only workaround today is monkey-patching a private method:
- Implement a
pyarrow.fs.FileSystemHandlersubclass wrapping the TBDSFS Python client, wrap it inpyarrow.fs.PyFileSystem, then patchPyArrowFileIO._initialize_fsto return that filesystem whenscheme == "tbdsfs". - pyarrow 21's
PyFileSystemcallback also has non-obvious contracts any custom handler must satisfy: single paths arrive as one-element lists;get_file_infomust return a one-element list; and for scheme'd URIs the netloc is prepended into the path (e.g.internal/usr/..., without a leading slash).
- Implement a
This works, but it depends on patching a private API (_initialize_fs), which is brittle across PyIceberg releases.
Suggested improvement: a documented, public extension point for registering an arbitrary pyarrow.fs.FileSystem (or a custom FileIO) per scheme — e.g. a register_file_system(scheme, factory) helper, or a scheme → FileSystem mapping read from FileIO/catalog properties — so non-standard object stores and filesystems can be integrated without touching internals. This would also naturally address the HadoopCatalog.__init__ filesystem-abstraction point above.
References
- Java:
org.apache.iceberg.hadoop.HadoopCatalog/HadoopTables - Downstream usage that currently breaks:
daft/catalog/__gravitino/_catalog.py→_open_iceberg_table
- Vorherrschende Sprache
- Python
- Sterne
- 1.1k
- Forks
- 606
- Ø Merge
- 1 T. 11 Std.
- Gemergte PRs (30 T.)
- 76
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Kein Beitragsleitfaden
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus apache/iceberg-python
-
View does not expose metadata_location: RestCatalog.load_view discards it from the server's responseEvtl. vergeben @Soumo-git-hub hat das vor 1 Tag übernommen. Offenkind:bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
apache/iceberg-python#4073 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
apache/iceberg-python#4010 · 3 Kommentare · 1 Reaktion ·
Maintainer antworten meist innerhalb von 1 Tag
-
to_bytes silently rescales a Decimal with a negative scaleEvtl. vergeben @Rodrigo-Palma hat das vor 17 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
apache/iceberg-python#3996 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Deletion vector bitmap count is read from the blob and used as a loop bound without validationEvtl. vergeben @ghoshp83 hat das vor 17 Tagen übernommen. Offenbug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
apache/iceberg-python#3979 ·
Maintainer antworten meist innerhalb von 1 Tag
-
FsspecFileIO: `_adls` mutates shared properties, so a second storage account gets the first account's filesystemEvtl. vergeben @krishnakaanchan-png hat das vor 34 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
apache/iceberg-python#3885 ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in apache/iceberg-python
Ähnliche Issues
-
feature:LinkChecker
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 66/100
digitalfabrik/integreat-cms#4594 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
EleutherAI/lm-evaluation-harness#4319 ·
Maintainer antworten meist innerhalb von 1 Tag
-
needs triage
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 76/100
Maintainer antworten meist innerhalb von 1 Tag
-
json_params_matcher fails on falsy top-level JSON primitives (0, False, "")Evtl. vergeben @mayureshsonawane17 hat das heute übernommen. OffenWaiting for: Product Owner
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
Maintainer antworten meist innerhalb von 5 Tagen