Replace the logical codec's object registry with durable metadata
@s5dsn-eqee arbeitet bereits daran.
Seit 21.9.2026.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 74/100
Rechercherichtung
Beginnen Sie in examples/datafusion-ffi-example/src/logical_extension_codec.rs bei try_encode_table_provider und vergleichen Sie den IPC-Ansatz in examples/distributed/storage-library/src/codec.rs. Aktualisieren Sie python/tests/_test_logical_extension_codec.py, einschließlich der Assertion für den früheren Codec und eines Tests zum zweimaligen Decodieren. Als erledigt gilt die Aufgabe, wenn der Registry- und Token-Code entfernt sind, provider_prefix verwendet wird, die angegebenen Tests bestehen und grep -in token nichts zurückgibt.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
examples/datafusion-ffi-example/src/logical_extension_codec.rs parks live table providers in a process-global HashMap and encodes an integer token into it. Encoding inserts, decoding removes, so the same bytes cannot be decoded twice, one plan cannot fan out to several readers, and a plan that never reaches a decoder keeps its provider alive for the life of the process. extension-guide/codecs.md tells authors not to do this.
Unlike the physical codec in the same crate (see the quarantine sub-issue), this one is fixable: try_encode_table_provider at line 148 claims node.downcast_ref::<MemTable>(), which is narrow, and a MemTable is fully describable by its schema and batches.
Pattern to copy: examples/distributed/storage-library/src/codec.rs — same Arrow IPC technique, same error convention (internal_datafusion_err! on encode, since this process holds the object; exec_datafusion_err! on decode, since those are foreign bytes).
Verified prerequisites: MemTable.batches is pub (datafusion-catalog/src/memory/table.rs:69), typed Vec<PartitionData> where PartitionData = Arc<tokio::sync::RwLock<Vec<RecordBatch>>>. MemTable::try_new rejects zero partitions (table.rs:84). arrow is already a dependency with IPC available, so no Cargo.toml change.
Proposed wire format, keeping the per-instance prefix the dispatch tests rely on:
<provider_prefix> | b"MEMTBL1" | u32 LE n_partitions | { u32 LE ipc_len | ipc stream }*
One stream per partition, because MemTable partition boundaries become output partitions. A stream carries its schema even when empty, so an empty partition round-trips.
Two traps worth writing down before someone hits them:
- Use
try_read()on each partition lock, notblocking_read(). The FFI codec runs with a tokio runtime handle installed, andblocking_readpanics in that context. - On decode, build
MemTable::try_newfrom the IPC schema, not theschema: SchemaRefargument.try_newvalidatesschema.contains(&batch.schema()), so metadata drift would surface as a spurious mismatch. This is the opposite choice fromstorage-library/src/codec.rs:376-380, which must honour the plan's schema because it re-reads files from disk; here the batches are the payload. Worth a comment noting the contrast, since the two codecs otherwise look alike.
Done when: the registry, token_id(), and the HashMap/Mutex/OnceLock/AtomicU64 imports are gone; the struct field token is renamed provider_prefix to match the Python kwarg that already uses that name; and grep -in token over the file returns nothing.
Tests: of 19 tests in python/tests/_test_logical_extension_codec.py, one changes. test_installing_a_codec_cannot_hijack_an_earlier_codecs_objects asserts len(before) == len(after) with a comment about tokens being minted per encode; that comment becomes false and the assertion becomes weaker than reality, so it should become assert before == after. Add one test for the property the guide claims and nothing currently covers: encode once, decode twice on one session, assert both produce the same rows. All 47 tests in the planner crate should be unaffected — every assertion there is on call counters, never on payload shape.
- Vorherrschende Sprache
- Python
- Sterne
- 605
- Forks
- 176
- Ø Merge
- 1 T. 23 Std.
- Gemergte PRs (30 T.)
- 8
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus apache/datafusion-python
-
enhancement
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
apache/datafusion-python#1757 ·
-
documentation
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
apache/datafusion-python#1726 ·
-
Schwierigkeit 2/5 Ein halber Tag Anfängerfreundlichkeit 88/100
apache/datafusion-python#1691 ·
-
bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
apache/datafusion-python#1644 ·
-
enhancement
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 30/100
apache/datafusion-python#1737 ·
Alle Issues in apache/datafusion-python
Ähnliche Issues
-
area: harness bug status: needs-triage
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
Human-Agent-Society/reef#625 ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 80/100
learningequality/kolibri#15351 · 2 Kommentare ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
-
Name consistency Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
eellak/triplestore#65 · 1 Kommentar ·