Report which extension codec handled each node of a serialized plan
還沒有人認領這個 Issue。
評估
- 難度
- 5/5
- 預估耗時
- 一週以上
- 新手友好度
- 35/100
- Issue 類型
- 功能
- 描述清晰度
- 基本清楚
- 活躍度
- 活躍
- 技術堆疊
- python
研究方向
從 SessionContext.logical_extension_codec_ids() 和 physical_extension_codec_ids() 開始,然後閱讀 #1678 中描述的擴充 codec 鏈。判定序列化 plan 如何為 EXPLAIN 表示,並定義如何將勝出的 codec 歸屬於每個節點;完成的標準是,針對特定 plan 的檢查或 EXPLAIN 結果能夠識別處理每個序列化節點的 codec。
由索引模型根據 Issue 內容生成。
描述
Is your feature request related to a problem or challenge? Please describe what you are trying to do.
Extension codecs compose as of #1678: a session holds a chain of them, and encoding walks the chain in install order until one claims an object. With several libraries installed there is currently no way to ask which codec handled a given node. Raised in https://github.com/apache/datafusion-python/pull/1678#pullrequestreview-4940081598.
Part of this is answered already. SessionContext.logical_extension_codec_ids() and physical_extension_codec_ids() list what is installed, in install order, and those ids are what a payload carries — so a decode failure names the codec that wrote the bytes and lists what the session actually has. What is missing is per-call attribution: which codec handled which node of a particular plan. Today the only way to find out is to check a codec's own call counters before and after, which requires the codec to expose them and tells you nothing about which node was involved.
Describe the solution you'd like
Surface the winning codec per node, most naturally in EXPLAIN output for a plan that has been serialized, or failing that as an inspection call that reports the encode decisions made for a given plan.
Describe alternatives you've considered
Leaving it to the decode error, which already names the responsible codec. That covers the case where something went wrong but not the case where someone is trying to understand a working setup, which is when a multi-library chain is most confusing.
Logging each claim at debug level. Cheap to add and much weaker: it is per-session rather than per-plan, and it puts the burden of correlating lines with nodes on the reader.
Additional context
Deliberately left out of #1678: recording the winner means threading it through plan formatting, which is a change to how plans are displayed rather than to how codecs compose.
- 主要語言
- Python
- 星號
- 605
- 分支
- 176
- 平均合併
- 1 天 23 小時
- 30 天內合併 PR
- 8
貢獻指南
這個儲存庫沒有索引到貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
apache/datafusion-python 的其他 Issue
-
enhancement
難度 2/5 1-3 小時 新手友好度 70/100
apache/datafusion-python#1757 ·
-
documentation
難度 2/5 1-3 小時 新手友好度 72/100
apache/datafusion-python#1726 ·
-
難度 2/5 半天 新手友好度 88/100
apache/datafusion-python#1691 ·
-
bug
難度 2/5 1-3 小時 新手友好度 78/100
apache/datafusion-python#1644 ·
-
enhancement
難度 5/5 一週以上 新手友好度 30/100
apache/datafusion-python#1737 ·
查看 apache/datafusion-python 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 75/100
anthropics/skills#1811 · 1 則留言 ·
-
難度 2/5 1-3 小時 新手友好度 75/100
speaches-ai/speaches#678 ·
-
bug
難度 2/5 1-3 小時 新手友好度 75/100
datalayer/mcp-compose#42 ·
-
難度 2/5 1-3 小時 新手友好度 75/100
conda-forge/spacy-feedstock#177 ·
-
難度 2/5 1-3 小時 新手友好度 70/100
UKGovernmentBEIS/inspect_evals#2523 ·