DataFrame subclass lost in `groupby.agg` with `split_out` set.
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 42/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 停滞
调研方向
首先运行提供的 test.py 复现代码,并比较有无 split_out 时计算出的类型。跟踪 DecomposableGroupbyAggregation.combine 和 aggregate,重点关注最终 aggregate 从 MyDataFrame 变为 pandas.DataFrame 的位置。当在 split_out=2 时保留子类,并添加覆盖这两种情况的回归测试后,即视为完成。
由索引模型根据 Issue 内容生成。
描述
Describe the issue:
As part of https://github.com/geopandas/dask-geopandas/pull/285, we found that dask-expr will lose the type of a pandas DataFrame subclass in groupby.agg if (and only if?) the split_out parameter is used.
Minimal Complete Verifiable Example:
Given this file:
# file: test.py
import dask.dataframe.backends
import pandas as pd
import dask_expr as dx
import dask.dataframe as dd
from dask.dataframe.dispatch import make_meta_dispatch, meta_nonempty
from dask.dataframe.core import get_parallel_type
import dask.dataframe.backends
dask.config.set(scheduler="single-threaded")
class MySeries(pd.Series):
@property
def _constructor(self):
return MySeries
@property
def _constructor_expanddim(self):
return MyDataFrame
class MyDataFrame(pd.DataFrame):
@property
def _constructor(self):
return MyDataFrame
@property
def _constructor_sliced(self):
return MySeries
class MyIndex(pd.Index): ...
class MyDaskSeries(dx.Series):
_partition_type = MySeries
class MyDaskDataFrame(dx.DataFrame):
_partition_type = MyDataFrame
class MyDaskIndex(dx.Index):
_partition_type = MyIndex
# Unclear if any of get_parallel_type and make_meta_dispatch are needed.
# Reproduces with or without them.
@get_parallel_type.register(MyDataFrame)
def get_parallel_type_dataframe(df):
return MyDataFrame
@get_parallel_type.register(MySeries)
def get_parallel_type_series(s):
return MyDaskSeries
@get_parallel_type.register(MyIndex)
def get_parallel_type_index(ind):
return MyDaskIndex
@make_meta_dispatch.register(MyDataFrame)
def make_meta_dataframe(df, index=None):
return df.head(0)
@make_meta_dispatch.register(MySeries)
def make_meta_series(s, index=None):
return s.head(0)
@make_meta_dispatch.register(MyIndex)
def make_meta_index(ind, index=None):
return ind[:0]
@meta_nonempty.register(MyDataFrame)
def make_meta_nonempty_dataframe(x):
return MyDataFrame(dask.dataframe.backends.meta_nonempty_dataframe(x))
df = dx.from_dict(
{"a": [1, 1, 2, 2], "b": [1, 2, 3, 4]}, npartitions=4, constructor=MyDataFrame
)
a = df.groupby("a").agg("first")
b = df.groupby("a").agg("first", split_out=2)
print("split-out=None", type(a.compute()))
print("split-out=2 ", type(b.compute()))
running that produces
$ python test.py
split-out=None <class '__main__.MyDataFrame'>
split-out=2 <class 'pandas.core.frame.DataFrame'>
I would expect the type there to be __main__.MyDataFrame regardless of split_out.
Anything else we need to know?:
Environment:
dask 2024.4.1
dask-expr 1.0.11
Edit: I made one addition to the script: adding a @meta_nonempty.register(MyDataFrame). I noticed that in DecomposableGroupbyAggregation.combine and DecomposableGroupbyAggregation.aggregate the types were regular pandas DataFrames, instead of the subclass.
Registering that meta_nonempty does keep it as MyDataFrame initially. I put some print statements in those methods to print the type of inputs[0] and type(_concat(inputs)) and get
combine <class '__main__.MyDataFrame'> <class '__main__.MyDataFrame'>
aggregate <class '__main__.MyDataFrame'> <class '__main__.MyDataFrame'>
aggregate <class '__main__.MyDataFrame'> <class '__main__.MyDataFrame'>
aggregate <class 'pandas.core.frame.DataFrame'> <class 'pandas.core.frame.DataFrame'>
aggregate <class 'pandas.core.frame.DataFrame'> <class 'pandas.core.frame.DataFrame'>
split-out=2 <class 'pandas.core.frame.DataFrame'>
So initially we're OK, but by the time we do the final aggregate we've lost the subclass.
- 主要语言
- Python
- 星标
- 89
- 派生
- 26
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
dask/dask-expr 的其他 Issue
-
难度 4/5 3-5 天 新手友好度 38/100
-
难度 3/5 1-2 天 新手友好度 45/100
-
难度 5/5 一周以上 新手友好度 30/100
-
难度 4/5 3-5 天 新手友好度 30/100
-
难度 3/5 1-2 天 新手友好度 25/100
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 78/100
conda-forge/conda-build-feedstock#289 · 1 条评论 · 1 个 reaction ·
-
`pulptest` no longer works in 4.0.0: `ImportError: Start directory is not importable: 'pulp/tests'`未关闭
难度 2/5 1-3 小时 新手友好度 75/100
-
难度 2/5 1-3 小时 新手友好度 72/100
-
难度 2/5 1-3 小时 新手友好度 78/100
-
难度 2/5 1-3 小时 新手友好度 72/100
TomCasavant/ohio-sites#224 ·