Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

DataFrame subclass lost in `groupby.agg` with `split_out` set.

未关闭
#1,024 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
42/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
停滞
技术栈
pandas, python

调研方向

首先运行提供的 test.py 复现代码,并比较有无 split_out 时计算出的类型。跟踪 DecomposableGroupbyAggregation.combine 和 aggregate,重点关注最终 aggregate 从 MyDataFrame 变为 pandas.DataFrame 的位置。当在 split_out=2 时保留子类,并添加覆盖这两种情况的回归测试后,即视为完成。

由索引模型根据 Issue 内容生成。

描述

Describe the issue:

As part of https://github.com/geopandas/dask-geopandas/pull/285, we found that dask-expr will lose the type of a pandas DataFrame subclass in groupby.agg if (and only if?) the split_out parameter is used.

Minimal Complete Verifiable Example:

Given this file:

# file: test.py
import dask.dataframe.backends
import pandas as pd

import dask_expr as dx
import dask.dataframe as dd
from dask.dataframe.dispatch import make_meta_dispatch, meta_nonempty
from dask.dataframe.core import get_parallel_type
import dask.dataframe.backends


dask.config.set(scheduler="single-threaded")

class MySeries(pd.Series):
    @property
    def _constructor(self):
        return MySeries

    @property
    def _constructor_expanddim(self):
        return MyDataFrame


class MyDataFrame(pd.DataFrame):
    @property
    def _constructor(self):
        return MyDataFrame

    @property
    def _constructor_sliced(self):
        return MySeries


class MyIndex(pd.Index): ...


class MyDaskSeries(dx.Series):
    _partition_type = MySeries


class MyDaskDataFrame(dx.DataFrame):
    _partition_type = MyDataFrame


class MyDaskIndex(dx.Index):
    _partition_type = MyIndex


# Unclear if any of get_parallel_type and make_meta_dispatch are needed.
# Reproduces with or without them.
@get_parallel_type.register(MyDataFrame)
def get_parallel_type_dataframe(df):
    return MyDataFrame


@get_parallel_type.register(MySeries)
def get_parallel_type_series(s):
    return MyDaskSeries


@get_parallel_type.register(MyIndex)
def get_parallel_type_index(ind):
    return MyDaskIndex


@make_meta_dispatch.register(MyDataFrame)
def make_meta_dataframe(df, index=None):
    return df.head(0)


@make_meta_dispatch.register(MySeries)
def make_meta_series(s, index=None):
    return s.head(0)


@make_meta_dispatch.register(MyIndex)
def make_meta_index(ind, index=None):
    return ind[:0]


@meta_nonempty.register(MyDataFrame)
def make_meta_nonempty_dataframe(x):
    return MyDataFrame(dask.dataframe.backends.meta_nonempty_dataframe(x))


df = dx.from_dict(
    {"a": [1, 1, 2, 2], "b": [1, 2, 3, 4]}, npartitions=4, constructor=MyDataFrame
)
a = df.groupby("a").agg("first")
b = df.groupby("a").agg("first", split_out=2)

print("split-out=None", type(a.compute()))
print("split-out=2   ", type(b.compute()))

running that produces

$ python test.py
split-out=None <class '__main__.MyDataFrame'>
split-out=2    <class 'pandas.core.frame.DataFrame'>

I would expect the type there to be __main__.MyDataFrame regardless of split_out.

Anything else we need to know?:

Environment:

dask               2024.4.1
dask-expr          1.0.11

Edit: I made one addition to the script: adding a @meta_nonempty.register(MyDataFrame). I noticed that in DecomposableGroupbyAggregation.combine and DecomposableGroupbyAggregation.aggregate the types were regular pandas DataFrames, instead of the subclass.

Registering that meta_nonempty does keep it as MyDataFrame initially. I put some print statements in those methods to print the type of inputs[0] and type(_concat(inputs)) and get

combine <class '__main__.MyDataFrame'> <class '__main__.MyDataFrame'>
aggregate <class '__main__.MyDataFrame'> <class '__main__.MyDataFrame'>
aggregate <class '__main__.MyDataFrame'> <class '__main__.MyDataFrame'>
aggregate <class 'pandas.core.frame.DataFrame'> <class 'pandas.core.frame.DataFrame'>
aggregate <class 'pandas.core.frame.DataFrame'> <class 'pandas.core.frame.DataFrame'>
split-out=2    <class 'pandas.core.frame.DataFrame'>

So initially we're OK, but by the time we do the final aggregate we've lost the subclass.

主要语言
Python
星标
89
派生
26
PR 合并指标
30 天内没有已合并 PR

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

dask/dask-expr 的其他 Issue

查看 dask/dask-expr 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。