rapidsai/cudf

[FEA] Support more dtypes in JIT GroupBy `apply`

开放

#12,608 创建于 2023年1月25日

 (6 条评论) (0 个反应) (0 位负责人)C++ (735 个派生)batch import
Pythonfeature requestgood first issuenumba

仓库指标

星标
 (6,000 个星标)
PR 合并指标
 (平均合并 17天 21小时) (30 天内合并 230 个 PR)

描述

Is your feature request related to a problem? Please describe. When https://github.com/rapidsai/cudf/pull/11452 lands, we'll get JIT Groupby.apply for a subset of UDFs and importantly, dtypes. However over the summer we only got as far as writing overloads for float64 and int64 dtypes in the users source data. It'd be nice if we could support more dtypes, starting at least with the rest of the numeric types.

Describe the solution you'd like Extend the existing groupby.apply, engine='jit' framework to support the following additional dtypes:

  • float32
  • int32
  • int16
  • int8
  • uint64
  • uint32
  • uint16
  • uint8
  • bool

A lot of the machinery in the original PR is fairly general and should make adding many of these easy- however there will undoubtedly be edge cases. As such it makes for a pretty good first issue for anyone jumping into the numba extension piece of the codebase.

贡献者指南