Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

How to use mutli-column functions in Patsy which return categorical output?

未关闭
#79 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
25/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
停滞
技术栈
numpy, pandas, python
领域
data

调研方向

从 issue 中展示的 patsy.dmatrix 调用和自定义 multival 函数开始,然后查看 statsmodels/statsmodels#2843 中链接的重复 issue。当自定义多列分类输出能够显示有意义的基于级别的列名,而不是数字索引时,工作即告完成。

由索引模型根据 Issue 内容生成。

描述

I am dealing with a situation where each item can have 1 or 2 labels both of which come from the same set. In this case I only want to have 1 set of categorical variables which encode if the values of each level are present or absent with respect to the reference. I also want to inversely weight the values of these categorical variables based on how many labels are present.
E.g. Let my categories be {A, B, C}

data = pd.DataFrame({"col1": ["A", "A", "B"], "col2": ["O", "B", "C"], "num_vals": [1,2,2]})

def inverse_val(x):
    return 1.0/x

# Using categorical coding will not give the correct output
X = patsy.dmatrix("(C(col1) + C(col2)):inverse_val(num_vals)", data, return_type="dataframe")

I tried to solve this issue using the following code:

def multival(*x, **kwargs):
  #raise Exception("Not Implemented")
  levels, reference = kwargs.get("levels", None), kwargs.get("reference", None)
  weights = kwargs.get("weights", None)
  if len(x[0].shape) != 1:
    raise Exception("Mismatching Shapes. All arrays should be 1d and should have the same shape")
  for k in x:
    if k.shape != x[0].shape:
      raise Exception("Mismatching Shapes. All arrays should be 1d and should have the same shape")
  if levels is None:
    levels = np.sort(np.unique(np.hstack(x))) # Sort the unique values and then use this ordering as levels
  if reference is None:
    reference = levels[0]
  #print "Levels: %s, reference: %s" % (levels, reference)
  levels = levels[levels != reference] # Remove reference from levels
  level_len = len(levels)
  #print x[0].shape[0], level_len
  out = np.zeros((x[0].shape[0], level_len))
  for i, v in enumerate(levels):
    # print i, v
    for col in x:
      out[np.where(np.array(col) == v), i] = 1
  #print "Created matrix with shape: ", out.shape
  colnames = ["T.%s" % k for k in levels]
  if weights is not None:
    weights = weights.values
    return pd.DataFrame(out, columns=colnames).divide(weights, axis=0)

  return pd.DataFrame(out, columns=colnames)

X = patsy.dmatrix("multival(col1, col2, weights=num_vals)", data, return_type="dataframe")

There is no way of identifying the column names generated by patsy:

E.g. The names are like this:

multival(col1, col2, weights=num_vals, reference='O')[0]
multival(col1, col2, weights=num_vals, reference='O')[1]
multival(col1, col2, weights=num_vals, reference='O')[2]

I would much rather prefer column names like the following:

multival(col1, col2, weights=num_vals, reference='O')[T.0]
multival(col1, col2, weights=num_vals, reference='O')[T.1]
multival(col1, col2, weights=num_vals, reference='O')[T.2]

OR if I am passing the levels variable:

multival(col1, col2, weights=num_vals, reference='O')[T.A]
multival(col1, col2, weights=num_vals, reference='O')[T.B]
multival(col1, col2, weights=num_vals, reference='O')[T.C]

Is there a way to achieve this in patsy ?

Duplicate of statsmodels/statsmodels#2843

主要语言
Python
星标
990
派生
106
平均合并
7 天 34 分钟
30 天内合并 PR
1

环境准备

我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

pydata/patsy 的其他 Issue

查看 pydata/patsy 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。