How to use mutli-column functions in Patsy which return categorical output?
还没有人认领这个 Issue。
评估
调研方向
从 issue 中展示的 patsy.dmatrix 调用和自定义 multival 函数开始,然后查看 statsmodels/statsmodels#2843 中链接的重复 issue。当自定义多列分类输出能够显示有意义的基于级别的列名,而不是数字索引时,工作即告完成。
由索引模型根据 Issue 内容生成。
描述
I am dealing with a situation where each item can have 1 or 2 labels both of which come from the same set. In this case I only want to have 1 set of categorical variables which encode if the values of each level are present or absent with respect to the reference. I also want to inversely weight the values of these categorical variables based on how many labels are present.
E.g. Let my categories be {A, B, C}
data = pd.DataFrame({"col1": ["A", "A", "B"], "col2": ["O", "B", "C"], "num_vals": [1,2,2]})
def inverse_val(x):
return 1.0/x
# Using categorical coding will not give the correct output
X = patsy.dmatrix("(C(col1) + C(col2)):inverse_val(num_vals)", data, return_type="dataframe")
I tried to solve this issue using the following code:
def multival(*x, **kwargs):
#raise Exception("Not Implemented")
levels, reference = kwargs.get("levels", None), kwargs.get("reference", None)
weights = kwargs.get("weights", None)
if len(x[0].shape) != 1:
raise Exception("Mismatching Shapes. All arrays should be 1d and should have the same shape")
for k in x:
if k.shape != x[0].shape:
raise Exception("Mismatching Shapes. All arrays should be 1d and should have the same shape")
if levels is None:
levels = np.sort(np.unique(np.hstack(x))) # Sort the unique values and then use this ordering as levels
if reference is None:
reference = levels[0]
#print "Levels: %s, reference: %s" % (levels, reference)
levels = levels[levels != reference] # Remove reference from levels
level_len = len(levels)
#print x[0].shape[0], level_len
out = np.zeros((x[0].shape[0], level_len))
for i, v in enumerate(levels):
# print i, v
for col in x:
out[np.where(np.array(col) == v), i] = 1
#print "Created matrix with shape: ", out.shape
colnames = ["T.%s" % k for k in levels]
if weights is not None:
weights = weights.values
return pd.DataFrame(out, columns=colnames).divide(weights, axis=0)
return pd.DataFrame(out, columns=colnames)
X = patsy.dmatrix("multival(col1, col2, weights=num_vals)", data, return_type="dataframe")
There is no way of identifying the column names generated by patsy:
E.g. The names are like this:
multival(col1, col2, weights=num_vals, reference='O')[0]
multival(col1, col2, weights=num_vals, reference='O')[1]
multival(col1, col2, weights=num_vals, reference='O')[2]
I would much rather prefer column names like the following:
multival(col1, col2, weights=num_vals, reference='O')[T.0]
multival(col1, col2, weights=num_vals, reference='O')[T.1]
multival(col1, col2, weights=num_vals, reference='O')[T.2]
OR if I am passing the levels variable:
multival(col1, col2, weights=num_vals, reference='O')[T.A]
multival(col1, col2, weights=num_vals, reference='O')[T.B]
multival(col1, col2, weights=num_vals, reference='O')[T.C]
Is there a way to achieve this in patsy ?
Duplicate of statsmodels/statsmodels#2843
- 主要语言
- Python
- 星标
- 990
- 派生
- 106
- 平均合并
- 7 天 34 分钟
- 30 天内合并 PR
- 1
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
pydata/patsy 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 72/100
-
难度 1/5 1 小时以内 新手友好度 72/100
-
难度 1/5 1 小时以内 新手友好度 68/100
-
难度 2/5 1-3 小时 新手友好度 55/100
-
难度 5/5 一周以上 新手友好度 25/100
相似的 Issue
-
customer-reported
难度 2/5 1-3 小时 新手友好度 68/100
Azure/azure-cli#34150 · 1 条评论 ·
维护者通常 1 天内回复
-
community-request
难度 1/5 1 小时以内 新手友好度 95/100
NVIDIA-NeMo/Curator#2464 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
WeblateOrg/translation-finder#1099 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 68/100
trezor/trezor-firmware#7997 ·
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复