Sigmoid test adequacy
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 25/100
- Issue 类型
- 功能
- 描述清晰度
- 需要澄清
- 活跃度
- 活跃
- 技术栈
- python
调研方向
The issue does not name files, tests, or entry points. Start by locating the current causal test adequacy calculation based on bootstrapped causal-effect kurtosis and the plots that display it; clarify the sigmoid transformation, score direction, and proposed stability labels before changing behavior.
由索引模型根据 Issue 内容生成。
描述
Summary
Current causal test adequacy measurement, based on the kurtosis of bootstrapped causal effect estimates is unintuitive. It can be arbitrarily negative or positive, and zero is the best possible score (although because it's a statistical value, this isn't ever perfectly achievable except with infinite resamples of and infinite dataset). To make it more intuitive, we discussed putting it through two "half sigmoids" to bring positive values into the range [0, 1] and negative values into the range [0, -1]. We also discussed adding the labels "suspiciously stable" and "suspiciously unstable" to the plots.
Considerations
- Target - Typically test adequacy is a "numbers go up" game where 100% is the goal. At the moment, we have 0 being the goal. The obvious easy thing to do here, based on the solution above, would be to simply transform the raw kurtosis number to bring it into the range [-1, 1], multiply by 100, and there's the "percentage" (although it's not a percentage of anything), so 0% is still the goal. If we can work out how, it'd be really cool to make it so that high numbers are better so that 100% is good (i.e., represents 0 kurtosis) and -100% is bad (i.e. represents -1 kurtosis). I'm not sure how you'd implement this conceptually, though.
- What should be exponentially harder to reach - @SylviaWhittle from your explanation, it seems like the "obvious easy solution" mentioned above would make it exponentially more difficult to reach the worse values of test adequacy (corresponding to kurtosis values of -1 and 1). From the perspective of "once the raw values get sufficiently large, we can't really scream YOU NEED MORE DATA any louder", this sort of makes sense,
but I can also see an argument for making it exponentially harder to achieve the best possible adequacy (corresponding to kurtosis of 0) to reflect the diminishing returns of additional data values. Or would this just serve to further diminish the diminishing returns.
Update: This was a stupid idea - of course it makes sense to have the "linear" part of the sigmoid be approaching zero kurtosis rather than the "exponential" bit since the "final approach" to zero (which can never actually be reached because statistics) will be more or less the same for every system. While I have seen kurtosis values of 50+, this isn't super common, and especially not if you've got anywhere near enough data. Based on "vibes", I'd guess that having the linear bit start around 1 (or -1) would be about right, but I have absolutely no formal basis for saying that. If you can find anything theoretical to justify that, that'd be cool. If not, we can look at this dataset and try to find a suitable value and empirical justification. I do still think 100% rather than 0% should represent "good" though if we can.
@SylviaWhittle, I'm happy to discuss either of these points further if you'd like, but I'm also keen not to micromanage you and to give you some creative freedom here to play with stuff and see what works for you. It's great to have your input here, since you're a much better representative of the sort of person who we're hoping will eventually use the framework than I am (albeit an extremely capable and eager one).
- 主要语言
- Python
- 星标
- 19
- 派生
- 7
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
CITCOM-project/CausalTestingFramework 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 68/100
-
难度 4/5 3-5 天 新手友好度 35/100
-
More tutorials未关闭
难度 4/5 3-5 天 新手友好度 45/100
-
难度 3/5 1-2 天 新手友好度 45/100
-
难度 5/5 一周以上 新手友好度 35/100
查看 CITCOM-project/CausalTestingFramework 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
approved correction metadata
难度 1/5 1 小时以内 新手友好度 88/100
acl-org/acl-anthology#10133 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
BasedHardware/omi#20084 ·
维护者通常 1 天内回复
-
bug needs-acceptance wg/evaluation-quality
难度 2/5 1-3 小时 新手友好度 76/100
vllm-project/semantic-router#4424 ·
维护者通常 1 天内回复