Sigmoid test adequacy
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- data-visualization, testing-qa
Hướng nghiên cứu
The issue does not name files, tests, or entry points. Start by locating the current causal test adequacy calculation based on bootstrapped causal-effect kurtosis and the plots that display it; clarify the sigmoid transformation, score direction, and proposed stability labels before changing behavior.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
Current causal test adequacy measurement, based on the kurtosis of bootstrapped causal effect estimates is unintuitive. It can be arbitrarily negative or positive, and zero is the best possible score (although because it's a statistical value, this isn't ever perfectly achievable except with infinite resamples of and infinite dataset). To make it more intuitive, we discussed putting it through two "half sigmoids" to bring positive values into the range [0, 1] and negative values into the range [0, -1]. We also discussed adding the labels "suspiciously stable" and "suspiciously unstable" to the plots.
Considerations
- Target - Typically test adequacy is a "numbers go up" game where 100% is the goal. At the moment, we have 0 being the goal. The obvious easy thing to do here, based on the solution above, would be to simply transform the raw kurtosis number to bring it into the range [-1, 1], multiply by 100, and there's the "percentage" (although it's not a percentage of anything), so 0% is still the goal. If we can work out how, it'd be really cool to make it so that high numbers are better so that 100% is good (i.e., represents 0 kurtosis) and -100% is bad (i.e. represents -1 kurtosis). I'm not sure how you'd implement this conceptually, though.
- What should be exponentially harder to reach - @SylviaWhittle from your explanation, it seems like the "obvious easy solution" mentioned above would make it exponentially more difficult to reach the worse values of test adequacy (corresponding to kurtosis values of -1 and 1). From the perspective of "once the raw values get sufficiently large, we can't really scream YOU NEED MORE DATA any louder", this sort of makes sense,
but I can also see an argument for making it exponentially harder to achieve the best possible adequacy (corresponding to kurtosis of 0) to reflect the diminishing returns of additional data values. Or would this just serve to further diminish the diminishing returns.
Update: This was a stupid idea - of course it makes sense to have the "linear" part of the sigmoid be approaching zero kurtosis rather than the "exponential" bit since the "final approach" to zero (which can never actually be reached because statistics) will be more or less the same for every system. While I have seen kurtosis values of 50+, this isn't super common, and especially not if you've got anywhere near enough data. Based on "vibes", I'd guess that having the linear bit start around 1 (or -1) would be about right, but I have absolutely no formal basis for saying that. If you can find anything theoretical to justify that, that'd be cool. If not, we can look at this dataset and try to find a suitable value and empirical justification. I do still think 100% rather than 0% should represent "good" though if we can.
@SylviaWhittle, I'm happy to discuss either of these points further if you'd like, but I'm also keen not to micromanage you and to give you some creative freedom here to play with stuff and see what works for you. It's great to have your input here, since you're a much better representative of the sort of person who we're hoping will eventually use the framework than I am (albeit an extremely capable and eager one).
- Ngôn ngữ chính
- Python
- Star
- 19
- Fork
- 7
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của CITCOM-project/CausalTestingFramework
-
Testing dashboardĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
-
More tutorialsĐang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 45/100
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Tất cả issue của CITCOM-project/CausalTestingFramework
Issue tương tự
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedĐang mởworkflow
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
New Submission: TropWATERĐang mởmetadata submission
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
-
Wrongly named dashboard variableĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
canonical/content-cache-operator#163 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[submission]Đang mởsubmission
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 65/100
leanprover/lean-eval-submissions#1852 ·
Maintainer thường phản hồi trong vòng 1 ngày