Sigmoid test adequacy
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- data-visualization, testing-qa
Direzione di ricerca
The issue does not name files, tests, or entry points. Start by locating the current causal test adequacy calculation based on bootstrapped causal-effect kurtosis and the plots that display it; clarify the sigmoid transformation, score direction, and proposed stability labels before changing behavior.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
Current causal test adequacy measurement, based on the kurtosis of bootstrapped causal effect estimates is unintuitive. It can be arbitrarily negative or positive, and zero is the best possible score (although because it's a statistical value, this isn't ever perfectly achievable except with infinite resamples of and infinite dataset). To make it more intuitive, we discussed putting it through two "half sigmoids" to bring positive values into the range [0, 1] and negative values into the range [0, -1]. We also discussed adding the labels "suspiciously stable" and "suspiciously unstable" to the plots.
Considerations
- Target - Typically test adequacy is a "numbers go up" game where 100% is the goal. At the moment, we have 0 being the goal. The obvious easy thing to do here, based on the solution above, would be to simply transform the raw kurtosis number to bring it into the range [-1, 1], multiply by 100, and there's the "percentage" (although it's not a percentage of anything), so 0% is still the goal. If we can work out how, it'd be really cool to make it so that high numbers are better so that 100% is good (i.e., represents 0 kurtosis) and -100% is bad (i.e. represents -1 kurtosis). I'm not sure how you'd implement this conceptually, though.
- What should be exponentially harder to reach - @SylviaWhittle from your explanation, it seems like the "obvious easy solution" mentioned above would make it exponentially more difficult to reach the worse values of test adequacy (corresponding to kurtosis values of -1 and 1). From the perspective of "once the raw values get sufficiently large, we can't really scream YOU NEED MORE DATA any louder", this sort of makes sense, but I can also see an argument for making it exponentially harder to achieve the best possible adequacy (corresponding to kurtosis of 0) to reflect the diminishing returns of additional data values. Or would this just serve to further diminish the diminishing returns.
@SylviaWhittle, I'm happy to discuss either of these points further if you'd like, but I'm also keen not to micromanage you and to give you some creative freedom here to play with stuff and see what works for you. It's great to have your input here, since you're a much better representative of the sort of person who we're hoping will eventually use the framework than I am (albeit an extremely capable and eager one).
- Lingua principale
- Python
- Stelle
- 19
- Fork
- 7
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di CITCOM-project/CausalTestingFramework
-
Testing dashboardAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
-
More tutorialsAperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 45/100
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
Tutte le issue di CITCOM-project/CausalTestingFramework
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
spec-kitty/spec-kitty#5319 ·
I maintainer di solito rispondono entro 1 giorno
-
backend::vllm diffusion multimodal
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
openai/openai-agents-python#5229 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno