[Feature]: Precision, Recall and F1Score metrics
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 25/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
- 技術スタック
- kotlin
調査の方向性
This is an umbrella issue: implementation is split across #1224–#1235, so start by checking those issues and their dependencies. For the Kotlin core, read skainet-lang/skainet-lang-core/src/commonMain/kotlin/sk/ainet/lang/nn/metrics/Accuracy.kt and resolve the averaging and constructor questions in #1224 before #1226 starts. The parent is done when the listed acceptance criteria are met, including the three metrics, unit tests, API dump, documentation, and review.
索引モデルが issue の本文から書いたものです。
説明
🧠 D: DOCUMENT — Problem & Opportunity
sk.ainet.lang.nn.metrics has exactly one metric: Accuracy. Precision, recall and F1 —
the next three metrics any classifier evaluation needs, and the ones that actually say
something on imbalanced data — are missing. Every framework SKaiNET is compared against
or ported from (KotlinDL, PyTorch/torchmetrics, scikit-learn, Keras) ships all four as the
baseline set; today a SKaiNET training loop has to leave the library to compute them.
Summary: Add Precision, Recall and F1Score to nn/metrics/, built on one
shared per-class confusion-matrix accumulator, following the existing Metric interface
(update / compute / reset) and Accuracy's dtype / dim / threshold handling.
Design decision (Assess, made here so Code doesn't have to): F1 depends on precision
and recall, which are also missing. Three classes each re-deriving TP/FP/FN would
triplicate the non-trivial iteration logic already in
skainet-lang/skainet-lang-core/src/commonMain/kotlin/sk/ainet/lang/nn/metrics/Accuracy.kt
(dtype dispatch, hard vs. soft targets, argmax over dim, binary threshold). So: extract
Accuracy's private helpers once, build one internal ConfusionMatrixAccumulator, and
let the three metrics be three compute() formulas over it. Ship all three from one DARC
cycle — same data, three views.
DARC or SKEEP? DARC only. These are additive classes behind the existing Metric
interface: no public-API-shape, DSL, storage, or compiler footprint. The one part of this
work that does trip a SKEEP trigger — teaching the ground-truth harness to validate
stateful metrics, not just stateless tensor ops — is split out as its own proposal
(#1223) rather than being settled inside a sub-issue. See the docs page
Contributing → Worked example: F1Score via DARC for the full reasoning.
🔍 A: ASSESS — Feasibility & Impact
✔️ Feasibility
Straightforward for binary classification (one confusion matrix). Multi-class needs an
explicit averaging-strategy decision (macro / micro / per-class) — that is where the design
risk lives, not the arithmetic.
✔️ Expected Impact
Unblocks any training/eval loop that needs more than raw accuracy — most imbalanced
classification workloads — and closes a gap against every framework in the comparison set.
✔️ Risks / Constraints
- Zero-division: F1 is undefined when TP+FP+FN = 0 for a class.
sklearndefaults to
0.0 with a warning;torchmetricshas azero_divisionparameter. Pick one, document
it, never leave it as an unspecifiedNaN. Recommendation: return0.0, matching
Accuracy.compute()'s empty-accumulator behaviour. - Averaging semantics: macro (unweighted mean over classes) vs. micro (pooled TP/FP/FN)
vs. per-class give different numbers for the same predictions. A cross-check that falls
out of the definitions: for single-label multi-class via argmax, micro precision ==
micro recall == accuracy — worth a test. - Which classes does "macro" average over?
sklearnaverages over the union of labels
present iny_true ∪ y_pred; averaging over the full class dimension (preds.shape[dim])
gives a different number when a class never appears in a batch. Needs a decision
(#1224). - Ground-truth validation has no home yet:
OperationExecutorin
skainet-test-groundtruthmaps a name to a statelessTensorOpscall and returns one
tensor; aMetricis stateful and returns a scalar. This is a process risk, resolved
by shipping v1 with unit tests only and tracking the harness extension as SKEEP-005
(#1223).
✔️ Dependencies
Nothing new. Built entirely from tensor element access already used by Accuracy.
📚 R: RESEARCH — What Must Be Understood First?
Research Tasks
- Reference implementations:
sklearn.metrics.precision_recall_fscore_support,
torchmetrics.classification.{Precision,Recall,F1Score} - Formula:
F1 = 2·P·R / (P + R), per class; macro = mean over classes; micro = pooled - Averaging modes for v1 — recommend binary + macro + micro, per-class (no
reduction) as a follow-up — #1224 - Zero-division convention — recommend
0.0— #1224 - Class set for macro averaging (seen classes vs. full class dim) — #1224
Open Questions
Does
F1Scoretake athreshold: Float?mirroringAccuracy, or is binary handled
entirely byaveraging = Averaging.BINARY? Affects the public constructor shape; resolve
in #1224 before #1226 starts.
🛠️ C: CODE — Lane breakdown
Lanes per Contributing → Issue taxonomy. Lane 0 (SKEEP) is skipped for the metrics
themselves and lives in #1223; Lane 4 (ground-truth wiring) is deliberately not a
sub-issue here — it becomes actionable only once SKEEP-005 is accepted.
Sub-issues
| Lane | Issue | Skill | Size | Entry point | Blocked by |
|---|---|---|---|---|---|
| 1 · Numerics research | #1224 | skill:numerics |
s | good first issue — no Kotlin |
— |
2 · Kotlin core: extract Accuracy helpers |
#1225 | skill:kotlin-core |
s | good first issue |
— |
2 · Kotlin core: Averaging + ConfusionMatrixAccumulator |
#1226 | skill:kotlin-core |
s | #1224, #1225 | |
2 · Kotlin core: Precision |
#1227 | skill:kotlin-core |
s | good first issue |
#1226 |
2 · Kotlin core: Recall |
#1228 | skill:kotlin-core |
s | good first issue |
#1226 |
2 · Kotlin core: F1Score |
#1229 | skill:kotlin-core |
s | good first issue |
#1227, #1228 |
| 3 · Platform: Android | #1230 | skill:android |
s | good first issue |
#1229 |
| 3 · Platform: iOS | #1231 | skill:ios |
xs | good first issue |
#1229 |
| 3 · Platform: JS/Wasm | #1232 | skill:js |
xs | good first issue |
#1229 |
| 3 · Platform: Kotlin/Native | #1233 | skill:native |
xs | good first issue |
#1229 |
| 4 · Ground-truth / CI | — | — | — | not a sub-issue — see SKEEP-005 #1223 | SKEEP-005 accepted |
| 5 · Docs partials | #1234 | skill:docs |
s | good first issue |
(#1229 for the examples tag) |
| 6 · DARC review | #1235 | skill:review |
s | #1229, #1234 | |
| 0 · Design (SKEEP) | #1223 | skill:design |
m | separate tracking issue, not a sub-issue | — |
To claim a lane: comment on the sub-issue; a maintainer assigns it. Report results back here when it closes.
Acceptance Criteria
-
Precision,Recall,F1Score+ factory functions insk.ainet.lang.nn.metrics, API dump updated - Unit tests: hand-computed binary, macro and micro multi-class cases, zero-division edge case,
micro == accuracycross-check - Doc partials under
docs/modules/ROOT/partials/ops/metrics/, included from the metrics how-to - Reviewed by someone other than the implementer;
@DarcValidatedadded post-review - Ground-truth coverage explicitly stated as "unit tests only, pending SKEEP-005" in the PRs — a reviewed trade-off, not an oversight
- 主要言語
- Kotlin
- スター
- 52
- フォーク
- 15
- 平均マージ
- 1日 15時間
- マージ済み PR(30日)
- 36
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
SKaiNET-developers/SKaiNET のほかの issue
-
coding good first issue size:xs skill:kotlin-core sub-issue
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
SKaiNET-developers/SKaiNET#1323 ·
メンテナーはふだん 1 日以内に返信
-
coding good first issue platform size:xs skill:js sub-issue
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
SKaiNET-developers/SKaiNET#1232 ·
メンテナーはふだん 1 日以内に返信
-
skeep tracking
難易度 5/5 1週間以上 初心者へのやさしさ 15/100
SKaiNET-developers/SKaiNET#1331 ·
メンテナーはふだん 1 日以内に返信
-
assessment size:s skill:review sub-issue
難易度 4/5 1〜2日 初心者へのやさしさ 18/100
SKaiNET-developers/SKaiNET#1330 ·
メンテナーはふだん 1 日以内に返信
-
documentation good first issue size:s skill:docs sub-issue
難易度 3/5 1〜2日 初心者へのやさしさ 78/100
SKaiNET-developers/SKaiNET#1329 ·
メンテナーはふだん 1 日以内に返信
SKaiNET-developers/SKaiNET の issue をすべて見る
似ている issue
-
[Submission] 抖音火山版オープンsubmit-adaption submit-adaption-pre
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
BetterAndroid/android-notification-icon-project#744 · コメント 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 64/100
utopia-rise/godot-jvm#1004 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
pedroSG94/RootEncoder#2213 ·
メンテナーはふだん 2 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
dzid26/TeslaBatteryBLE#183 ·
メンテナーはふだん 1 日以内に返信