Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Feature]: Precision, Recall and F1Score metrics

オープン
#1,222 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
25/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
静か
技術スタック
kotlin

調査の方向性

This is an umbrella issue: implementation is split across #1224–#1235, so start by checking those issues and their dependencies. For the Kotlin core, read skainet-lang/skainet-lang-core/src/commonMain/kotlin/sk/ainet/lang/nn/metrics/Accuracy.kt and resolve the averaging and constructor questions in #1224 before #1226 starts. The parent is done when the listed acceptance criteria are met, including the three metrics, unit tests, API dump, documentation, and review.

索引モデルが issue の本文から書いたものです。

説明

darc enhancement tracking training

🧠 D: DOCUMENT — Problem & Opportunity

sk.ainet.lang.nn.metrics has exactly one metric: Accuracy. Precision, recall and F1 —
the next three metrics any classifier evaluation needs, and the ones that actually say
something on imbalanced data — are missing. Every framework SKaiNET is compared against
or ported from (KotlinDL, PyTorch/torchmetrics, scikit-learn, Keras) ships all four as the
baseline set; today a SKaiNET training loop has to leave the library to compute them.

Summary: Add Precision, Recall and F1Score to nn/metrics/, built on one
shared per-class confusion-matrix accumulator, following the existing Metric interface
(update / compute / reset) and Accuracy's dtype / dim / threshold handling.

Design decision (Assess, made here so Code doesn't have to): F1 depends on precision
and recall, which are also missing. Three classes each re-deriving TP/FP/FN would
triplicate the non-trivial iteration logic already in
skainet-lang/skainet-lang-core/src/commonMain/kotlin/sk/ainet/lang/nn/metrics/Accuracy.kt
(dtype dispatch, hard vs. soft targets, argmax over dim, binary threshold). So: extract
Accuracy's private helpers once, build one internal ConfusionMatrixAccumulator, and
let the three metrics be three compute() formulas over it. Ship all three from one DARC
cycle — same data, three views.

DARC or SKEEP? DARC only. These are additive classes behind the existing Metric
interface: no public-API-shape, DSL, storage, or compiler footprint. The one part of this
work that does trip a SKEEP trigger — teaching the ground-truth harness to validate
stateful metrics, not just stateless tensor ops — is split out as its own proposal
(#1223) rather than being settled inside a sub-issue. See the docs page
Contributing → Worked example: F1Score via DARC for the full reasoning.

🔍 A: ASSESS — Feasibility & Impact

✔️ Feasibility

Straightforward for binary classification (one confusion matrix). Multi-class needs an
explicit averaging-strategy decision (macro / micro / per-class) — that is where the design
risk lives, not the arithmetic.

✔️ Expected Impact

Unblocks any training/eval loop that needs more than raw accuracy — most imbalanced
classification workloads — and closes a gap against every framework in the comparison set.

✔️ Risks / Constraints
  • Zero-division: F1 is undefined when TP+FP+FN = 0 for a class. sklearn defaults to
    0.0 with a warning; torchmetrics has a zero_division parameter. Pick one, document
    it, never leave it as an unspecified NaN. Recommendation: return 0.0, matching
    Accuracy.compute()'s empty-accumulator behaviour.
  • Averaging semantics: macro (unweighted mean over classes) vs. micro (pooled TP/FP/FN)
    vs. per-class give different numbers for the same predictions. A cross-check that falls
    out of the definitions: for single-label multi-class via argmax, micro precision ==
    micro recall == accuracy
    — worth a test.
  • Which classes does "macro" average over? sklearn averages over the union of labels
    present in y_true ∪ y_pred; averaging over the full class dimension (preds.shape[dim])
    gives a different number when a class never appears in a batch. Needs a decision
    (#1224).
  • Ground-truth validation has no home yet: OperationExecutor in
    skainet-test-groundtruth maps a name to a stateless TensorOps call and returns one
    tensor; a Metric is stateful and returns a scalar. This is a process risk, resolved
    by shipping v1 with unit tests only and tracking the harness extension as SKEEP-005
    (#1223).
✔️ Dependencies

Nothing new. Built entirely from tensor element access already used by Accuracy.

📚 R: RESEARCH — What Must Be Understood First?

Research Tasks
  • Reference implementations: sklearn.metrics.precision_recall_fscore_support,
    torchmetrics.classification.{Precision,Recall,F1Score}
  • Formula: F1 = 2·P·R / (P + R), per class; macro = mean over classes; micro = pooled
  • Averaging modes for v1 — recommend binary + macro + micro, per-class (no
    reduction) as a follow-up — #1224
  • Zero-division convention — recommend 0.0 — #1224
  • Class set for macro averaging (seen classes vs. full class dim) — #1224
Open Questions

Does F1Score take a threshold: Float? mirroring Accuracy, or is binary handled
entirely by averaging = Averaging.BINARY? Affects the public constructor shape; resolve
in #1224 before #1226 starts.

🛠️ C: CODE — Lane breakdown

Lanes per Contributing → Issue taxonomy. Lane 0 (SKEEP) is skipped for the metrics
themselves and lives in #1223; Lane 4 (ground-truth wiring) is deliberately not a
sub-issue here
— it becomes actionable only once SKEEP-005 is accepted.

Sub-issues

Lane Issue Skill Size Entry point Blocked by
1 · Numerics research #1224 skill:numerics s good first issue — no Kotlin —
2 · Kotlin core: extract Accuracy helpers #1225 skill:kotlin-core s good first issue —
2 · Kotlin core: Averaging + ConfusionMatrixAccumulator #1226 skill:kotlin-core s #1224, #1225
2 · Kotlin core: Precision #1227 skill:kotlin-core s good first issue #1226
2 · Kotlin core: Recall #1228 skill:kotlin-core s good first issue #1226
2 · Kotlin core: F1Score #1229 skill:kotlin-core s good first issue #1227, #1228
3 · Platform: Android #1230 skill:android s good first issue #1229
3 · Platform: iOS #1231 skill:ios xs good first issue #1229
3 · Platform: JS/Wasm #1232 skill:js xs good first issue #1229
3 · Platform: Kotlin/Native #1233 skill:native xs good first issue #1229
4 · Ground-truth / CI — — — not a sub-issue — see SKEEP-005 #1223 SKEEP-005 accepted
5 · Docs partials #1234 skill:docs s good first issue (#1229 for the examples tag)
6 · DARC review #1235 skill:review s #1229, #1234
0 · Design (SKEEP) #1223 skill:design m separate tracking issue, not a sub-issue —

To claim a lane: comment on the sub-issue; a maintainer assigns it. Report results back here when it closes.

Acceptance Criteria
  • Precision, Recall, F1Score + factory functions in sk.ainet.lang.nn.metrics, API dump updated
  • Unit tests: hand-computed binary, macro and micro multi-class cases, zero-division edge case, micro == accuracy cross-check
  • Doc partials under docs/modules/ROOT/partials/ops/metrics/, included from the metrics how-to
  • Reviewed by someone other than the implementer; @DarcValidated added post-review
  • Ground-truth coverage explicitly stated as "unit tests only, pending SKEEP-005" in the PRs — a reviewed trade-off, not an oversight
主要言語
Kotlin
スター
52
フォーク
15
平均マージ
1日 15時間
マージ済み PR(30日)
36

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

SKaiNET-developers/SKaiNET のほかの issue

SKaiNET-developers/SKaiNET の issue をすべて見る

似ている issue

Kotlin の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。