Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Proposal] Sparse probing: optional groups argument so rows from one prompt can't straddle the split

オープン
#1,813 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

@lorenzozanee がすでに取り組んでいます。

2026年9月26日 から。

  • #1824 @lorenzozanee による — オープン

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
25/100
issue の種類
機能追加
明瞭さ
明確に書かれている
活発さ
停滞
技術スタック
python, pytorch

調査の方向性

Start with sparse_probing.py:251 and inspect fit_sparse_probe and sweep_sparse_probe, then review the guide’s leakage section and existing split tests. An open linked pull request (#1824) is already working on this proposal, so check its changes before considering any contribution. Done means group IDs do not cross the split, invalid groupings raise clearly, the supplied fixture scores near chance, the guide covers grouped splitting, and the listed checks pass.

索引モデルが issue の本文から書いたものです。

説明

complexity-simple enhancement help wanted TransformerBridge
Proposal

Add an optional groups: Integer[torch.Tensor, "example"] to fit_sparse_probe and sweep_sparse_probe that holds out whole groups, defaulting to today's row-level split when omitted (sparse_probing.py:251).

Motivation

The guide already warns that rows sharing a source prompt must not straddle the split, but nothing in the API lets a caller act on it. Flattening [batch, pos, d_model], where every position of a document carries that document's label, is the normal way to build features and has no safe form today.

Pitch

On 40 groups of 8 rows: each group a shared identity vector plus noise, labels assigned per group at random, so the honest answer is no signal:

def grouped(n_groups=40, per_group=8, d=64, seed=0):
    g = torch.Generator().manual_seed(seed)
    ident = torch.randn(n_groups, d, generator=g) * 3.0
    rows = ident.repeat_interleave(per_group, 0) + torch.randn(n_groups * per_group, d, generator=g)
    labels = (torch.rand(n_groups, generator=g) < 0.5).long().repeat_interleave(per_group)
    return rows, labels, torch.arange(n_groups).repeat_interleave(per_group)

X, y, groups = grouped(seed=0)
fit_sparse_probe(X, y, k=8, seed=0).metrics.f1   # 0.805

Row-level F1 is 0.72–0.81 across seeds 0-3 where a group-held-out split gives 0.38–0.70. The probe is reading group identity out of the training rows of the same group, and nothing in the result says so.

  • groups assigns each row a group id; the stratified split partitions groups instead of rows, both classes still on both sides.
  • Omitting it changes nothing, so no existing result moves.

Acceptance:

  • No group id appears in both train_indices and test_indices
  • Clear raise when the grouping can't keep both classes on both sides
  • The fixture above scores near chance with groups supplied
  • Guide's leakage section shows the groups form
  • make unit-test passes
  • uv run mypy . passes
Checklist
  • I have checked that there is no similar issue in the repo (required)
主要言語
Python
スター
3.9k
フォーク
708
平均マージ
1日 18時間
マージ済み PR(30日)
65

環境構築

Codespaces で開く

このプロジェクトの開発コンテナを、あなたの GitHub アカウントでブラウザ上に起動します。

  • Dockerfile・Docker Compose ファイルなし
  • プルリクエストのテンプレートあり
  • コントリビューションガイドなし

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

TransformerLensOrg/TransformerLens のほかの issue

TransformerLensOrg/TransformerLens の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。