BertScorer early-stopping default ([email protected]) collapses on sparse multilabel
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python
- Domain
- machine-learning
Research direction
Start in autointent/modules/scoring/_bert.py, reading the default EarlyStoppingConfig, _get_compute_metrics(), and _train(). Reproduce the GoEmotions multilabel behavior described in the issue, then compare the stopping metric with the available threshold-free or disabled-early-stopping options. Done means the default no longer stops immediately or restores a near-random checkpoint on sparse multilabel data while early stopping remains useful for multiclass tasks.
Written by the indexing model from the issue text.
Description
Summary
BertScorer's default early stopping uses scoring_f1 (F1 thresholded at 0.5) as metric_for_best_model. On sparse multilabel data, calibrated sigmoid probabilities sit well below 0.5, so this metric reads ≈0 from the very first eval, the EarlyStoppingCallback (patience 3) fires almost immediately, and load_best_model_at_end=True restores a near-random early checkpoint. The fitted model then outputs near-constant low probabilities → after thresholding the pipeline predicts essentially all-positive (degenerate).
Where
autointent/modules/scoring/_bert.py:
- Default
EarlyStoppingConfig→{'val_fraction': 0.2, 'patience': 3, 'threshold': 0.0, 'metric': 'scoring_f1'}. _get_compute_metrics()computesscoring_f1on the raw eval predictions (0.5-thresholded for multilabel)._train()setsmetric_for_best_model=self.early_stopping_config.metricandload_best_model_at_end=(metric is not None), plus anEarlyStoppingCallback.
Evidence
Fine-tuning bert-base-uncased on GoEmotions (28 classes, ~4% positive rate, ~2.5k balanced rows):
- With the default early stopping,
eval_scoring_f1is stuck at ~0.034 every epoch and the pipeline scores macrodecision_f1≈ 0.107 (degenerate all-positive; recall→1.0, precision→base rate). - With early stopping disabled (train full epochs) at the same LR, the same model learns and reaches macro
decision_f1≈ 0.17–0.41 depending on data size.
So the model is trainable; the threshold-0.5 early-stopping metric is what breaks it.
Suggested fix
Use a threshold-free signal for metric_for_best_model on multilabel:
eval_loss, or a ranking metric (e.g.neg_coverage/ MAP), or compute F1 at a tuned/optimal threshold rather than a fixed 0.5; or- disable early stopping by default for multilabel tasks.
Any of these would prevent the immediate-stop collapse while keeping early stopping useful for multiclass.
Environment
AutoIntent 0.3.1, MPS, transformers-no-hpo preset, GoEmotions multilabel.
- Dominant language
- Python
- Stars
- 53
- Forks
- 16
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from deeppavlov/AutoIntent
-
`OptimizationConfig.seed` is `PositiveInt` — `seed=0` is rejected while `Pipeline(seed=0)` accepts itPossibly taken @Gambit-Checkmate claimed this 21 days ago. Openbug good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
deeppavlov/AutoIntent#352 ·
-
Extract `BaseAPIDescriptionScorer` shared by `LLMDescriptionScorer` and `TypeSafeDescriptionScorer`Openenhancement
Difficulty 4/5 3-5 days Newbie friendliness 48/100
deeppavlov/AutoIntent#357 ·
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 52/100
deeppavlov/AutoIntent#356 ·
-
enhancement
Difficulty 3/5 1-2 days Newbie friendliness 72/100
deeppavlov/AutoIntent#355 · 1 comment ·
-
`LLMDescriptionScorer` runs cache hits through the `max_per_second` limiter — 10 s per 100 cached utterancesPossibly taken @kayaal34 claimed this 21 days ago. Openenhancement
Difficulty 4/5 3-5 days Newbie friendliness 58/100
deeppavlov/AutoIntent#354 ·
All issues in deeppavlov/AutoIntent
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
topoteretes/cognee#5647 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Sendspin/sendspin-python-cli#291 ·
Maintainers usually reply within 6 days
-
Assertion-shape guard fails on development: vacuous recorded-iteration assertion in the assetLinks batch_get helper testPossibly taken A pull request linked to this issue is open or already merged. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
awslabs/visual-asset-management-system#414 ·
Maintainers usually reply within 1 day
-
bug v1 v2
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
modelcontextprotocol/python-sdk#3670 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
aicell-lab/bioengine#232 ·
Maintainers usually reply within 1 day