Model2Vec backends share generated distillation vocabulary across independent corpora
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 88/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- backend, machine-learning
Research direction
Start with bertopic/backend/_model2vec.py and reproduce the shared-vocabulary case using the provided recording distiller. Verify default and explicit distill_kwargs ownership, explicit vocabularies, custom vectorizers, one-time distillation, and None settings against the described 12 passing cases. Done means independent backends build independent vocabularies without mutating caller dictionaries.
Written by the indexing model from the issue text.
Description
Problem
Model2VecBackend retains a shared default distill_kwargs dictionary and writes the corpus vocabulary into it during embed. A second default-constructed backend then reuses the first corpus vocabulary rather than building its own. Explicit dictionaries supplied by callers are modified too, and reusing one dictionary reproduces the contamination.
Reproduction
Against the inspected source revision:
from types import SimpleNamespace
from unittest.mock import patch
import numpy as np
from bertopic.backend._model2vec import Model2VecBackend
vocabularies = []
def recording_distiller(model, **kwargs):
vocabularies.append(list(kwargs["vocabulary"]))
return SimpleNamespace(encode=lambda docs, **options: np.zeros((len(docs), 3)))
# Record the vocabulary sent to the external distiller; no model is downloaded.
with patch("model2vec.distill.distill", side_effect=recording_distiller):
Model2VecBackend("source", distill=True).embed(["apple apple banana"])
Model2VecBackend("source", distill=True).embed(["orange orange lemon"])
print("first vocabulary:", vocabularies[0])
print("second vocabulary:", vocabularies[1])
Observed before the candidate change:
first vocabulary: ['apple', 'banana']
second vocabulary: ['apple', 'banana']
Proposed change
Use None as the default and take a fresh top-level dictionary copy for each backend. Keep explicit nonempty vocabulary authoritative and retain distillation-once behavior.
Verification
Before: 8 failed, 4 passed. Candidate: 12 passed, no skipped cases.
Backends created before and after the first distillation; default/shared/explicit settings; caller ownership; explicit vocabulary; custom bigram vectorizer; one-time distillation; None settings.
Complete backend and BaseEmbedder modules with real CountVectorizer were exercised. The external distiller and returned encoder are recording doubles; no model weights were downloaded and embedding quality was not measured.
Source: bertopic/backend/_model2vec.py, Git blob 425675802e9427246f53c4b74438cf83f0db7310; branch master. Python 3.13.5, Linux; exact dependency versions accompany the test logs.
Searches: PR: distill_kwargs; Issue: distill_kwargs. PR #2245 introduces the backend and provides usage context; it does not address shared configuration. This case replaces the duplicated c-TF-IDF investigation (#2034).
Candidate source patch
diff --git a/bertopic/backend/_model2vec.py b/bertopic/backend/_model2vec.py
--- a/bertopic/backend/_model2vec.py
+++ b/bertopic/backend/_model2vec.py
@@ -55,13 +55,13 @@
self,
embedding_model: Union[str, StaticModel],
distill: bool = False,
- distill_kwargs: dict = {},
+ distill_kwargs: dict | None = None,
distill_vectorizer: str | None = None,
):
super().__init__()
self.distill = distill
- self.distill_kwargs = distill_kwargs
+ self.distill_kwargs = dict(distill_kwargs) if distill_kwargs is not None else {}
self.distill_vectorizer = distill_vectorizer
self._has_distilled = False
Could a maintainer confirm this scope for a follow-up PR?
- Dominant language
- Python
- Stars
- 7.8k
- Forks
- 921
- Avg merge
- 5d 16h
- Merged PRs (30d)
- 1
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from MaartenGr/BERTopic
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Add context to theOpen
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
-
bug
Difficulty 3/5 1-2 days Newbie friendliness 68/100
All issues in MaartenGr/BERTopic
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
letsencrypt/cp-cps#353 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
DOI-USGS/pywatershed#421 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
python-pillow/Pillow#10087 · 1 comment ·
Maintainers usually reply within 1 day