Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Model2Vec backends share generated distillation vocabulary across independent corpora

Open Beginner friendly
#2,544 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
88/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Start with bertopic/backend/_model2vec.py and reproduce the shared-vocabulary case using the provided recording distiller. Verify default and explicit distill_kwargs ownership, explicit vocabularies, custom vectorizers, one-time distillation, and None settings against the described 12 passing cases. Done means independent backends build independent vocabularies without mutating caller dictionaries.

Written by the indexing model from the issue text.

Description

Problem

Model2VecBackend retains a shared default distill_kwargs dictionary and writes the corpus vocabulary into it during embed. A second default-constructed backend then reuses the first corpus vocabulary rather than building its own. Explicit dictionaries supplied by callers are modified too, and reusing one dictionary reproduces the contamination.

Reproduction

Against the inspected source revision:

from types import SimpleNamespace
from unittest.mock import patch
import numpy as np
from bertopic.backend._model2vec import Model2VecBackend

vocabularies = []
def recording_distiller(model, **kwargs):
    vocabularies.append(list(kwargs["vocabulary"]))
    return SimpleNamespace(encode=lambda docs, **options: np.zeros((len(docs), 3)))

# Record the vocabulary sent to the external distiller; no model is downloaded.
with patch("model2vec.distill.distill", side_effect=recording_distiller):
    Model2VecBackend("source", distill=True).embed(["apple apple banana"])
    Model2VecBackend("source", distill=True).embed(["orange orange lemon"])
print("first vocabulary:", vocabularies[0])
print("second vocabulary:", vocabularies[1])

Observed before the candidate change:

first vocabulary: ['apple', 'banana']
second vocabulary: ['apple', 'banana']

Proposed change

Use None as the default and take a fresh top-level dictionary copy for each backend. Keep explicit nonempty vocabulary authoritative and retain distillation-once behavior.

Verification

Before: 8 failed, 4 passed. Candidate: 12 passed, no skipped cases.

Backends created before and after the first distillation; default/shared/explicit settings; caller ownership; explicit vocabulary; custom bigram vectorizer; one-time distillation; None settings.

Complete backend and BaseEmbedder modules with real CountVectorizer were exercised. The external distiller and returned encoder are recording doubles; no model weights were downloaded and embedding quality was not measured.

Source: bertopic/backend/_model2vec.py, Git blob 425675802e9427246f53c4b74438cf83f0db7310; branch master. Python 3.13.5, Linux; exact dependency versions accompany the test logs.

Searches: PR: distill_kwargs; Issue: distill_kwargs. PR #2245 introduces the backend and provides usage context; it does not address shared configuration. This case replaces the duplicated c-TF-IDF investigation (#2034).

Candidate source patch

diff --git a/bertopic/backend/_model2vec.py b/bertopic/backend/_model2vec.py
--- a/bertopic/backend/_model2vec.py
+++ b/bertopic/backend/_model2vec.py
@@ -55,13 +55,13 @@
         self,
         embedding_model: Union[str, StaticModel],
         distill: bool = False,
-        distill_kwargs: dict = {},
+        distill_kwargs: dict | None = None,
         distill_vectorizer: str | None = None,
     ):
         super().__init__()
 
         self.distill = distill
-        self.distill_kwargs = distill_kwargs
+        self.distill_kwargs = dict(distill_kwargs) if distill_kwargs is not None else {}
         self.distill_vectorizer = distill_vectorizer
         self._has_distilled = False
 

Could a maintainer confirm this scope for a follow-up PR?

Dominant language
Python
Stars
7.8k
Forks
921
Avg merge
5d 16h
Merged PRs (30d)
1

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from MaartenGr/BERTopic

All issues in MaartenGr/BERTopic

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.