One on-device model: Gemma 2 (2B) for every WebLLM feature; remove the model picker
Maintainers usually reply within 1 day
@s-annam is already working on this.
Since Sep 24, 2026.
Assessment
This issue has not been assessed yet.
Description
Summary
Ship one on-device model, Gemma 2 (2B) (gemma-2-2b-it-q4f16_1-MLC), for every WebLLM feature, and remove the model picker. The user never chooses or needs to understand a model: there is one download, one license acceptance, and the same behaviour everywhere.
Why:
- In the maintainer's hands-on runs, Gemma 2 gave better rewrites than Qwen 2.5 (1.5B) and than Gemma 3 (1B). Gemma 3 1B produced degenerate output: it echoed the instructions and looped (see comments).
- The picker was a source of confusion. Each "which model am I on, and what happens when the default changes" question (see the earlier Part 2 of this issue) goes away when there's nothing to choose.
This supersedes the earlier plan for this issue (add Gemma 3 to the picker, then decide the default). The earlier body is in the edit history.
Decisions (maintainer, 2026-09-24)
- One model for all features. Rewrite (section and whole résumé), critique, the parse escape hatch (
useLlmEscapeHatch), JD match (useJdMatch,run-llm-match.ts) and sector classification (src/lib/job-search/sector.ts) all run on Gemma 2. - Consent comes first for everyone. Gemma is Restricted-Community (Gemma Terms of Use, https://ai.google.dev/gemma/terms). No model file is downloaded or loaded until the user has accepted the terms.
- A user-initiated feature (rewrite, critique, escape hatch, JD match, "download ahead") opens the existing
ConsentDialogon first use. Accepting it continues the action; declining leaves everything unchanged. - A background feature (sector classification) never prompts. Without recorded consent it takes its existing heuristic fallback.
- Consent is recorded against the shipped model id, not the license type. That way a user who accepted Llama's license earlier isn't treated as having accepted Gemma's. Old
offlinecv:webllm:consent:<LicenseType>keys are ignored and removed in the cleanup below.
- A user-initiated feature (rewrite, critique, escape hatch, JD match, "download ahead") opens the existing
- One-time, silent cleanup. On first load after the change, delete the cached weights of models no longer offered, using web-llm's
deleteModelAllInfoInCache: Qwen 2.5 1.5B, Llama 3.2 3B and Gemma 3 1B. Also remove the obsoleteofflinecv:webllm:modelIdkey. Mark the cleanup done so it runs once. It fails open: an error is logged and never shown. - Eval stays a dev-time tool. By default the rewrite eval measures the shipped model. The eval harness keeps a dev-only way to add candidate models (web-llm
model_ids) so a future model can be compared before it ships. That list is not in the product bundle. - The Gemma 3 row and the
ModelMetadata.chatOptionsseam from the first version of #1016 are dropped, since nothing uses them.
Plan
src/lib/webllm/models.ts: replace the registry and the picker concepts (MODEL_REGISTRY,DEFAULT_MODEL_ID, tiers,isRegisteredModelId) with the one shipped model's metadata: id, name, license URL, download size (1895 MB). Rewrite the docblock: why one model, why Gemma 2, why consent is required, and what the #65 eval said (Gemma 2 scored 54% against Qwen's 67%). Record the hands-on reason behind the decision, and point to the eval as how a future change should be measured.- Delete the model picker (
ModelSelector.tsx, 556 lines) and replace it with a small status line in the same place (ReconstructedResume.tsx:439-443), next to the "Rewrite full résumé" trigger. It shows the model name; Download · ~1.9 GB (through consent), progress, or Ready · runs offline; Remove from this device; and the load error with Try again. KeepModelLoadProgress. useModelSelection.ts: shrink to consent state for the shipped model, or replace it. Every consumer that readselectedModelIduses the shipped id.- Put the consent gate in one shared place (a hook or helper). Don't copy it into each feature.
sector.ts: gate the LLM path on recorded consent, with the heuristic fallback otherwise.- Eval harness (
src/lib/webllm/eval/run-eval-browser.ts, and the parse-eval and jd-spike harnesses if they read the registry): default to the shipped model, plus a dev-only candidate list documented insrc/lib/webllm/eval/README.md. - Update
docs/ai-usage.mdand any copy or docblocks that describe a picker, several models, or Qwen as the model.
Acceptance criteria
- No model picker is rendered anywhere, and no UI copy asks the user to choose a model.
- Every WebLLM feature loads
gemma-2-2b-it-q4f16_1-MLC. No code path loads any other model in the product build. - No model download or load starts before consent is recorded. User-initiated features show
ConsentDialogfirst, and accepting continues the action. Sector classification falls back to its heuristic without consent. Each path has a test. - Consent is keyed to the shipped model id. An old license-type consent key does not count as consent.
- The one-time cleanup removes the retired models' caches and the old
modelIdkey, runs once, and fails open. It has tests. - The eval page measures the shipped model by default, and candidate models can be added through a documented dev-only list that isn't in the product bundle.
-
docs/ai-usage.mdand the docblocks describe one model. No stale "picker", "registry of three/four" or "Qwen is the default" copy remains in current-tense docs. Historical eval reports are left unchanged. - typecheck, lint and the affected tests pass.
Out of scope
- Prompt tuning for Gemma 2, and guardrails against degenerate output (instruction echo, repetition loops). The Gemma 3 run showed they're missing for every model, so they get their own issue.
- Hardware auto-routing.
- Dominant language
- TypeScript
- Stars
- 11
- Forks
- 4
- Avg merge
- 1d 24m
- Merged PRs (30d)
- 71
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from offlinecv/OfflineCV
-
chore gaal
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 92/100
Maintainers usually reply within 1 day
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 92/100
Maintainers usually reply within 1 day
-
Download PDF: preview the exact exported PDF, all pages, before saving itPossibly taken @s-annam claimed this today. Openfeature gaal ready-for-agent ux:edit-export
offlinecv/OfflineCV#1077 · 1 assignee ·
Maintainers usually reply within 1 day
-
chore gaal refactor
Difficulty 4/5 3-5 days Newbie friendliness 48/100
Maintainers usually reply within 1 day
All issues in offlinecv/OfflineCV
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
rohitg00/agentmemory#1428 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
boxlite-ai/boxlite#1729 ·
Maintainers usually reply within 1 day
-
detectors enhancement good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
SM260845/readme-gen#1 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
angular/angularfire#3774 ·
Maintainers usually reply within 2 days