Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

One on-device model: Gemma 2 (2B) for every WebLLM feature; remove the model picker

Closed
#1,015 3 comments 0 reactions 1 assignee View on GitHub

Maintainers usually reply within 1 day

@s-annam is already working on this.

Since Sep 24, 2026.

Assessment

This issue has not been assessed yet.

Description

feature improvement

Summary

Ship one on-device model, Gemma 2 (2B) (gemma-2-2b-it-q4f16_1-MLC), for every WebLLM feature, and remove the model picker. The user never chooses or needs to understand a model: there is one download, one license acceptance, and the same behaviour everywhere.

Why:

  • In the maintainer's hands-on runs, Gemma 2 gave better rewrites than Qwen 2.5 (1.5B) and than Gemma 3 (1B). Gemma 3 1B produced degenerate output: it echoed the instructions and looped (see comments).
  • The picker was a source of confusion. Each "which model am I on, and what happens when the default changes" question (see the earlier Part 2 of this issue) goes away when there's nothing to choose.

This supersedes the earlier plan for this issue (add Gemma 3 to the picker, then decide the default). The earlier body is in the edit history.

Decisions (maintainer, 2026-09-24)

  1. One model for all features. Rewrite (section and whole résumé), critique, the parse escape hatch (useLlmEscapeHatch), JD match (useJdMatch, run-llm-match.ts) and sector classification (src/lib/job-search/sector.ts) all run on Gemma 2.
  2. Consent comes first for everyone. Gemma is Restricted-Community (Gemma Terms of Use, https://ai.google.dev/gemma/terms). No model file is downloaded or loaded until the user has accepted the terms.
    • A user-initiated feature (rewrite, critique, escape hatch, JD match, "download ahead") opens the existing ConsentDialog on first use. Accepting it continues the action; declining leaves everything unchanged.
    • A background feature (sector classification) never prompts. Without recorded consent it takes its existing heuristic fallback.
    • Consent is recorded against the shipped model id, not the license type. That way a user who accepted Llama's license earlier isn't treated as having accepted Gemma's. Old offlinecv:webllm:consent:<LicenseType> keys are ignored and removed in the cleanup below.
  3. One-time, silent cleanup. On first load after the change, delete the cached weights of models no longer offered, using web-llm's deleteModelAllInfoInCache: Qwen 2.5 1.5B, Llama 3.2 3B and Gemma 3 1B. Also remove the obsolete offlinecv:webllm:modelId key. Mark the cleanup done so it runs once. It fails open: an error is logged and never shown.
  4. Eval stays a dev-time tool. By default the rewrite eval measures the shipped model. The eval harness keeps a dev-only way to add candidate models (web-llm model_ids) so a future model can be compared before it ships. That list is not in the product bundle.
  5. The Gemma 3 row and the ModelMetadata.chatOptions seam from the first version of #1016 are dropped, since nothing uses them.

Plan

  • src/lib/webllm/models.ts: replace the registry and the picker concepts (MODEL_REGISTRY, DEFAULT_MODEL_ID, tiers, isRegisteredModelId) with the one shipped model's metadata: id, name, license URL, download size (1895 MB). Rewrite the docblock: why one model, why Gemma 2, why consent is required, and what the #65 eval said (Gemma 2 scored 54% against Qwen's 67%). Record the hands-on reason behind the decision, and point to the eval as how a future change should be measured.
  • Delete the model picker (ModelSelector.tsx, 556 lines) and replace it with a small status line in the same place (ReconstructedResume.tsx:439-443), next to the "Rewrite full résumé" trigger. It shows the model name; Download · ~1.9 GB (through consent), progress, or Ready · runs offline; Remove from this device; and the load error with Try again. Keep ModelLoadProgress.
  • useModelSelection.ts: shrink to consent state for the shipped model, or replace it. Every consumer that read selectedModelId uses the shipped id.
  • Put the consent gate in one shared place (a hook or helper). Don't copy it into each feature.
  • sector.ts: gate the LLM path on recorded consent, with the heuristic fallback otherwise.
  • Eval harness (src/lib/webllm/eval/run-eval-browser.ts, and the parse-eval and jd-spike harnesses if they read the registry): default to the shipped model, plus a dev-only candidate list documented in src/lib/webllm/eval/README.md.
  • Update docs/ai-usage.md and any copy or docblocks that describe a picker, several models, or Qwen as the model.

Acceptance criteria

  • No model picker is rendered anywhere, and no UI copy asks the user to choose a model.
  • Every WebLLM feature loads gemma-2-2b-it-q4f16_1-MLC. No code path loads any other model in the product build.
  • No model download or load starts before consent is recorded. User-initiated features show ConsentDialog first, and accepting continues the action. Sector classification falls back to its heuristic without consent. Each path has a test.
  • Consent is keyed to the shipped model id. An old license-type consent key does not count as consent.
  • The one-time cleanup removes the retired models' caches and the old modelId key, runs once, and fails open. It has tests.
  • The eval page measures the shipped model by default, and candidate models can be added through a documented dev-only list that isn't in the product bundle.
  • docs/ai-usage.md and the docblocks describe one model. No stale "picker", "registry of three/four" or "Qwen is the default" copy remains in current-tense docs. Historical eval reports are left unchanged.
  • typecheck, lint and the affected tests pass.

Out of scope

  • Prompt tuning for Gemma 2, and guardrails against degenerate output (instruction echo, repetition loops). The Gemma 3 run showed they're missing for every model, so they get their own issue.
  • Hardware auto-routing.
Dominant language
TypeScript
Stars
11
Forks
4
Avg merge
1d 24m
Merged PRs (30d)
71

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from offlinecv/OfflineCV

All issues in offlinecv/OfflineCV

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.