Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Application catalog + persistent context store (serve many applications from pre-computed contexts)

Open
#6 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
python

Research direction

Start with ROADMAP.md item 1, then read cache/manager.py, cache/spill.py, backends/tabicl.py, serve/app.py, and client.py to map the existing cache and serialization paths. The work is done when persisted contexts can be shared across replicas, version mismatches fail safely and trigger re-fitting, and the persistence, catalog, deletion, preload, and batch-fit surfaces meet the listed acceptance criteria.

Written by the indexing model from the issue text.

Description

enhancement roadmap

Why

The first prospective users described their deployment shape (2026-09-01): an admin uploads a cohort table per application, TabICLv2 encodes it once, the encoded context (kv cache) is saved, and end users only ever call predict against their application -- one deployment serving many applications from pre-computed contexts.

The README section Serving many applications from pre-computed contexts states exactly how far v0.10.0 gets you. In short: publishing a table per application (upload_csv + fit_uploaded), selecting it by dataset_id under a tenant, mixing both models under one budget, and reloading evicted contexts from the spill tier all work. What does not exist is any notion of a context as a durable, addressable artifact: spilling is eviction-only, replica-local, best-effort and LRU-bounded; a context that was never evicted is lost on restart; nothing can list, pin, persist or delete a context through the API.

Build order

  • ContextStore protocol (persist / load / list / delete by id) with a local-directory implementation first and an object-storage (GCS/S3) one second. The existing DiskSpillStore becomes the eviction tier on top of it rather than a parallel mechanism. Contexts gain a pinned flag (exempt from eviction) and a persisted state.
  • Versioned, checked serialization. Serialized contexts are pickled estimators tied to the tabicl/tabpfn/torch versions and device they were fit on. Files must carry tabctx + backend + torch versions and the device; a mismatch re-fits from the stored table instead of loading garbage.
  • API surface: POST /v1/tabctx/datasets/{id}/persist, DELETE /v1/tabctx/datasets/{id} (today nothing but LRU pressure removes a context), GET /v1/tabctx/datasets per tenant (task, model, shape, feature names, residency: GPU / spilled / persisted). Feature names already travel with the context since v0.10.0, so the catalog is mostly an endpoint. Client methods to match.
  • Warm-load manifest at replica start (contexts to preload from the store) and a batch tabctx fit entry point that writes to the store without a serving replica (offline pre-computation).
  • Roles at the gateway. Tenancy is namespacing, not authorization: a tenant that can predict can also re-fit. "Admin publishes, users predict" needs read-only tenant credentials at the authenticating proxy (see serve/tenancy.py's trust model); per-application aliases (/v1/apps/<app>/predict) stay a gateway concern.

Acceptance

  • A context persisted by one replica is loadable by a fresh deployment (different pod, same store) and predicts identically to the original -- the same identity check benchmarks/probe_deployment.py already runs for spill restore.
  • Multi-replica: the consistent-hash router may send a dataset to any replica; each must be able to load it from the shared store, so the store is shared by construction.
  • A version-mismatched artifact is refused with a clear error and falls back to re-fit, never a silent wrong prediction.

Pointers

cache/manager.py (CachedContext, ContextCacheManager), cache/spill.py (DiskSpillStore, the serializer hooks dumps_payload/loads_payload on backends/tabicl.py), serve/app.py, client.py, ROADMAP.md item 1.

Dominant language
Python
Stars
1
Forks
0
Avg merge
12h 16m
Merged PRs (30d)
14

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from VectorInstitute/tabctx

All issues in VectorInstitute/tabctx

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.