enhancement: migrate pre-existing GPUStack model files into the node cache
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 30/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- go
- Domain
- infrastructure
Research direction
Start with pkg/modelmanager/store and the modelManager.rootPath chart value, then trace the existing ModelArtifact resolution and publish pipeline. Compare an offline import tool with plugin RPC, copy versus move, and shared-NFS handling; done means an agreed migration design that preserves hashing, manifest verification, and the store's single-writer invariant.
Written by the indexing model from the issue text.
Description
Background
The operator's node model cache is content-addressed and only admits bytes it has
verified itself. Hosts that already run GPUStack (server/worker product) commonly
hold large model caches downloaded by the legacy worker. Today there is no path
for those pre-existing files to enter the node cache: every model is re-downloaded
from the Hub (or pulled from a peer node) even when the same bytes already sit on
the same disk.
This issue collects the two designs as-found and sketches how a migration could
work.
Current GPUStack design (as-found)
Verified against the GPUStack repository (main branch) and a four-host GPU lab
running a mix of all-in-one and server/worker deployments.
- Default paths:
data_dir=/var/lib/gpustack,cache_dir=<data_dir>/cache
(gpustack/config/config.py), overridable with--data-dir/--cache-dir
orGPUSTACK_DATA_DIR/GPUSTACK_CACHE_DIR. - On-disk layout is flat:
<cache>/huggingface/<org>/<name>/<repo files>
viahf_hub_download(local_dir=...)(gpustack/worker/downloaders.py), with
a<name>.locksibling and partial files under.cache/huggingface/download/.
ModelScope is the same shape:<cache>/model_scope/<org>/<name>/with a
._____temp/directory. This is not themodels--<org>--<name>blobs/
snapshots layout — that was the v0.x-era layout, and an alembic migration
moved old installs to the flat one. - The
ModelFiletable (gpustack/schemas/model_files.py) recordssource,
local_dir,worker_id,resolved_paths(JSON),size,state,
source_index,cleanup_on_delete— no per-file digest. - Engines consume the path directly (vLLM positional argument, SGLang
--model-path); there is no snapshot/symlink/copy indirection. Mirrored
deployments bind-mount all of/var/lib/gpustackinto inference containers
(VOLUME /var/lib/gpustackinpack/Dockerfile). The GPUStack helm chart
mountsworker.dataDir(default/var/lib/gpustack) as a hostPath. - Model evaluation pulls config skeletons in the native
models--layout
into the samecache/huggingface/directory, so two layout generations
coexist in one cache. - Observed in the wild: caches ranging from empty to ~5.1 TB living on a shared
NFS export consumed by two hosts at once, both layout generations mixed, many
config-only skeletons, single-file GGUFs underollama/, and several
processes (pipx systemd service + docker containers) reading and writing the
same tree concurrently.
Operator design (after the model-artifact work)
rootPathdefaults to/var/lib/gpustack/models(chart value
modelManager.rootPath).- Layout (
pkg/modelmanager/store): alayoutmarker file, plus
published/<hex>/{marker,manifest.json,tree/...}(read-only after publish),
partial/<hex>/<attempt>/,trash/, andledger/{refs,digests}/.
Single-writer, atomic rename on publish, per-file SHA-256 in the manifest. store.Open()reads only thelayoutmarker: foreign files anywhere under
the root are not scanned, not adopted, and not deleted. A layout mismatch
refuses to start. This is deliberate — the store only trusts bytes it
verified during its own write pipeline.- GC accounts capacity filesystem-wide (statfs), so foreign files consume
capacity but are never evicted. - The store forbids shared inodes, so a migration cannot hard-link files into
the tree.
The two roots do not collide: the legacy product never writes under
/var/lib/gpustack/models, so old and new caches can coexist on one host.
The gap
There is no adoption path from cache/ into models/. Everything the node
cache serves is downloaded (or peer-pulled) and hashed by the plugin itself,
even when identical bytes are already local.
Options
A. Offline import tool (lazy, verify-then-publish). For each candidate
model on the host: enumerate the flat files, apply the artifact's file filters,
hash while copying into a partial/ attempt, verify the tree against the
resolved ModelArtifact's manifestDigest, then publish through the store's
own pipeline (same code path as a network download, source = local disk).
"Lazy" means only models actually referenced by a ModelArtifact are imported
— never bulk-copy a 5 TB cache.
B. Do nothing — re-download / peer-pull. Zero new code. Correct but wastes
bandwidth and time on hosts that already hold the bytes; on a shared NFS cache
serving several hosts the waste is multiplied per host.
C. Rename-in-place after retiring the legacy worker. Once the old
deployment is gone, move the flat tree into a store attempt and publish in
place — no copy, no double occupancy. Breaks the legacy install, so it only
fits decommissioning migrations, and still needs hashing + manifest
verification before publish.
Recommended shape: A with lazy scoping, keeping B as the always-available
baseline.
Blockers / dependencies
- No digests in the legacy DB. Every imported file must be hashed in full.
Measurements on disk-bound nodes put a post-hoc re-read-and-hash pass at
roughly 44–46% of the original download duration; hashing while copying
avoids the second read but still costs one full read. - Resolution needs the Hub. Verifying against
manifestDigestrequires a
resolvedModelArtifact, which today means Hub reachability. Air-gapped
hosts would need the deferredexpectedDigest/pre-resolved artifact path. - Single writer. The store has exactly one writer (the plugin). An import
tool must either run while the plugin is stopped, or go through a new plugin
RPC so the single-writer invariant holds. - GC candidacy. An imported tree with zero references becomes a GC
candidate after the grace period. That is fine — a deployment that needs it
will claim it — but import tooling should not fight the collector for
just-imported trees. - Shared caches. Where several hosts share one NFS export, per-host import
duplicates work; the tool needs a story for dedup/locking, or should simply
declare shared-cache hosts out of scope and let peer sync propagate.
Questions to decide
- One-shot CLI run by an operator, or a plugin RPC so migration is online?
- Copy (safe, doubles peak usage) vs move (cheap, breaks the legacy install)?
- Shared-NFS caches: dedup/lock, or explicitly out of scope?
- Dominant language
- Go
- Stars
- 4
- Forks
- 6
- Avg merge
- 1h 52m
- Merged PRs (30d)
- 375
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from gpustack/gpustack-operator
-
enhancement: A recreated Devices ledger never restores the node's accelerator counting capacityOpenkind/enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
gpustack/gpustack-operator#146 ·
Maintainers usually reply within 1 day
-
todo
Difficulty 5/5 Over a week Newbie friendliness 35/100
gpustack/gpustack-operator#601 ·
Maintainers usually reply within 1 day
-
kind/enhancement
Difficulty 5/5 Over a week Newbie friendliness 35/100
gpustack/gpustack-operator#599 ·
Maintainers usually reply within 1 day
-
area/worker kind/bug
Difficulty 4/5 3-5 days Newbie friendliness 45/100
gpustack/gpustack-operator#547 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
gpustack/gpustack-operator#512 ·
Maintainers usually reply within 1 day
All issues in gpustack/gpustack-operator
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
gruntwork-io/boilerplate#329 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
prime-radiant-inc/evener#3291 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Netcracker/qubership-apihub-backend#582 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day