Add `swift` family support for `IQ3_S` (Swift 1.5 now has an IQ3_S tier)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 66/100
Research direction
The change is in the model table in setup.py: the IQ3_S entry has families set to ("qwen",), which blocks the swift family. Check how FAMILIES["swift"] builds shard names and how the Q2_0 swift exclusion (#171) is handled, to confirm the IQ3_S shards need no extra pack-tool handling. Done when --family swift --model IQ3_S appears in the model choices, the stale comment is removed, and the qwen choices are unchanged.
Written by the indexing model from the issue text.
Description
Summary
IQ3_S is currently restricted to the qwen family in setup.py, based on the
assumption that "Swift 1.5 has no IQ3_S". That assumption is now out of date:
UkisAI released an IQ3_S tier for Swift 1.5 on 2026-10-08. Please allow
--family swift --model IQ3_S.
Evidence
setup.py:
# the original model only (Swift 1.5 has no IQ3_S): matches the full BF16 model on the published benchmarks
"IQ3_S": {"about": "3.5-bit i-quant, the best quality (matches the full model), the slowest; needs a 64 GB PC "
"with little else running", "download_gb": 83.6, "ram_gb": 62, "arena_gb": 50.3,
"families": ("qwen",)},
Because families defaults to ("qwen", "swift"), the explicit ("qwen",) on
IQ3_S is what blocks the Swift family. With the current release:
--family swift -> --model choices: IQ2_XS, IQ3_XXS
--family qwen -> --model choices: Q2_0, IQ2_XS, IQ3_XXS, IQ3_S
The files already exist and match your naming convention
Model: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF
The two shards match the Swift pattern in FAMILIES["swift"]
(Swift-Qwen3.8-Flash-Next-GSQ-RCO-{q}-0000{i}-of-00002.gguf) exactly:
| File | Bytes | SHA-256 |
|---|---|---|
Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf |
44,900,816,736 | 46cb3996cd1a17fe8bc5ae1f9bb16869480de5a03f17f868d6047a20a46edd3b |
Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf |
38,842,920,384 | 70ce6fb48a94fe34bd7e9a53a526be68066c64b5a5077cdc82fb31350b66b608 |
Combined size: 83.74 GB (decimal). The BF16 mmproj is the same file already
shared by the other Swift tiers
(mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf,
SHA-256 cd1140f4abba943ce8d9fa92fed8407e389954d26907efde7b72153c75bd8a60).
Upstream reports IQ3_S as the lowest-KLD of the four tiers (development KLD
0.144674, vs 0.240139 for IQ3_XXS), so it is the quality-ceiling option for
this family.
Request
- Add
"swift"toIQ3_S'sfamilies(and drop the stale comment), so
--family swift --model IQ3_Sis selectable. - Please sanity-check the Swift-specific packaging constraints that already
affected this family — e.g. the SwiftQ2_0exclusion in the same dict
(#171, one layer's experts split across the two shards) and the
--ple-gguf/ PLE-table placement, since these differ between tiers. - If IQ3_S's Swift shards turn out to need extra pack-tool handling, please
keep it gated rather than silently enabling it.
Environment (where this was hit)
- Strata
v0.1.41 - 2× Tesla T10 16 GB (sm_75, Turing), 247 GiB RAM
- Weights already downloaded and SHA-256 verified against the official
SHA256SUMS/release-manifest.json.
Happy to test a patch on this hardware and report back.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)Possibly taken @nekomario28 claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
Maintainers usually reply within 1 day
-
hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7Possibly taken @nekomario28 claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Niko1221/Strata#1463 · 2 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Niko1221/Strata#1322 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Maintainers usually reply within 1 day
Similar issues
-
Status: Awaiting triage
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
espressif/arduino-esp32#12984 ·
Maintainers usually reply within 1 day
-
torch_ops/logprob.cu does not compile with the serving container's nvcc (13.3.73); check_torch_ops.py cannot run as shippedPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 Under an hour Newbie friendliness 72/100
ashhart/TensorFold#535 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 Under an hour Newbie friendliness 78/100
sudoevolve/EUI-NEO#95 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day