[ET-VK] Embedding with vocab > 16384 silently returns wrong values on the texture path; force_fp16 breaks all-MiniLM-L6-v2
Mantenedores costumam responder em até 1 dia
@SS-JIA já está trabalhando nisso.
Desde 10/9/2026.
Avaliação
Esta issue ainda não foi avaliada.
Descrição
🐛 Describe the bug
aten.embedding returns wrong values on Vulkan when the vocabulary exceeds 16384 rows and the output is a texture. There is no error; the model produces well-formed but semantically wrong embeddings.
This breaks sentence-transformers/all-MiniLM-L6-v2 (vocab 30522) under force_fp16.
Threshold
nn.Embedding(V, 384), 64 random indices, cosine of the Vulkan output against the CPU reference on an Adreno 840:
| vocab | fp32 | fp16 |
|---|---|---|
| 2048 | 1.000000 | 1.000000 |
| 4096 | 1.000000 | 1.000000 |
| 8192 | 1.000000 | 1.000000 |
| 16384 | 1.000000 | 1.000000 |
| 16385 | 0.255916 | 0.255916 |
| 20000 | 0.256048 | 0.256048 |
| 30522 | 0.259350 | 0.259349 |
The break is exactly at 16384, which is this device's maxImageDimension2D, and it is precision-independent. All of these dispatch embedding_texture3d_{float,half}.
Why force_fp16 turns this into a real-model failure
all-MiniLM-L6-v2 at the published 254-token shape, 8 sentences, embeddings compared against the CPU reference:
| mean cosine | pairwise similarity max abs diff | top-1 nearest neighbour preserved | |
|---|---|---|---|
| fp32 | 0.999999 | 0.00031 | 8 / 8 |
| fp16 | 0.219010 | 0.68950 | 1 / 8 |
One embedding comes back at cosine -0.016 against its reference. The semantic structure is destroyed, so this silently breaks retrieval rather than merely degrading it.
The two differ only in which embedding kernel they reach:
fp32: embedding_buffer_float
fp16: embedding_texture3d_half
force_fp16 biases the graph toward texture storage (TagMemoryMetaPass.constrain_op_arg_repset calls try_constrain_with_arg_repset(arg_i, utils.ANY_TEXTURE) unconditionally when force_fp16 is set), which moves the embedding output from buffer to texture and onto the broken path. fp32 escapes only by landing on the buffer kernel.
What I ruled out
For the MiniLM failure specifically: dynamic shapes (a static export is bit-identical), the padding and attention mask (a fully unmasked 254-token input fails the same, cosine 0.169), and the mean-pool/normalise tail (the pre-pool token output is already wrong). Exporting only model.embeddings reproduces it at cosine 0.33, before any encoder layer.
For the isolated case: it is not the legacy texture-weight path. embedding() in Embedding.cpp prepacks the weight as kBuffer for these models, and the dispatched kernel is the non-legacy embedding_texture3d_*. I tried guarding the legacy branch on max_texture2d_dim() and it changed nothing, confirming that branch is not involved.
I did not find the exact mechanism inside embedding_texture.glsl. load_weight_texel() computes embedding_idx * width(weight) + dim_idx, which does not overflow int32 at these sizes, so the 16384 boundary most likely comes from how the indices or the weight are addressed rather than from that multiply.
Suggested direction
Two things seem worth separating:
force_fp16should not push a tensor toward texture storage without consulting texture limits, so oversized cases keep the working buffer kernel.- Exceeding texture extents should fail loudly rather than silently returning garbage.
Repro
Scripts are straightforward to reconstruct from the table above: export nn.Embedding(V, 384) with VulkanPartitioner() for V on either side of 16384 and compare against the CPU reference.
Versions
ExecuTorch 1.4.1 for the export, runtime at c27baa8031. Device: Samsung Galaxy S26 Ultra, Snapdragon SM8850, Adreno 840, Android 16.
cc @SS-JIA @manuelcandales @digantdesai @cbilgin @mergennachin @kimishpatel @iseeyuan
- Linguagem predominante
- Python
- Estrelas
- 5.1k
- Forks
- 1.2k
- Merge médio
- 2d 6h
- PRs com merge (30d)
- 535
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Tem um modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de pytorch/executorch
-
enhancement triaged
Dificuldade 2/5 Meio dia Facilidade para iniciantes 68/100
pytorch/executorch#21640 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 68/100
pytorch/executorch#23262 ·
Mantenedores costumam responder em até 1 dia
-
XNNPACK never delegates slice_copy with stride != 1, but op-support.csv only excludes zero-dim/dynamic shapesTalvez já em andamento @JakeStevens assumiu há 1 dia. Abertamodule: xnnpack
pytorch/executorch#23258 · 1 responsável ·
Mantenedores costumam responder em até 1 dia
-
[RFC] ExecuTorch Persisting Device Specialized Delegate ArtifactsTalvez já em andamento @JacobSzwejbka assumiu há 2 dias. Aberta
pytorch/executorch#23192 · 5 comentários · 1 reação · 1 responsável ·
Mantenedores costumam responder em até 1 dia
-
[QNN] Enable ConvTranspose + BatchNorm fusion after #23170Talvez já em andamento @psiddh assumiu há 2 dias. Abertamodule: qnn partner: qualcomm
pytorch/executorch#23185 · 1 comentário · 1 reação · 1 responsável ·
Mantenedores costumam responder em até 1 dia
Todas as issues de pytorch/executorch
Issues semelhantes
-
repo-audit
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
scverse/repo-health#20 ·
Mantenedores costumam responder em até 1 dia
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memoriesAberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 85/100
phasespace-labs/palinode#232 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 82/100
collective/icalendar#1858 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
Mantenedores costumam responder em até 1 dia
-
lfx-mcp cannot supply global variables: LangflowClient drops X-LANGFLOW-GLOBAL-VAR-* from envAbertabug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
langflow-ai/langflow#15496 ·
Mantenedores costumam responder em até 1 dia