Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

[Bug] Subgraph startup rehashes the WASM module once per data source (minutes for subgraphs with millions of dynamic data sources)

Aberta
#6,731 1 comentário 0 reações 0 responsáveis Ver no GitHub

@madumas já está trabalhando nisso.

Desde 9/10/2026.

  • #6732 de @madumas — aberto

Avaliação

Dificuldade
3/5
Tempo estimado
Meio dia
Facilidade para iniciantes
22/100
Tipo de issue
Bug
Clareza
Claramente especificada
Status de atividade
Estagnada
Stack de tecnologia
rust, wasm
Domínio
backend, performance

Direção de pesquisa

Start in SubgraphInstance::new_host in core/src/subgraph/context/instance/mod.rs, where keccak256 runs over the full module bytes before the module_cache lookup. Check whether the bytes can be identified once per shared Arc or per template instead of per data source. Pull request #6732 is already open against this issue, so review or build on it rather than starting a parallel fix. Done means each distinct module is hashed once per runner start.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

Bug report

When a subgraph runner starts, SubgraphInstance::new_host
(core/src/subgraph/context/instance/mod.rs) computes keccak256 over the
full WASM module bytes for every data source, and only then looks the hash
up in module_cache:

let module_hash = alloy::primitives::keccak256(module_bytes.as_ref()).0;
if let Some(sender) = self.module_cache.get(&module_hash) {
    sender.clone()
} else { /* spawn_mapping … */ }

module_cache avoids recompiling the module, but not rehashing it. All data
sources created from the same template share one Arc of module bytes, so the
same bytes are hashed again for each of them. Startup time is therefore
O(number of data sources × module size), where it could be
O(number of data sources + number of distinct modules × module size).

This is paid on every runner start: node restart, subgraph restart after an
error, unassign/reassign, etc. Hashing the module for every data source dates
back to v0.19.0 (in another form since v0.17), and is still on master
(core/src/subgraph/context/instance/mod.rs:107 at 6838f4e3c).

Expected: each distinct module is hashed once per runner start.
Actual: each module is hashed once per data source.

Impact

On the Uniswap v3 subgraph on Base (deployment QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ),
with 1,770,112 data sources (all but one created from the pool template, whose
module is 66,103 bytes), startup takes about 4 minutes, nearly all of it in this
loop: graph-node logs nothing between Data source count at start and
forcing subgraph to use static filters (~236 s in the log below).

Relevant log output
2026-10-09T19:27:41.705 INFO Resolve subgraph files using IPFS, n_templates: 1, n_data_sources: 1, runner_index: 0, subgraph_id: QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
2026-10-09T19:27:45.167 INFO Data source count at start: 1770112, runner_index: 0, subgraph_id: QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
2026-10-09T19:31:41.628 INFO forcing subgraph to use static filters., runner_index: 0, subgraph_id: QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
2026-10-09T19:31:41.688 INFO Start processing block, triggers: 66, runner_index: 0, subgraph_id: QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
IPFS hash

QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ

Subgraph name or link to explorer

Uniswap v3 (Base)

Some information to help us out
  • Tick this box if this bug is caused by a regression found in the latest release.
  • Tick this box if this bug is specific to the hosted service.
  • I have searched the issue tracker to make sure this issue is not a duplicate.
OS information

Linux

Linguagem predominante
Rust
Estrelas
3.2k
Forks
1.1k
Métricas de merge de PRs
Nenhum PR com merge em 30d

Preparar o ambiente

Abrir no Codespaces

Inicia o contêiner de desenvolvimento do projeto no navegador, com a sua própria conta do GitHub.

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de graphprotocol/graph-node

Todas as issues de graphprotocol/graph-node

Issues semelhantes

Mais issues de Rust

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.