[Bug] Subgraph startup rehashes the WASM module once per data source (minutes for subgraphs with millions of dynamic data sources)
Valutazione
- Difficoltà
- 3/5
- Tempo stimato
- Mezza giornata
- Idoneità per principianti
- 22/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Ferma
- Stack tecnologico
- rust, wasm
- Ambito
- backend, performance
Direzione di ricerca
Start in SubgraphInstance::new_host in core/src/subgraph/context/instance/mod.rs, where keccak256 runs over the full module bytes before the module_cache lookup. Check whether the bytes can be identified once per shared Arc or per template instead of per data source. Pull request #6732 is already open against this issue, so review or build on it rather than starting a parallel fix. Done means each distinct module is hashed once per runner start.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Bug report
When a subgraph runner starts, SubgraphInstance::new_host
(core/src/subgraph/context/instance/mod.rs) computes keccak256 over the
full WASM module bytes for every data source, and only then looks the hash
up in module_cache:
let module_hash = alloy::primitives::keccak256(module_bytes.as_ref()).0;
if let Some(sender) = self.module_cache.get(&module_hash) {
sender.clone()
} else { /* spawn_mapping … */ }
module_cache avoids recompiling the module, but not rehashing it. All data
sources created from the same template share one Arc of module bytes, so the
same bytes are hashed again for each of them. Startup time is therefore
O(number of data sources × module size), where it could be
O(number of data sources + number of distinct modules × module size).
This is paid on every runner start: node restart, subgraph restart after an
error, unassign/reassign, etc. Hashing the module for every data source dates
back to v0.19.0 (in another form since v0.17), and is still on master
(core/src/subgraph/context/instance/mod.rs:107 at 6838f4e3c).
Expected: each distinct module is hashed once per runner start.
Actual: each module is hashed once per data source.
Impact
On the Uniswap v3 subgraph on Base (deployment QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ),
with 1,770,112 data sources (all but one created from the pool template, whose
module is 66,103 bytes), startup takes about 4 minutes, nearly all of it in this
loop: graph-node logs nothing between Data source count at start and
forcing subgraph to use static filters (~236 s in the log below).
Relevant log output
2026-10-09T19:27:41.705 INFO Resolve subgraph files using IPFS, n_templates: 1, n_data_sources: 1, runner_index: 0, subgraph_id: QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
2026-10-09T19:27:45.167 INFO Data source count at start: 1770112, runner_index: 0, subgraph_id: QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
2026-10-09T19:31:41.628 INFO forcing subgraph to use static filters., runner_index: 0, subgraph_id: QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
2026-10-09T19:31:41.688 INFO Start processing block, triggers: 66, runner_index: 0, subgraph_id: QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
IPFS hash
QmTEVcvbvWdW1LhNCdGpUFTxjxGmA34aRUksvX8TPjrQkZ
Subgraph name or link to explorer
Uniswap v3 (Base)
Some information to help us out
- Tick this box if this bug is caused by a regression found in the latest release.
- Tick this box if this bug is specific to the hosted service.
- I have searched the issue tracker to make sure this issue is not a duplicate.
OS information
Linux
- Lingua principale
- Rust
- Stelle
- 3.2k
- Fork
- 1.1k
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di graphprotocol/graph-node
-
current: include emits an all-null bucket for dimensionless aggregations, nulling the whole responseAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
graphprotocol/graph-node#6719 ·
-
RUSTSEC-2026-0194: Quadratic run time when checking a start tag for duplicate attribute namesForse già presa @szupzj18 l’ha presa 59 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
graphprotocol/graph-node#6673 ·
-
RUSTSEC-2026-0185: Remote memory exhaustion in quinn-proto from unbounded out-of-order stream reassemblyForse già presa @abisheik687 l’ha presa 101 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
graphprotocol/graph-node#6650 · 1 commento ·
-
`loadRelated` can return stale children while blocks are queued for writing (`!= any` in `FindDerivedQuery`; unrelated queued writes not excluded)Forse già presa @madumas l’ha presa 6 giorni fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 50/100
graphprotocol/graph-node#6726 ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
graphprotocol/graph-node#6722 ·
Tutte le issue di graphprotocol/graph-node
Issue simili
-
C-bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
rust-lang/rust-analyzer#23501 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
[Bug]: Web chat input doesn't regain focus after a reply finishesForse già presa @GaijinSystems l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
zeroclaw-labs/zeroclaw#11658 ·
I maintainer di solito rispondono entro 2 giorni
-
good first issue help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
bytecodealliance/wasm-tools#2768 ·
I maintainer di solito rispondono entro 1 giorno