Paper ↔ implementation comparison for RPG-Encoder (arXiv:2602.02084) + optional gap notes
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 20/100
- Issue type
- Documentation
- Clarity
- Needs clarification
- Activity status
- Stale
- Domain
- documentation
Research direction
Start by reading the paper alongside the referenced implementation files, including rpg-core/src/graph.rs, rpg-encoder/src/evolution.rs, and rpg-nav/src/search.rs. Verify which listed deltas are relevant to RPG-ZeroRepo and identify the project’s intended incremental-update behavior. Done would require an agreed comparison scope and documented, reproducible findings.
Written by the indexing model from the issue text.
Description
Hi RPG-ZeroRepo team — I read the RPG-Encoder paper (arXiv:2602.02084). Really impressive work.
I built a from-scratch implementation because I wanted something I could use immediately in my own workflow across multiple repos, and I made it MCP-ready so it can be plugged into agent/tool setups out of the box.
Repo: https://github.com/userFRM/rpg-encoder (MIT)
If you think it’s useful for others while your official release lands, feel free to link it from here / issues / discussions — totally fine by me.
What’s implemented (mapped to the paper at a high level)
RPG graph structure
- Paper:
G = (V_H ∪ V_L, E_feat ∪ E_dep) - Impl:
RPGraph,HierarchyNode,Entity,EdgeKind - Code:
rpg-core/src/graph.rs
Phase 1 — Semantic Lifting
- token-aware batching, verb–object feature extraction, file synthesis flow
- Code:
rpg-encoder/src/lift.rs,semantic_lifting.rs
Phase 2 — Hierarchy Recovery
- domain discovery + 3-level semantic path assignment
- Code:
rpg-encoder/src/hierarchy.rs,prompts/*.md
Phase 3 — Artifact Grounding
- trie-based LCA grounding + AST dependency resolution
- Code:
rpg-core/src/lca.rs,rpg-encoder/src/grounding.rs
Evolution / incremental maintenance
- git diff detection + delete/modify/insert application + drift detection
- Code:
rpg-encoder/src/evolution.rs
Agent tools described in the paper
SearchNode: feature/snippet search, auto modes, scope/type/line filters (rpg-nav/src/search.rs)FetchNode: fetch entity/hierarchy node + source (rpg-nav/src/fetch.rs)ExploreRPG: upstream/downstream BFS, edge/type filtering, depth control (rpg-nav/src/explore.rs)
Paper-specific details
- file-level module entities (
V_L):EntityKind::Module,create_module_entities()(rpg-core/src/graph.rs, ~619) - feature edge materialization (
E_feat):materialize_containment_edges()(rpg-core/src/graph.rs, ~698) - 8 language parsers: Python, Rust, Java, Go, C, C++, JS, TS (
rpg-parser/src/)
Notes / deltas vs the paper text (kept short)
- Benchmarks / eval: I haven’t reproduced SWE-bench / RepoCraft metrics, and I don’t have your internal evaluation harness/scripts, so I can’t do an apples-to-apples benchmark comparison between implementations.
- Reconstruction traversal export: no explicit “topo-ordered node list” export yet (exports are dot/mermaid, etc.).
- Incremental semantic insertion (Algorithm 3): incremental additions currently land via structural/file-path placement in
apply_additions(). Semantic placement is achieved via a full hierarchy rebuild (build_semantic_hierarchy+submit_hierarchy), not per-entity LLM routing during every update. - Drift judgment: paper mentions LLM-based “intent shift”; this uses a deterministic Jaccard-based drift metric (
compute_drift()).
Extras (practical additions not in the paper)
- stale graph detection on MCP server startup
- graph backup before destructive ops
- optional zstd compression
- pre-commit hook to auto-update on commit
rpg-encoder diffpreview +rpg-encoder validateintegrity checks- schema versioning + migration (semver)
- search quality benchmarks + npm distribution conveniences
If you ever want to use any part of what I built, I’m happy to help however is easiest on your side — I can share a couple tiny sample repos + their expected graph outputs (so you can sanity-check behavior), or split out specific pieces (parsers / grounding / nav tools) if that’s more useful. And if you prefer PRs or just issues, I’ll follow whatever contribution process you use.
One thing I’m curious about: for incremental updates, do you plan to do the paper’s “semantic routing per new entity” (Algorithm 3), or is your intended approach more like “put new stuff in structurally, then rebuild the semantic hierarchy in batches”?
- Dominant language
- Python
- Stars
- 553
- Forks
- 42
- PR merge metrics
- No merged PRs in 30d
Getting set up
- Ships a Dockerfile or Docker Compose file
- No pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/RPG-ZeroRepo
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
microsoft/RPG-ZeroRepo#100 ·
-
Difficulty 3/5 1-2 days Newbie friendliness 35/100
microsoft/RPG-ZeroRepo#101 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 28/100
microsoft/RPG-ZeroRepo#3 · 2 reactions ·
-
Difficulty 5/5 Over a week Newbie friendliness 15/100
microsoft/RPG-ZeroRepo#2 · 3 comments ·
-
Difficulty 1/5 Under an hour Newbie friendliness 1/100
microsoft/RPG-ZeroRepo#1 · 7 comments ·
All issues in microsoft/RPG-ZeroRepo
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
aicell-lab/bioengine#232 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
modelscope/evalscope#1836 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
jbaruch/speaker-toolkit#480 ·
Maintainers usually reply within 1 day