Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

An unresolvable call is indistinguishable from no call

Aperta
#3,665 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

@nikhilsaxena04 ci sta già lavorando.

Dal 27/9/2026.

  • #3885 di @nikhilsaxena04 — aperta

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
65/100
Tipo di issue
Funzionalità
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
devtools

Direzione di ricerca

Start in extract.py at the for rc in all_raw_calls loop, then read _park_unresolved_member_call around L3225 and the callee handling in extractors/engine.py. Reproduce with graphify . --code-only --no-viz and verify that each resolver drop records caller, callee, location, and reason in metadata without emitting an edge.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Version: 0.9.63 (verified; also reproduced on 0.9.49)
Language: Python; the mechanism is language-independent

Summary

The cross-file call resolution loop in extract.py discards a call site with a bare continue in fourteen places. Each decision is correct. None of them is recorded.

The result is that a caller whose calls could not be resolved has no outgoing calls edge — which is byte-identical, in graph.json, to a caller that contains no calls at all. A consumer cannot tell "does not call X" from "I could not tell whether it calls X".

That matters because the two answers have opposite consequences. Anything built on reachability can go green on the second while believing it got the first.

Related but distinct

I searched open and closed issues before filing. #2447 and #3405 both observe this same ambiguity — #2447 puts it as "the failure is silent, and callers of the graph can't distinguish 'no edge' from 'no call'" — but the ask in each is better resolution, and each could be closed in full without a single drop ever being recorded. #2672 raises the same shape one layer up, in the skill.md Honesty Rules, for empty query / path results.

This request is independent of how much recall improves: however many shapes resolve, the ones that still don't should stop being invisible.

Reproduction

# demo.py
import pytest

RAW = "data/raw"

def reads_pinned_bytes():
    return open(RAW).read()

BUILDERS = {"cash": reads_pinned_bytes}

@pytest.fixture
def pinned_leg():
    return reads_pinned_bytes()

def test_via_registry():
    return BUILDERS["cash"]()

def test_via_fixture(pinned_leg):
    return pinned_leg

def test_via_getattr():
    return getattr(BUILDERS["cash"], "__call__")()
$ graphify . --code-only --no-viz

All three, on stock 0.9.63:

test_via_registry()   edges: NONE   unresolved_calls: NONE
test_via_fixture()    edges: NONE   unresolved_calls: NONE
test_via_getattr()    edges: NONE   unresolved_calls: NONE

reads_pinned_bytes is in the same file, and is reached by all three at run time.

Where it happens

Not in the callee scan. callee_name is set — the generic fallback in extractors/engine.py reads the call's function node directly, so a subscript callee yields the whole expression text and is queued normally:

raw_call | L15 | 'BUILDERS["cash"]'
raw_call | L21 | 'getattr(BUILDERS["cash"], "__call__")'

It is discarded in extract.py, in the for rc in all_raw_calls: loop (L7575-7784 in 0.9.63), at L7616-7620:

candidates = global_label_to_nids.get(callee, [])
if not candidates and _lang_is_case_insensitive(rc.get("source_file")):
    candidates = global_label_to_nids_ci.get(callee.lower(), [])
if not candidates:
    continue          # <- dies here, leaving no trace

Thirteen other continue statements in the same loop do the same thing for their own reasons — in 0.9.63 they sit at L7578, 7580, 7584, 7590, 7598, 7604, 7611, 7637, 7650, 7712, 7719, 7747 and 7760, covering empty callee, _LANGUAGE_BUILTIN_GLOBALS, is_member_call, Ruby mixin markers, language == "bash", Markdown, Go predeclared, cross-language family mismatch, Go import-path mismatch, unresolved ambiguity, source-less stubs, indirect dispatch and missing JS/TS import evidence.

The ask

Record the drop. Do not emit an edge — there is no target, and inventing one would be a guess presented as extraction.

The information needed is already in hand at every one of those fourteen sites: caller, callee text, source file, line, and the reason the resolver is giving up. So this is not an inference feature. It is the removal of a discard.

0.9.63 already has the right container. _park_unresolved_member_call (L3225) writes a per-node metadata["unresolved_calls"] list — names rather than node ids, deduped, capped at 64, for cross-repo member calls awaiting merge-graphs. This request is to widen that existing mechanism, not to add a new one: park every drop, with its reason, not only the ones that carry a known receiver_type.

A sketch of the shape (illustrative, not a proposed patch):

def _note_unresolved(rc, reason):
    ...  # same parking as _park_unresolved_member_call, plus:
    #   {"callee": rc["callee"], "reason": reason,
    #    "line": rc.get("source_location"), "confidence": "AMBIGUOUS"}

Why AMBIGUOUS rather than a low-confidence edge

The confidence vocabulary (EXTRACTED / INFERRED / AMBIGUOUS) already supports three-valued answers; the resolver currently collapses the third into silence. Reachability has to stay three-valued:

answer may a consumer treat it as proof?
reaches (EXTRACTED) yes
does not reach yes
cannot tell (INFERRED / AMBIGUOUS) no — abstain, never green

More edges is not automatically better. An INFERRED guess consumed as proof is an instrument reporting a verdict it never earned. The value here is precisely that the record is not an edge.

Measured impact

A private Python repository, 1476 nodes / 2667 edges, built with 0.9.63, instrumented to log each drop. The counts below are corroboration only — the reproduction above is self-contained and needs nothing from it.

call sites silently discarded: 2775
   2106  is_member_call
    374  not candidates
    215  language builtin
     33  bash
     19  indirect (reached the continue having emitted no edge)
     16  unresolved ambiguity
     12  source-less stub

Restricted to callable nodes in test files:

callable test-file nodes     : 449
  with no outgoing edge      : 115   (26%)
     ... explained by a drop record : 100
     ... genuinely silent           :  15

So 100 of 115 currently-silent callers gain a stated reason. The remaining 15 are mostly true negatives — callers that really make no resolvable call.

For comparison, the same repo on 0.9.49 had 196/424 (46%) edgeless, so cross-file recall has improved substantially in between. The ambiguity is unaffected by that improvement: a caller in the remaining 26% is exactly as unreadable as before.

Out of scope for this request

Two related gaps, deliberately not asked for here:

  1. pytest fixture injection. def test_x(some_fixture) is not a call, so it never reaches this loop and no record would be produced. Modelling it is a separate (and, for Python test corpora, valuable) feature. It is why exactly 5 of the 15 "genuinely silent" nodes above are in fact false silences — each takes a fixture parameter and reaches real code through it.
  2. Resolving dict-literal dispatch. BUILDERS["cash"] could in principle be resolved when the dict is a module-level literal of plain function references. That would yield INFERRED at best, since the dict can be mutated at run time — useful for a human asking "what calls this", not sufficient for anything gate-grade. Independent of this request.

This request is only that a drop stop being invisible.

One incidental finding

A cached rebuild skips the resolution loop entirely. Measuring this on a repo with an existing graphify-out/ produced zero records until the output directory was rebuilt from scratch. Not a bug report, but worth knowing for anyone reproducing the numbers above.

Lingua principale
Python
Stelle
124k
Fork
11.9k
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Graphify-Labs/graphify

Tutte le issue di Graphify-Labs/graphify

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.