Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[Bug]: tokenize emits a non-existent new line token for empty input

Aperta
#1,133 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
3/5
Tempo stimato
1-2 giorni
Idoneità per principianti
74/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
backend

Direzione di ricerca

Inizia eseguendo l’esempio repro.py con GraalPy e CPython, quindi traccia l’implementazione GraalPy di _tokenize.TokenizerIter utilizzata da tokenize.generate_tokens. La correzione è completata quando un input vuoto produce solo l’ENDMARKER compatibile con CPython, con posizioni corrispondenti, mentre i casi esistenti di commenti e newline rimangono corretti.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug
Describe the bug

When using token.generate_tokens() on an empty string, a NL token is returned when no such token exists for CPython. The return token also spans a non-existent length, which caused stubgen-pyx to crash. Only effects "", # ... and "\n" were fine. This caused stubgen-pyx to crash on an otherwise empty file.

Likely occurs in GraalPy's re-implementation of _tokenize.TokenizerIter

Operating system

Linux

CPU architecture

x86_64

GraalPy version

GraalPy 3.13.14 (GraalVM CE Native 25.3.4.1)

JDK version

openjdk 25.0.4 2026-07-21 LTS (Temurin-25.0.4+7)

Context configuration

N/A

Steps to reproduce
# repro.py
import io, tokenize
for tok in tokenize.generate_tokens(io.StringIO("").readline):
    print(tok)
$ graalpy repro.py
TokenInfo(type=63 (NL), string='', start=(1, 0), end=(1, 1), line='')
TokenInfo(type=0 (ENDMARKER), string='', start=(2, 0), end=(2, 0), line='')
$ python3.13 repro.py          # and python3.14, identical
TokenInfo(type=0 (ENDMARKER), string='', start=(1, 0), end=(1, 0), line='')
Native version
import io, _tokenize
for t in _tokenize.TokenizerIter(io.StringIO("").readline, extra_tokens=True):
    print(t)
#GraalPy 3.13.14:  (63, '', (1, 0), (1, 1), '')      ← phantom NL
#                  (0,  '', (2, 0), (2, 0), '')
#CPython 3.13:     (0,  '', (1, 0), (1, 0), '')
Expected behavior

Should match the same output as CPython 3.13, and not emit non-existent spans.

Stack trace
N/A
Additional context

N/A

Lingua principale
Python
Stelle
1.6k
Fork
155
Merge medio
8h 20m
PR unite (30g)
43

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di oracle/graalpython

Tutte le issue di oracle/graalpython

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.