Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

TokenReader subrange bounds enforced inconsistently: peekToken/peekPreviousTokenKind/backtrackToMarker ignore embedded window

未关闭
#481 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
72/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
typescript
领域
compilers

调研方向

从 tsdoc/src/parser/TokenReader.ts 开始,将 peekToken()、peekPreviousTokenKind() 和 backtrackToMarker() 与受保护的同级方法进行比较。检查嵌入式 reader 的行为,以及 issue 中提到的 NodeParser 调用点。完成标准是:三个方法都遵守 [_readerStartIndex, _readerEndIndex) 窗口,并且通过测试覆盖所选择的输入结束行为。

由索引模型根据 Issue 内容生成。

描述

Summary

TokenReader respects its embedded subrange (_readerStartIndex/_readerEndIndex) in peekTokenKind()/peekTokenAfterKind()/peekTokenAfterAfterKind(), but three sibling methods ignore those bounds: peekToken() has no end guard at all, and peekPreviousTokenKind() / backtrackToMarker() compare against 0 / allow rewinding before the subrange start. This leaks outer tokens into embedded parses and can return undefined where Token is promised.

Location

  • File: tsdoc/src/parser/TokenReader.ts
  • Class: TokenReader
  • Methods:
    • peekToken(): Token — return this.tokens[this._currentIndex]; with no _readerEndIndex check
    • peekPreviousTokenKind() — if (this._currentIndex === 0) instead of === this._readerStartIndex
    • backtrackToMarker(marker) — checks marker > this._currentIndex but not marker < this._readerStartIndex

Contrast with the guarded siblings in the same file:

public peekTokenKind(): TokenKind {
  if (this._currentIndex >= this._readerEndIndex) {
    return TokenKind.EndOfInput;
  }
  ...
}

Problem

  1. peekToken() past end returns undefined: return type claims Token, but past _readerEndIndex the index expression yields undefined. Callers doing peekToken().range / peekToken().kind then throw TypeError: Cannot read properties of undefined instead of getting a clean EndOfInput signal. Current internal hot paths in NodeParser.ts happen to call peekTokenKind() first (e.g. block/inline tag badCharacter paths), which masks the hole, but the public API contract is broken for any external consumer or future internal use.
  2. peekPreviousTokenKind() leaks across embedded start: for an embedded reader starting at index N>0, calling it at the embedded start returns tokens[N-1].kind (outer context) instead of EndOfInput. NodeParser._parseBlockTag and _parseFencedCode switch on this to decide start-of-input behavior (AtSignInWord, CodeFenceOpeningIndent); an embedded @ or fence at the subrange start can therefore be misclassified using the token before the subrange.
  3. backtrackToMarker() can rewind before subrange start: nothing prevents marker < _readerStartIndex, letting a later readToken() consume outer tokens that the embedded reader was explicitly scoped to exclude.

Trigger / Reproduction

Based on static analysis (no execution performed):

  • Construct new TokenReader(parserContext, embeddedSequence) where embeddedSequence.startIndex > 0, advance to the embedded start, and call peekPreviousTokenKind() — returns the outer predecessor kind instead of EndOfInput.
  • Call peekToken() when _currentIndex === _readerEndIndex — returns undefined instead of signaling end (compare peekTokenKind() returning EndOfInput in the same state).
  • Call backtrackToMarker(outerMarker) with outerMarker < _readerStartIndex on an embedded reader — accepted, subsequent reads escape the subrange.

Note: this is a static-analysis finding; I did not execute a parser fixture.

Expected Behavior

All TokenReader accessors/mutators should honor the same [ _readerStartIndex, _readerEndIndex ) window: peekToken() should signal end (or throw a clear parser-bug error like readToken() does) instead of returning undefined; peekPreviousTokenKind() should return EndOfInput at the subrange start; backtrackToMarker() should reject markers below the subrange start.

Actual Behavior

Subrange bounds are enforced inconsistently, so embedded parses can observe outer tokens and end-of-input handling differs by method.

Impact

  • Potential spurious AtSignInWord / fence-indent errors for embedded constructs starting at a subrange boundary.
  • undefined-token TypeErrors for API consumers using peekToken() at end, bypassing the clean EndOfInput protocol every sibling method follows.

Suggested Direction

  • Mirror the existing peekTokenKind() guard in peekToken() (return the EndOfInput token or throw a parser-bug error — maintainer's choice, but document it), change the peekPreviousTokenKind() zero-check to _readerStartIndex, and add a marker < _readerStartIndex rejection in backtrackToMarker(). No grammar changes needed.

Evidence

  • Source via API: tsdoc/src/parser/TokenReader.ts shows guarded peekTokenKind/peekTokenAfterKind/peekTokenAfterAfterKind vs unguarded peekToken/peekPreviousTokenKind/backtrackToMarker; tsdoc/src/parser/NodeParser.ts shows peekPreviousTokenKind gating AtSignInWord/CodeFenceOpeningIndent and embedded TokenReader construction for scoped parses.
  • Duplicate check: issue search for peekToken TokenReader bounds returns total_count: 0, and open issues contain no TokenReader-boundary report — no apparent duplicate. This is a non-security correctness finding, so Microsoft's private-security-disclosure requirement does not apply.

Classification

  • FACT: three TokenReader methods ignore the subrange window that three sibling methods enforce (verified in source via API).
  • INFERENCE: embedded parses can observe out-of-range tokens; end-of-input peekToken() yields undefined.
  • HYPOTHESIS: aligning all methods on the same window fixes the misclassification/TypeError risk with no grammar change.
主要语言
TypeScript
星标
5k
派生
162
平均合并
2 天 4 小时
30 天内合并 PR
12

环境准备

我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

microsoft/tsdoc 的其他 Issue

查看 microsoft/tsdoc 的全部 Issue

相似的 Issue

更多 TypeScript Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。