Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

[FR] Support zoekt indexing of binary files

未關閉
#575 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

  • #7 來自 @sudhanshu112233shukla —— 已關閉,未合併

評估

難度
5/5
預估耗時
一週以上
新手友好度
25/100
Issue 類型
功能
描述清晰度
需要釐清
活躍度
停滯
技術堆疊
go
領域
search

研究方向

Start with IndexBuilder.Add and DocChecker.Check in indexbuilder.go, especially their null-byte checks and skip-marker handling. First determine the intended behavior for PDFs and other binary files; done should include an agreed indexing approach and tests showing the desired result without incorrectly treating binary content as text.

由索引模型根據 Issue 內容生成。

描述

Feature zoekt

When trying to index binary files (like PDFs), they are skipped with the following reason: binary content at byte offset x. Seems like binary files with a byte offset don't play nice with zoekt for whatever reason.


Diagnosing with Ask:

You're absolutely correct! The primary reason PDFs and other binary files are skipped is because they contain null bytes (0x00), not just because of size limits or trigram counts.

The Null Byte Check

In IndexBuilder.Add() indexbuilder.go:428-431:

if idx := bytes.IndexByte(doc.Content, 0); idx >= 0 {
    doc.SkipReason = fmt.Sprintf("binary content at byte offset %d", idx)
    doc.Language = "binary"
}

This checks if there's a null byte anywhere in the file content. If found, it:

  • Sets the SkipReason to indicate where the binary content was detected
  • Marks the language as "binary"

Then the content is replaced indexbuilder.go:433-440:

if doc.SkipReason != "" {
    doc.Content = []byte(notIndexedMarker + doc.SkipReason)
    doc.Symbols = nil
    doc.SymbolsMetaData = nil
    if doc.Language == "" {
        doc.Language = "skipped"
    }
}

The file content is replaced with a marker string like "NOT-INDEXED: binary content at byte offset 5" indexbuilder.go:386.

Similar Check in DocChecker

The DocChecker.Check() method has an identical check indexbuilder.go:544-546:

if index := bytes.IndexByte(content, 0); index > 0 {
    return fmt.Errorf("binary data at byte offset %d", index)
}

Bottom line: PDFs cannot be indexed in Zoekt because they contain null bytes, which are fundamental to binary file formats. The presence of even a single null byte causes the file to be marked as binary and excluded from text-based indexing. There's no configuration option to override this behavior since Zoekt is designed specifically for text-based source code search.

主要語言
TypeScript
星號
3.9k
分支
374
平均合併
2 天 13 小時
30 天內合併 PR
48

環境準備

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

sourcebot-dev/sourcebot 的其他 Issue

查看 sourcebot-dev/sourcebot 的全部 Issue

相似的 Issue

更多 TypeScript Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。