Make Cache.Open return io.ReadSeekCloser to support Range requests
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- 機能追加
- 明瞭さ
- 明確に書かれている
- 活発さ
- 静か
- 技術スタック
- go
- 領域
- api, backend, infrastructure, testing
調査の方向性
internal/cache/api.go から始め、disk.go、memory.go、noop.go、s3.go、remote.go にある Open の実装を追い、その後 httputil.ServeCacheHit と階層型キャッシュのコードを調べます。まず cache、S3、remote、ServeCacheHit、tiered、cachetest の指定されたテストを実行します。完了の条件は、単一範囲のレスポンスと seek 可能な読み取りがすべてのバックエンドで動作し、範囲指定の読み取りで部分オブジェクトが保存されず、go test ./... とリポジトリの lint ターゲットが成功することです。
索引モデルが issue の本文から書いたものです。
説明
Summary
Change the Cache interface so that Open() returns an io.ReadSeekCloser instead of an io.ReadCloser, in order to support HTTP Range requests when serving cached objects.
For most backends this is trivial. For backends that stream over the network (S3 and the Remote cache client), we introduce a wrapper that supports a single Seek() to set the start offset, followed by purely sequential reads, backed by a range request. The higher-level Range-serving code is written to use exactly this access pattern.
Interface change
In internal/cache/api.go, change:
Open(ctx context.Context, key Key) (io.ReadCloser, http.Header, error)
to return io.ReadSeekCloser. Document that callers serving ranges MUST use a single seek-to-start followed by sequential reads (no seek-to-end probing). Because io.ReadSeekCloser is a superset of io.ReadCloser, all existing sequential consumers (http.Fetch, git snapshot/bundle, gomod cacher, cachetest suite, etc.) continue to compile and work unchanged.
Shared seek helper
Add one reusable "seek-once, lazily open at offset, then sequential" wrapper implementing io.ReadSeekCloser, parameterised by an "open underlying stream at offset" function:
- Holds a pending start offset (default 0).
Seekis only meaningful before the firstRead: it sets the start offset, resolvingio.SeekStart/io.SeekCurrent/io.SeekEndagainst the known object size. After reading begins,Seekreturns an error.- On first
Read, lazily opens the underlying stream at the offset, then reads sequentially. Closetears down the underlying stream.
This helper is shared by the S3 and Remote backends (DRY).
Backend changes
- disk (
disk.go): return*os.Filedirectly — already anio.ReadSeekCloser. Signature only. - memory (
memory.go): wrap the existing*bytes.Reader(already seekable) in a no-op-close wrapper instead ofio.NopCloser. - noop (
noop.go): signature only (always returns a cache miss). - s3 (
s3.go): implement the helper's "open at offset" using the existingparallelGet/GetObjectpath, starting from the seek offset instead of 0. - remote (
remote.go+client/*.go): implement "open at offset" via a rangedGET. This requires:- a new
Range(start)RequestOptionin theclientpackage that setsRange: bytes=start-; client.Openaccepting206 Partial Contentin addition to200 OK.
- a new
Server-side Range support
Add single-range support to httputil.ServeCacheHit (shared by the API handler and the generic caching handler). Because the S3/Remote readers only support seek-to-start (not seek-to-end), parse the Range header manually rather than using http.ServeContent (which probes the end via Seek(0, io.SeekEnd)):
- Use the existing
Content-Lengthheader for the object size (no seek-to-end). - For a satisfiable single range:
Seek(start, io.SeekStart)once, thenio.CopyN, emitting206 Partial Content,Content-Range,Content-Length, andAccept-Ranges: bytes. - For an unsatisfiable range:
416 Range Not SatisfiablewithContent-Range: bytes */size. - No range / full request: behave as today (advertise
Accept-Ranges: bytes). - Preserve existing conditional (
If-Match/If-None-Match) handling.
This change powers both the API endpoint and, transitively, the Remote backend's ranged reads.
Tiered cache behaviour
The tiered backfill must not commit a truncated object when a range request reads only a slice.
- Full sequential read from a higher tier: keep today's free tee-backfill into tier 0 (no extra GET).
- Ranged read from a higher tier (a non-trivial
Seek): abandon the tee (cancel the tier-0 write so the partial entry is discarded) and kick off a singleton full copy — a background, request-independent (context.WithoutCancel) download of the whole object from the hitting tier into tier 0, deduplicated so N concurrent range readers trigger at most one copy. - A
bytes=0-whole-object range (Seek to current position 0 before any read) is treated as a no-op and keeps the cheap tee path.
Mechanics:
backfillReadCloserbecomes seekable and tracks bytes read.Seekto the current position before reading delegates to the source and keeps teeing; any otherSeekcancels the tee, fires the singleton-copy trigger once, then delegates the seek to the source.- Singleton copy dedup lives on
Tieredvia a shared*sync.Mapkeyed bynamespace + "/" + key. SinceTiered.Namespace()returns a fresh value per request, this map (and anamespacefield) must be carried throughNamespace()by pointer so dedup spans requests. - On trigger:
LoadOrStorethe key; if present, no-op. Otherwise spawn a goroutine that re-Opens the object from the hitting tier (full, unseeked read), writes it to tier 0 viaWriteFunc, and deletes the dedup entry on completion. Errors are logged, not returned (best-effort warming).
Consequence: a ranged read against a cold local tier causes two reads from the higher tier (the range plus the deduplicated background full copy). This is the cost of warming tier 0 on range access.
Tests
- S3 seekable reader: seek-then-sequential-read, error on seek-after-read.
ServeCacheHitranges:206+Content-Range,416unsatisfiable,Accept-Rangesadvertised, full request unchanged.- Remote range round-trip (client
Rangeoption +206handling end-to-end). - Tiered: ranged read does not commit a truncated tier-0 entry; ranged read triggers a (deduplicated) full singleton copy that warms tier 0; full read still tees as before.
- Add a Range case to the
cachetestsuite so every backend is exercised.
Validation
justtasks /go test ./...- linters (golangci-lint via the repo's
justtarget)
Out of scope / notes
- Multi-range (
multipart/byteranges) responses are not supported; only single ranges. - The "warm tier 0 on range access" copy is best-effort and fire-and-forget.
- 主要言語
- Go
- スター
- 41
- フォーク
- 14
- 平均マージ
- 19時間 28分
- マージ済み PR(30日)
- 3
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
block/cachew のほかの issue
-
etag-range-followup
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 30/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 38/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 56/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
NVIDIA/k8s-device-plugin#2076 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
-
agentic-workflows
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
area/release kind/bug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
kubernetes-sigs/kueue#16455 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信