Make Cache.Open return io.ReadSeekCloser to support Range requests
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bien especificado
- Estado de actividad
- Tranquilo
- Stack tecnológico
- go
- Área
- api, backend, infrastructure, testing
Línea de trabajo
Comienza con internal/cache/api.go y sigue las implementaciones de Open en disk.go, memory.go, noop.go, s3.go y remote.go; después inspecciona httputil.ServeCacheHit y el código de caché por niveles. Ejecuta primero las pruebas indicadas de cache, S3, remote, ServeCacheHit, tiered y cachetest; se considera terminado cuando las respuestas de rango único y las lecturas con búsqueda funcionan en todos los backends, las lecturas por rango no almacenan objetos parciales y go test ./... y el objetivo de lint del repositorio pasan.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
Change the Cache interface so that Open() returns an io.ReadSeekCloser instead of an io.ReadCloser, in order to support HTTP Range requests when serving cached objects.
For most backends this is trivial. For backends that stream over the network (S3 and the Remote cache client), we introduce a wrapper that supports a single Seek() to set the start offset, followed by purely sequential reads, backed by a range request. The higher-level Range-serving code is written to use exactly this access pattern.
Interface change
In internal/cache/api.go, change:
Open(ctx context.Context, key Key) (io.ReadCloser, http.Header, error)
to return io.ReadSeekCloser. Document that callers serving ranges MUST use a single seek-to-start followed by sequential reads (no seek-to-end probing). Because io.ReadSeekCloser is a superset of io.ReadCloser, all existing sequential consumers (http.Fetch, git snapshot/bundle, gomod cacher, cachetest suite, etc.) continue to compile and work unchanged.
Shared seek helper
Add one reusable "seek-once, lazily open at offset, then sequential" wrapper implementing io.ReadSeekCloser, parameterised by an "open underlying stream at offset" function:
- Holds a pending start offset (default 0).
Seekis only meaningful before the firstRead: it sets the start offset, resolvingio.SeekStart/io.SeekCurrent/io.SeekEndagainst the known object size. After reading begins,Seekreturns an error.- On first
Read, lazily opens the underlying stream at the offset, then reads sequentially. Closetears down the underlying stream.
This helper is shared by the S3 and Remote backends (DRY).
Backend changes
- disk (
disk.go): return*os.Filedirectly — already anio.ReadSeekCloser. Signature only. - memory (
memory.go): wrap the existing*bytes.Reader(already seekable) in a no-op-close wrapper instead ofio.NopCloser. - noop (
noop.go): signature only (always returns a cache miss). - s3 (
s3.go): implement the helper's "open at offset" using the existingparallelGet/GetObjectpath, starting from the seek offset instead of 0. - remote (
remote.go+client/*.go): implement "open at offset" via a rangedGET. This requires:- a new
Range(start)RequestOptionin theclientpackage that setsRange: bytes=start-; client.Openaccepting206 Partial Contentin addition to200 OK.
- a new
Server-side Range support
Add single-range support to httputil.ServeCacheHit (shared by the API handler and the generic caching handler). Because the S3/Remote readers only support seek-to-start (not seek-to-end), parse the Range header manually rather than using http.ServeContent (which probes the end via Seek(0, io.SeekEnd)):
- Use the existing
Content-Lengthheader for the object size (no seek-to-end). - For a satisfiable single range:
Seek(start, io.SeekStart)once, thenio.CopyN, emitting206 Partial Content,Content-Range,Content-Length, andAccept-Ranges: bytes. - For an unsatisfiable range:
416 Range Not SatisfiablewithContent-Range: bytes */size. - No range / full request: behave as today (advertise
Accept-Ranges: bytes). - Preserve existing conditional (
If-Match/If-None-Match) handling.
This change powers both the API endpoint and, transitively, the Remote backend's ranged reads.
Tiered cache behaviour
The tiered backfill must not commit a truncated object when a range request reads only a slice.
- Full sequential read from a higher tier: keep today's free tee-backfill into tier 0 (no extra GET).
- Ranged read from a higher tier (a non-trivial
Seek): abandon the tee (cancel the tier-0 write so the partial entry is discarded) and kick off a singleton full copy — a background, request-independent (context.WithoutCancel) download of the whole object from the hitting tier into tier 0, deduplicated so N concurrent range readers trigger at most one copy. - A
bytes=0-whole-object range (Seek to current position 0 before any read) is treated as a no-op and keeps the cheap tee path.
Mechanics:
backfillReadCloserbecomes seekable and tracks bytes read.Seekto the current position before reading delegates to the source and keeps teeing; any otherSeekcancels the tee, fires the singleton-copy trigger once, then delegates the seek to the source.- Singleton copy dedup lives on
Tieredvia a shared*sync.Mapkeyed bynamespace + "/" + key. SinceTiered.Namespace()returns a fresh value per request, this map (and anamespacefield) must be carried throughNamespace()by pointer so dedup spans requests. - On trigger:
LoadOrStorethe key; if present, no-op. Otherwise spawn a goroutine that re-Opens the object from the hitting tier (full, unseeked read), writes it to tier 0 viaWriteFunc, and deletes the dedup entry on completion. Errors are logged, not returned (best-effort warming).
Consequence: a ranged read against a cold local tier causes two reads from the higher tier (the range plus the deduplicated background full copy). This is the cost of warming tier 0 on range access.
Tests
- S3 seekable reader: seek-then-sequential-read, error on seek-after-read.
ServeCacheHitranges:206+Content-Range,416unsatisfiable,Accept-Rangesadvertised, full request unchanged.- Remote range round-trip (client
Rangeoption +206handling end-to-end). - Tiered: ranged read does not commit a truncated tier-0 entry; ranged read triggers a (deduplicated) full singleton copy that warms tier 0; full read still tees as before.
- Add a Range case to the
cachetestsuite so every backend is exercised.
Validation
justtasks /go test ./...- linters (golangci-lint via the repo's
justtarget)
Out of scope / notes
- Multi-range (
multipart/byteranges) responses are not supported; only single ranges. - The "warm tier 0 on range access" copy is best-effort and fire-and-forget.
- Lenguaje dominante
- Go
- Estrellas
- 41
- Forks
- 14
- Merge medio
- 19 h 28 min
- PR fusionados (30 d)
- 3
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de block/cachew
-
etag-range-followup
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
-
Azure Blob as a storage backendAbierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 30/100
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 38/100
-
Dificultad 4/5 3-5 días Aptitud para principiantes 56/100
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
Todos los issues de block/cachew
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
gruntwork-io/boilerplate#329 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
prime-radiant-inc/evener#3291 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Netcracker/qubership-apihub-backend#582 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Los mantenedores suelen responder en 1 día