[coverage] Conformance findings: CLOUDFETCH-018

Open
#520 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
nodejs, typescript
Domain
backend, databases

Research direction

Start with the two named xfail tests in the coverage PR under tests/ and trace the CloudFetch drain paths for the Thrift and SEA backends, including the retry and link-refresh behavior described here. Done means a stalled download makes at least one cloud request, returns within 150 seconds, and surfaces an error containing timeout or timed out instead of blocking.

Written by the indexing model from the issue text.

Description

Summary

Surfaced by the multi-language coverage fan-out while conformance-testing these SPEC-IDs against databricks/databricks-sql-nodejs. Each finding is committed as an expected-failure (xfail) test in the coverage PR — the test asserts the CORRECT (post-fix) behavior and stays red until THIS driver (databricks/databricks-sql-nodejs) is fixed, then flips green as a tripwire.

Findings

  • CLOUDFETCH-018 [thrift]: Thrift CloudFetch downloader enforces no absolute end-to-end per-chunk deadline: with every cloud GET stalled 180s the drain does not return within the 150s budget and no timeout-classified error surfaces (audit finding H05).
    • failing test: CloudFetch — Stalled Download Deadline > stalled chunk download — drain raises a timeout-classified error within an absolute per-chunk deadline [thrift] (see the coverage PR diff under tests/)
  • CLOUDFETCH-018 [sea]: SEA kernel bounds a stalled CloudFetch chunk only by a static per-request timeout inside a 5-attempt retry loop, so the two multiply instead of bounding the chunk: with every GET stalled 180s the drain does not return within the 150s budget and no timeout-classified error surfaces (audit finding H05).
    • failing test: CloudFetch — Stalled Download Deadline > stalled chunk download — drain raises a timeout-classified error within an absolute per-chunk deadline [sea] (see the coverage PR diff under tests/)
  • CLOUDFETCH-018: A stalled CloudFetch chunk download is not bounded by any absolute end-to-end per-chunk deadline on either backend: with every cloud GET held 180s the drain does not return within 150s and no error surfaces. The SEA kernel bounds each attempt only by a static per-request timeout under a 5-attempt loop ("Chunk N download failed (attempt 1/5): NetworkError … after 1 attempts … retrying"), so the two multiply instead of bounding the chunk; Thrift shows the same non-termination. Fix: enforce a wall-clock deadline over the whole chunk (connect + response headers + body + retry backoff + link refresh) and surface it as a timeout-classified terminal error (audit finding H05).

Reproduce & Expected

CLOUDFETCH-018 — A CloudFetch chunk download that STALLS -- the cloud-storage GET is accepted but response headers/body never arrive -- must be abandoned under an ABSOLUTE end-to-end wall-clock budget for that chunk,…

Reproduce:

  • Stall EVERY CloudFetch download: the proxy accepts each GET and holds it for
    180s, so no attempt ever completes and the chunk can only finish by the driver
    giving up.
  • A result large enough to be delivered via CloudFetch external links; drain it
    and expect a terminal timeout rather than a multi-minute block.

Expected (per the shared spec):

  • full assertion contract:
result:
- label: stalled_drain
  exception_thrown: true
- label: stalled_drain
  elapsed_seconds_range:
    max: 150
- label: stalled_drain
  error:
    contains:
    - timeout
    - timed out
    - deadline
protocol:
  thrift:
  - label: stalled_drain
    cloud_downloads_min: 1
  sea:
  - label: stalled_drain
    cloud_downloads_min: 1

Context

Dominant language
TypeScript
Stars
36
Forks
50
Avg merge
13h 46m
Merged PRs (30d)
9

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from databricks/databricks-sql-nodejs

All issues in databricks/databricks-sql-nodejs

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.