[coverage] Conformance findings: CLOUDFETCH-018
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- nodejs, typescript
Research direction
Start with the two named xfail tests in the coverage PR under tests/ and trace the CloudFetch drain paths for the Thrift and SEA backends, including the retry and link-refresh behavior described here. Done means a stalled download makes at least one cloud request, returns within 150 seconds, and surfaces an error containing timeout or timed out instead of blocking.
Written by the indexing model from the issue text.
Description
Summary
Surfaced by the multi-language coverage fan-out while conformance-testing these SPEC-IDs against databricks/databricks-sql-nodejs. Each finding is committed as an expected-failure (xfail) test in the coverage PR — the test asserts the CORRECT (post-fix) behavior and stays red until THIS driver (databricks/databricks-sql-nodejs) is fixed, then flips green as a tripwire.
Findings
- CLOUDFETCH-018 [thrift]: Thrift CloudFetch downloader enforces no absolute end-to-end per-chunk deadline: with every cloud GET stalled 180s the drain does not return within the 150s budget and no timeout-classified error surfaces (audit finding H05).
- failing test:
CloudFetch — Stalled Download Deadline > stalled chunk download — drain raises a timeout-classified error within an absolute per-chunk deadline [thrift](see the coverage PR diff undertests/)
- failing test:
- CLOUDFETCH-018 [sea]: SEA kernel bounds a stalled CloudFetch chunk only by a static per-request timeout inside a 5-attempt retry loop, so the two multiply instead of bounding the chunk: with every GET stalled 180s the drain does not return within the 150s budget and no timeout-classified error surfaces (audit finding H05).
- failing test:
CloudFetch — Stalled Download Deadline > stalled chunk download — drain raises a timeout-classified error within an absolute per-chunk deadline [sea](see the coverage PR diff undertests/)
- failing test:
- CLOUDFETCH-018: A stalled CloudFetch chunk download is not bounded by any absolute end-to-end per-chunk deadline on either backend: with every cloud GET held 180s the drain does not return within 150s and no error surfaces. The SEA kernel bounds each attempt only by a static per-request timeout under a 5-attempt loop ("Chunk N download failed (attempt 1/5): NetworkError … after 1 attempts … retrying"), so the two multiply instead of bounding the chunk; Thrift shows the same non-termination. Fix: enforce a wall-clock deadline over the whole chunk (connect + response headers + body + retry backoff + link refresh) and surface it as a timeout-classified terminal error (audit finding H05).
Reproduce & Expected
CLOUDFETCH-018 — A CloudFetch chunk download that STALLS -- the cloud-storage GET is accepted but response headers/body never arrive -- must be abandoned under an ABSOLUTE end-to-end wall-clock budget for that chunk,…
Reproduce:
- Stall EVERY CloudFetch download: the proxy accepts each GET and holds it for
180s, so no attempt ever completes and the chunk can only finish by the driver
giving up. - A result large enough to be delivered via CloudFetch external links; drain it
and expect a terminal timeout rather than a multi-minute block.
Expected (per the shared spec):
- full assertion contract:
result:
- label: stalled_drain
exception_thrown: true
- label: stalled_drain
elapsed_seconds_range:
max: 150
- label: stalled_drain
error:
contains:
- timeout
- timed out
- deadline
protocol:
thrift:
- label: stalled_drain
cloud_downloads_min: 1
sea:
- label: stalled_drain
cloud_downloads_min: 1
Context
- The behavior was first fixed in a DIFFERENT driver — reference PR: https://github.com/databricks/databricks-sql-kernel/pull/323 — which seeded the shared language-neutral spec. This issue tracks the same conformance gap in databricks/databricks-sql-nodejs; the reference PR is for cross-referencing the intended behavior, NOT a change to this repo.
- Coverage PR carrying the reproducing xfail test(s): https://github.com/databricks/databricks-driver-test/pull/1577
- Dominant language
- TypeScript
- Stars
- 36
- Forks
- 50
- Avg merge
- 13h 46m
- Merged PRs (30d)
- 9
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from databricks/databricks-sql-nodejs
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
engineer-bot
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
databricks/databricks-sql-nodejs#274 · 1 comment · 1 reaction ·
-
Difficulty 3/5 1-2 days Newbie friendliness 68/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 4/5 3-5 days Newbie friendliness 52/100
All issues in databricks/databricks-sql-nodejs
Similar issues
-
Browser Waiting for: Product Owner
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
getsentry/sentry-javascript#24577 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
agilepathway/label-checker#640 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
copse-dev/agent-pane#2953 ·
-
agentic-workflows
Difficulty 1/5 Under an hour Newbie friendliness 85/100
githubnext/rig#534 ·
-
automation missing-model model-sync provider:pioneer
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
anomalyco/models.dev#7701 ·