Retry: sleep_for_retry uses max(wait, delay_max) — inflates small Retry-After to delay_max (60s floor)
還沒有人認領這個 Issue。
評估
研究方向
閱讀 src/databricks/sql/auth/retry.py,從 DatabricksRetryPolicy.sleep_for_retry 以及相鄰的 get_backoff_time clamp 開始。重現這個小型 Retry-After 案例,並檢查 tests/test_error_recovery.py,尤其是漸進式 Retry-After 情境。完成標準是:Retry-After 值受到限制,而不是提升到 delay_max,且受影響的重試路徑不再等待 60 秒的下限。
由索引模型根據 Issue 內容生成。
描述
Summary
DatabricksRetryPolicy.sleep_for_retry applies delay_max as a floor instead of a ceiling, so a small server Retry-After (e.g. 2) is inflated to the full delay_max (default 60s) on every retry. This makes retryable errors that advertise a short Retry-After sleep far longer than the server requested.
Location
src/databricks/sql/auth/retry.py, in sleep_for_retry:
retry_after = self.get_retry_after(response)
if retry_after:
proposed_wait = retry_after
else:
proposed_wait = self.get_backoff_time()
proposed_wait = max(proposed_wait, self.delay_max) # <-- BUG: floor, not ceiling
...
time.sleep(proposed_wait)
max(proposed_wait, self.delay_max) guarantees the sleep is at least delay_max. With the default _retry_delay_max = 60, a server response of Retry-After: 2 results in max(2, 60) = 60s.
Why it's a bug
The sibling method get_backoff_time in the same file does the opposite (and correct) clamp, with a docstring that states the intent:
# get_backoff_time():
# "Never returns a value larger than self.delay_max"
proposed_backoff = min(proposed_backoff, self.delay_max)
So delay_max is intended as a ceiling on the wait. sleep_for_retry inverts it. The fix is to cap (not floor) the proposed wait — min(proposed_wait, self.delay_max) — or to not clamp an explicit server Retry-After upward at all.
Impact / repro
A server that returns 503 with a small Retry-After (say 2s, then 4s, then success) is honored as 60s, then 60s — 120s total instead of the intended ~6s.
Observed in the driver-test conformance suite (ERRORRECOV-001, "HTTP 503 with progressive Retry-After"): the request-executing test blocked in retry.py sleep_for_retry -> time.sleep(60) twice and hit the 120s pytest-timeout. faulthandler stack (SEA backend):
tests/test_error_recovery.py:70 cur.execute(SIMPLE_QUERY)
-> databricks/sql/backend/sea/backend.py execute_command
-> .../sea/utils/http_client.py _make_request
-> urllib3 connectionpool.urlopen -> retries.sleep(response)
-> databricks/sql/auth/retry.py:~301 sleep_for_retry -> time.sleep(proposed_wait)
Backend-agnostic: it's in the shared DatabricksRetryPolicy, so both the SEA and Thrift HTTP paths are affected (the kernel path uses a different retry mechanism and is unaffected).
Suggested fix
proposed_wait = min(proposed_wait, self.delay_max)
(and confirm delay_max is the intended upper bound on an honored Retry-After, matching get_backoff_time).
- 主要語言
- Python
- 星號
- 233
- 分支
- 152
- 平均合併
- 21 小時 5 分鐘
- 30 天內合併 PR
- 10
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
databricks/databricks-sql-python 的其他 Issue
-
難度 2/5 1-3 小時 新手友好度 78/100
-
難度 2/5 1-3 小時 新手友好度 76/100
-
難度 2/5 1-3 小時 新手友好度 78/100
-
難度 2/5 1-3 小時 新手友好度 72/100
-
難度 2/5 1-3 小時 新手友好度 84/100
查看 databricks/databricks-sql-python 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 88/100
-
難度 2/5 1-3 小時 新手友好度 82/100
-
難度 2/5 1-3 小時 新手友好度 78/100
-
enhancement
難度 2/5 1-3 小時 新手友好度 72/100
-
難度 2/5 1-3 小時 新手友好度 74/100