DefaultHttpDestination.equals() change in 5.34.0 causes MT sidecar 408 timeouts under load
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 48/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- java
- Lĩnh vực
- backend, networking
Hướng nghiên cứu
Bắt đầu bằng cách so sánh các thay đổi của DefaultHttpDestination.equals() và hashCode() trong PR #1095 với hành vi của DefaultApacheHttpClient5Cache. Tái hiện lỗi thông qua com.sap.mtx.multitenancy.SubscribeAndUnsubscribeTest.onBoardAndOffBoardNewTenant và truy vết destination được truyền vào ApacheHttpClient5Accessor.getHttpClient(). Hoàn thành khi kịch bản subscribe/unsubscribe MT không còn tạo ra các phản hồi 408 hoặc HTTP 500 dưới tải.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
After upgrading from 5.33.0 to 5.34.0, a Java CAP application that uses the MTX sidecar (@sap/cds-mtxs) for multi-tenancy experiences HTTP 408 Request Timeout responses from the sidecar during tenant subscribe/unsubscribe. The issue is consistently reproducible on a loaded CI environment (Jenkins), though not locally where the sidecar responds fast enough.
Affected versions
- Broken:
5.34.0 - Working:
5.33.0
Root cause analysis
PR #1095 ("Fix unexpected connection-pool shut-down") changed DefaultHttpDestination.equals() and hashCode() to now include customHeaderProviders and headerProvidersFromClassLoading in the comparison (previously these were excluded).
DefaultHttpDestination equality is the cache key for the Cloud SDK HTTP client cache (DefaultApacheHttpClient5Cache). MT scenarios attach per-tenant header providers (auth/token headers) to destinations. Before this change, all tenants pointing at the same sidecar URI shared one HttpClient and one connection pool. After this change, each tenant's destination is a distinct cache key → a new HttpClient + connection pool is created per subscribe/unsubscribe call → rapid connection pool churn against a single sidecar process → the sidecar starts returning 408.
The relevant call chain (in com.sap.cds:cds-feature-mt):
ProvisioningService.subscribe()
→ ServiceCallImpl.execute()
→ HttpClientFactory.getHttpClient(destination) // calls ApacheHttpClient5Accessor
→ sidecar PUT /-/cds/saas-provisioning/tenant/{id}
← HTTP 408
→ InternalError("Unexpected return code 408") // 408 is not in the retry set {500,502,503,504}
→ MtxSidecarDeploymentHandler.onSubscribe() throws
→ HTTP 500 to the subscribe caller
Evidence
Two consecutive CI builds on the same PR branch both failed with the exact same stack:
Caused by: com.sap.cds.feature.mt.lib.subscription.exceptions.InternalError: Unexpected return code 408
at com.sap.cds.feature.mt.lib.subscription.ProvisioningService.lambda$new$1(ProvisioningService.java:80)
...
Caused by: com.sap.cds.feature.mt.lib.subscription.exceptions.InternalError: Unexpected return code 408
at com.sap.cds.feature.mt.lib.subscription.ProvisioningService.lambda$new$1(ProvisioningService.java:80)
Failing test: com.sap.mtx.multitenancy.SubscribeAndUnsubscribeTest.onBoardAndOffBoardNewTenant (expects HTTP 201, gets 500). All other 13 commits in the 5.33.0→5.34.0 range are dependency bumps that are already overridden by the consuming project's own version pins — the only behaviour-changing commit is #1095.
Steps to reproduce
- Run an MTX-sidecar-based CAP Java application's integration tests that subscribe/unsubscribe multiple tenants in rapid succession (e.g. the
mtx-localmodule ofcds-services). - Each subscribe/unsubscribe call hits the sidecar via
ApacheHttpClient5Accessor.getHttpClient(destination)where the destination carries tenant-specific header providers. - With
5.34.0, a new HttpClient + connection pool is allocated per tenant on every call → pool exhaustion / timeout after several tenants → 408 from the sidecar. - With
5.33.0, all same-URI destinations share one HttpClient and pool → no exhaustion → sidecar responds 200/202.
Suggested fix
Options:
- On the cloud-sdk side: consider whether the HTTP-client cache should key on something coarser (e.g. URI only, or URI + a stable identity of the header providers) rather than full header provider equality, especially for the per-request-dynamic providers used in MT scenarios.
- As a workaround: the consuming application can pin
cloud.sdk.version=5.33.0until this is resolved.
- Ngôn ngữ chính
- Java
- Star
- 41
- Fork
- 33
- Merge trung bình
- 1 ngày 57 phút
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của SAP/cloud-sdk-java
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
SAP/cloud-sdk-java#1291 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
SAP/cloud-sdk-java#1290 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
SAP/cloud-sdk-java#1289 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
SAP/cloud-sdk-java#1280 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Newly introduced `CsrfTokenInterceptor` is producing large amount of warning logsCó thể đã có người làm @ricardosrib đã nhận 17 ngày trước. Đang mởbug
SAP/cloud-sdk-java#1270 · 3 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của SAP/cloud-sdk-java
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
apache/rocketmq-dashboard#5358 ·
Maintainer thường phản hồi trong vòng 3 ngày
-
area:cpan-port area:database bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
fglock/PerlOnJava#1605 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
1.0.0-rc2
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
wso2/dpdp-accelerator#377 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area/dependencies backport/26.4 kind/cve severity/high source/scan-dependencies status/triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày