HttpClient connection pool missing idle timeout
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- java
- Ambito
- backend, networking
Direzione di ricerca
Start with DefaultApacheHttpClient5Factory.java and DefaultHttpClientFactory.java, then review the Apache HttpClient connection-management references and existing configuration patterns. Determine how idle timeout, time-to-live, and validation settings are handled in both factories; done means the Cloud SDK team’s intended behavior and either a safe default or supported application configuration are documented and covered by relevant tests.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
On BTP CF (US30, GCP), we observed 1000+ requests in the last 24 hours taking ~120 seconds across multiple CAP Java services. Each request eventually succeeds after an automatic retry, but the ~120s stall causes significant user-visible latency.
Symptom
A CAP Java service calls an external service via OData. The first attempt fails after ~120 seconds with:
{
"msg": "I/O exception (java.net.SocketException) caught when processing request to {s}->https://<masked>.us30.<masked>.cloud.sap:443: Connection timed out",
"logger": "org.apache.http.impl.execchain.RetryExec",
"level": "INFO",
"written_at": "2026-09-29T15:14:49.995Z"
}
RetryExec then immediately retries with a fresh connection, which succeeds. The target service logs show no trace of the first attempt — the request never arrived.
This pattern is observed across 7–8 different service pairs, all running CAP Java on CF/BTP, all showing the same ~120s delay fingerprint.
Current Assumption (to be confirmed by Cloud SDK team)
This is our working assumption based on log analysis and Cloud SDK Java source code inspection
Network layer hypothesis
CF/BTP egress traffic passes through a NAT gateway. NAT gateways have an idle connection timeout — the exact value likely differs between hyperscalers (AWS, Azure, GCP) and is not publicly documented for BTP. When a pooled TCP connection sits idle past this threshold, the NAT silently drops its flow table entry without sending a TCP RST or FIN. The socket remains in ESTABLISHED state on the Java side (half-open socket).
When RetryExec tries to reuse this stale socket, it calls socket.write() to send the HTTP request. Java's SO_TIMEOUT only applies to socket.read() — there is no application-level timeout on write(). The OS TCP stack retransmits until it exhausts its retry budget (on the observed Diego cells), then raises ETIMEDOUT → SocketException: Connection timed out.
Cloud SDK source — missing idle timeout configuration
Inspecting the Cloud SDK Java source (github.com/SAP/cloud-sdk-java) shows that DefaultApacheHttpClient5Factory only sets a single 2-minute overall timeout, applied to connectTimeout and socketTimeout. The available ConnectionConfig methods for controlling connection pool lifetime are not used:
// DefaultApacheHttpClient5Factory.java — current state
PoolingHttpClientConnectionManagerBuilder.create()
.setDefaultConnectionConfig(ConnectionConfig.custom()
.setConnectTimeout(timeout) // 2 minutes
.setSocketTimeout(timeout) // 2 minutes
// setIdleTimeout(...) ← not set
// setTimeToLive(...) ← not set
// setValidateAfterInactivity(...) ← not set
.build())
...
.build();
Apache HttpClient's default when these are not set is no idle timeout — connections in the pool live indefinitely. Combined with the HttpClient cache TTL of 1 hour (expireAfterAccess), the same connection pool (and its potentially stale connections) is reused for up to 1 hour.
The 2-minute overall timeout is larger than the observed ~120s OS TCP timeout, meaning the OS always fires first and the application-level timeout never has a chance to act.
The same gap exists in DefaultHttpClientFactory (HC4).
Questions
- Can you confirm that
DefaultApacheHttpClient5FactoryandDefaultHttpClientFactory(Cloud SDK Java) intentionally do not setsetIdleTimeout/setTimeToLive/setValidateAfterInactivityon the connection pool? - Do you know the NAT idle timeout values for CF/BTP on GCP, AWS and Azure? This would help determine the right default value.
- Can Cloud SDK either:
- Set a safe default idle timeout that works across all hyperscalers, or
- Expose a configuration property (e.g. via
application.yaml) so applications can set it without overriding internal SDK internals?
A custom workaround at the application level (overriding HttpClientFactory) appears technically possible but has unclear side effects given the complexity of the SDK's connection management — a proper fix at the Cloud SDK level is needed.
References
- Apache HttpClient 5 connection management docs: https://hc.apache.org/httpcomponents-client-5.6.x/connection-management.html
- Cloud SDK Java source
DefaultApacheHttpClient5Factory.java:cloudplatform/connectivity-apache-httpclient5/src/main/java/com/sap/cloud/sdk/cloudplatform/connectivity/DefaultApacheHttpClient5Factory.java - Cloud SDK Java source
DefaultHttpClientFactory.java:cloudplatform/connectivity-apache-httpclient4/src/main/java/com/sap/cloud/sdk/cloudplatform/connectivity/DefaultHttpClientFactory.java
Steps to Reproduce
Since it does not happen on every single request, there is no way for us to reproduce it.
Expected Behavior
Cloud SDK could either:
- Set a safe default idle timeout that works across all hyperscalers, or
- Expose a configuration property (e.g. via
application.yaml) so applications can set it without overriding internal SDK internals?
Used Versions
- Java and Maven version via
mvn --version: 21.0.+ - SAP Cloud SDK version: >5.31
- CAP version: 4.9.4
- Lingua principale
- Java
- Stelle
- 41
- Fork
- 33
- Merge medio
- 22h 16m
- PR unite (30g)
- 19
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di SAP/cloud-sdk-java
-
question
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
SAP/cloud-sdk-java#1301 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
SAP/cloud-sdk-java#1300 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
SAP/cloud-sdk-java#1297 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
SAP/cloud-sdk-java#1291 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
SAP/cloud-sdk-java#1289 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di SAP/cloud-sdk-java
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
-
Make branch and label autocomplete matching locale-independentForse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 83/100
jenkinsci/gitlab-plugin#1950 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
commons-app/apps-android-commons#6984 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
It's not necessary to copy the memory block in the readWrite() of org.h2.store.fs.mem.FileMemDataAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
h2database/h2database#4435 ·
I maintainer di solito rispondono entro 1 giorno