Activity/Workflow TaskProcessor opens a brand-new gRPC channel per task, causing native grpc-core crashes (SIGABRT) under load
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 2/5
- Tiempo estimado
- 1-3 horas
- Aptitud para principiantes
- 82/100
- Tipo de issue
- Error
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Stack tecnológico
- grpc, ruby
- Área
- backend-api-design
Línea de trabajo
Lee lib/temporal/activity/poller.rb alrededor de las líneas 110-114 y lib/temporal/activity/task_processor.rb alrededor de las líneas 71-73; después, sigue el flujo para ver cómo el poller y el task processor obtienen sus conexiones. Se considera terminado cuando los task processors pueden reutilizar la conexión de larga duración del poller sin cambiar el comportamiento existente para los callers que no proporcionen una; ejecuta la suite de tests relevante y verifica que el procesamiento concurrente de tareas siga estando cubierto.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Environment
- temporal-ruby: current
master(b5efd2cef8) - grpc gem: 1.66.0
- Ruby 3.3.6, Rails 7
What happened
Our Rails-based Temporal workers (Temporal::Worker, activity_thread_pool_size: 10, workflow_thread_pool_size: 6) crash intermittently with exit code 134 (SIGABRT) — a native abort inside the grpc gem's C++ core, uncatchable from Ruby. Two different internal failure signatures observed on different occasions:
terminate called after throwing an instance of 'std::logic_error'
what(): basic_string::_S_construct null not valid
terminate called recursively
Aborted (core dumped)
and, separately:
F0000 ... work_stealing_thread_pool.cc:186] Check failed: pool_->IsQuiesced()
*** Check failure stack trace: ***
Aborted (core dumped)
Root cause
Activity::Poller#process(task) / Workflow::Poller#process(task) construct a brand-new TaskProcessor for every polled task:
TaskProcessor#connection memoizes its own Temporal::Connection::GRPC:
So every single activity/workflow task opens a brand-new gRPC channel (fresh DNS resolution, fresh TLS handshake, fresh subchannels/LB policy) and tears it down again right after finishing. Under load (several tasks/sec with a non-trivial activity_thread_pool_size), this produces heavy concurrent gRPC channel churn. We confirmed this directly via GRPC_TRACE=call_error,client_channel: 48 distinct channel handles, 33 creating client_channel / 89 destroying subchannel wrapper lines in a single ~3 minute window on one worker pod.
That churn races grpc-core's internal C++ lifecycle bookkeeping (subchannel refcounting, LB policy teardown, and the EventEngine thread pool's shutdown/quiescence accounting) and trips different fatal internal assertions depending on timing — which is why we saw two different crash signatures for what appears to be the same underlying stressor.
This may also explain, or be related to, #291 ("Unable to poll" / GRPC::Unavailable: Socket closed errors happening frequently under similar thread-pool concurrency), and possibly #280.
Proposed fix
TaskProcessor should reuse the Poller's own long-lived connection instead of building its own per task. gRPC channels are explicitly designed to be shared across concurrent calls, so this is safe even with several TaskProcessors running concurrently on the poller's thread pool — the connection's only Mutex (poll_mutex) guards solely the long-poll bookkeeping (poll_activity_task_queue/poll_workflow_task_queue/cancel_polling_request), not the respond_*_task_completed/respond_*_task_failed calls concurrent task processors make, so no new lock contention is introduced by sharing.
We've deployed this exact fix downstream (as a monkeypatch, since we can't modify the gem source directly in our app) and confirmed 0 crashes over a multi-day soak in an environment that was previously crash-looping every ~10 minutes to a few hours.
Happy to open a PR with this fix — backward-compatible, adds an optional connection: keyword arg to TaskProcessor#initialize defaulting to nil, so existing behavior is unchanged for anyone not passing it. Let me know if that's welcome.
- Lenguaje dominante
- Ruby
- Estrellas
- 288
- Forks
- 113
- Merge medio
- 10 d 15 h
- PR fusionados (30 d)
- 2
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de coinbase/temporal-ruby
-
Dificultad 3/5 1-2 días Aptitud para principiantes 35/100
coinbase/temporal-ruby#341 ·
-
Emitting Metrics for PrometheusAbierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
coinbase/temporal-ruby#328 · 1 comentario ·
-
Dificultad 3/5 1-2 días Aptitud para principiantes 45/100
coinbase/temporal-ruby#326 · 1 comentario ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
coinbase/temporal-ruby#324 · 3 comentarios · 2 reacciones ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 20/100
coinbase/temporal-ruby#322 ·
Todos los issues de coinbase/temporal-ruby
Issues similares
-
P2 testing
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
Los mantenedores suelen responder en 1 día
-
performance
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
Los mantenedores suelen responder en 1 día
-
Blame view link labels are wrongAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
openSUSE/open-build-service#20338 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
DB上でコメント本文がNULLを許容しているAbiertoバグ
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día