Make zombie loggers logic more robust
Los mantenedores suelen responder en 5 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Error
- Claridad
- Necesita aclaración
- Estado de actividad
- Estancado
- Stack tecnológico
- cpp, ios
- Área
- mobile-dev, observability-sre
Línea de trabajo
Comienza con Logger::RecordShutdown en lib/api/Logger.cpp alrededor de la línea 948 y sigue la protección del logger zombi utilizada durante FlushAndTeardown. Revisa las rutas de LogManager Initialize/FlushTeardown y GetLogger, y luego considera la prueba de estrés propuesta con registro concurrente durante 100.000 iteraciones. Se considera terminado cuando la condición de carrera ya no provoca un deadlock ni un bloqueo de la terminación sin introducir un fallo.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Describe your environment.
This issue is reproducible in one popular app on older models of iOS devices with slower processor.
Steps to reproduce.
Steps:
- application exiting.
- main thread is calling
FlushAndTeardown. - at about the same time another thread is scheduled to perform logging on
ILogger. - both clash with a deadlock in zombie logger protection code in
Logger::RecordShutdown()method.
What is the expected behavior?
Well, it is expected that applications do not abuse the logging API that way.. At the same time we have some protection mechanism in place, to allow the safe use-after-free. Just that protection mechanism is failing at extremely low rate, unique to the concurrent-use-during-free.
What did you expect to see?
I expect:
- the app should avoid doing what it is doing.
- the zombie logger logic MAY be improved to handle this race condition / deadlock in zombie-logger protection code in a better way.
What is the actual behavior?
Deadlock and hang on app termination, hang in the fool-proof code that is supposed to prevent a crash due to use-after-free. As of note, the code very reliably preventing the crash ... by hanging instead. Unfortunately that hang is eventually reported as a crash.
Additional context.
The crash rate right now is extremely low. It does not seem to affect newer devices.
I think we need to add the following stress test:
Initialize/FlushTeardownin a tight loop onLogManagerinstance.- rogue thread(s) attempting to obtain loggers via
GetLoggerand log massive volumes of data
Basic expectation here that the app should not crash after a 100,000 iterations like this. I am not sure if we can use some other fuzzy testing tools to artificially cause the deadlock.
Solution could be to perform timed-wait on mutex here:
https://github.com/microsoft/cpp_client_telemetry/blob/a924650883ecfd44f12dba131ca117f502f372b9/lib/api/Logger.cpp#L948
And when we see that the timeout happened, we return status back, and we avoid doing anything on that ILogger instance - discarding events that are timing out on that path.
- Lenguaje dominante
- C
- Estrellas
- 102
- Forks
- 67
- Merge medio
- 5 d 5 h
- PR fusionados (30 d)
- 8
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/cpp_client_telemetry
-
C API enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
microsoft/cpp_client_telemetry#628 ·
Los mantenedores suelen responder en 5 días
-
OneDS C++ SDK retries already-ingested iOS events, causing duplicate telemetry recordsPosiblemente ocupada @bmehta001 la tomó hace 5 días. Abiertobug
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
microsoft/cpp_client_telemetry#1542 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 5 días
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
microsoft/cpp_client_telemetry#1504 ·
Los mantenedores suelen responder en 5 días
-
Dropping of telemetry eventsAbiertobug
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
microsoft/cpp_client_telemetry#1413 · 2 comentarios ·
Los mantenedores suelen responder en 5 días
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 38/100
microsoft/cpp_client_telemetry#1387 ·
Los mantenedores suelen responder en 5 días
Todos los issues de microsoft/cpp_client_telemetry
Issues similares
-
area:http-gateway good first issue priority:low type:docs
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
crazy-goat/php-fpm-ng#828 ·
Los mantenedores suelen responder en 1 día
-
area:docs good first issue type:docs
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
-
CVE-2026-18839 popt: size_t underflow in `singleOptionHelp`Posiblemente ocupada @pmatilai la tomó hace 34 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
rpm-software-management/popt#143 ·
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100