Make zombie loggers logic more robust
Maintainer antworten meist innerhalb von 5 Tagen
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Anfängerfreundlichkeit
- 25/100
- Issue-Typ
- Bug
- Klarheit
- Muss geklärt werden
- Aktivitätsstatus
- Veraltet
- Tech-Stack
- cpp, ios
- Bereich
- mobile-dev, observability-sre
Rechercherichtung
Beginne mit Logger::RecordShutdown in lib/api/Logger.cpp um Zeile 948 und verfolge den Zombie-Logger-Schutz, der während FlushAndTeardown verwendet wird. Untersuche die Pfade LogManager Initialize/FlushTeardown und GetLogger und ziehe anschließend den vorgeschlagenen Stresstest mit nebenläufiger Protokollierung über 100.000 Iterationen in Betracht. Als erledigt gilt die Aufgabe, wenn die Race Condition keinen Deadlock oder Hänger bei der Beendigung mehr verursacht, ohne einen Absturz einzuführen.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Describe your environment.
This issue is reproducible in one popular app on older models of iOS devices with slower processor.
Steps to reproduce.
Steps:
- application exiting.
- main thread is calling
FlushAndTeardown. - at about the same time another thread is scheduled to perform logging on
ILogger. - both clash with a deadlock in zombie logger protection code in
Logger::RecordShutdown()method.
What is the expected behavior?
Well, it is expected that applications do not abuse the logging API that way.. At the same time we have some protection mechanism in place, to allow the safe use-after-free. Just that protection mechanism is failing at extremely low rate, unique to the concurrent-use-during-free.
What did you expect to see?
I expect:
- the app should avoid doing what it is doing.
- the zombie logger logic MAY be improved to handle this race condition / deadlock in zombie-logger protection code in a better way.
What is the actual behavior?
Deadlock and hang on app termination, hang in the fool-proof code that is supposed to prevent a crash due to use-after-free. As of note, the code very reliably preventing the crash ... by hanging instead. Unfortunately that hang is eventually reported as a crash.
Additional context.
The crash rate right now is extremely low. It does not seem to affect newer devices.
I think we need to add the following stress test:
Initialize/FlushTeardownin a tight loop onLogManagerinstance.- rogue thread(s) attempting to obtain loggers via
GetLoggerand log massive volumes of data
Basic expectation here that the app should not crash after a 100,000 iterations like this. I am not sure if we can use some other fuzzy testing tools to artificially cause the deadlock.
Solution could be to perform timed-wait on mutex here:
https://github.com/microsoft/cpp_client_telemetry/blob/a924650883ecfd44f12dba131ca117f502f372b9/lib/api/Logger.cpp#L948
And when we see that the timeout happened, we return status back, and we avoid doing anything on that ILogger instance - discarding events that are timing out on that path.
- Vorherrschende Sprache
- C
- Sterne
- 102
- Forks
- 67
- Ø Merge
- 5 T. 5 Std.
- Gemergte PRs (30 T.)
- 8
Entwicklungsumgebung
Dieses Projekt bietet weder Dev-Container noch Dockerfile noch Beitragsleitfaden – die Einrichtung liegt bei Ihnen. Beginnen Sie mit der README; die allgemeinen Schritte stehen in unserem Leitfaden für den ersten Beitrag.
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus microsoft/cpp_client_telemetry
-
C API enhancement
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 62/100
microsoft/cpp_client_telemetry#628 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
OneDS C++ SDK retries already-ingested iOS events, causing duplicate telemetry recordsEvtl. vergeben @bmehta001 hat das vor 5 Tagen übernommen. Offenbug
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
microsoft/cpp_client_telemetry#1542 · 1 Kommentar · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 5 Tagen
-
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
microsoft/cpp_client_telemetry#1504 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
bug
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 25/100
microsoft/cpp_client_telemetry#1413 · 2 Kommentare ·
Maintainer antworten meist innerhalb von 5 Tagen
-
bug
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 38/100
microsoft/cpp_client_telemetry#1387 ·
Maintainer antworten meist innerhalb von 5 Tagen
Alle Issues in microsoft/cpp_client_telemetry
Ähnliche Issues
-
area:http-gateway good first issue priority:low type:docs
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
crazy-goat/php-fpm-ng#828 ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:docs good first issue type:docs
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 92/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
-
CVE-2026-18839 popt: size_t underflow in `singleOptionHelp`Evtl. vergeben @pmatilai hat das vor 34 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
rpm-software-management/popt#143 ·
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 85/100