Make zombie loggers logic more robust
Mantenedores costumam responder em até 5 dias
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Facilidade para iniciantes
- 25/100
- Tipo de issue
- Bug
- Clareza
- Precisa de esclarecimento
- Status de atividade
- Estagnada
- Stack de tecnologia
- cpp, ios
- Domínio
- mobile-dev, observability-sre
Direção de pesquisa
Comece com Logger::RecordShutdown em lib/api/Logger.cpp por volta da linha 948 e rastreie a proteção contra logger zumbi usada durante FlushAndTeardown. Revise os caminhos LogManager Initialize/FlushTeardown e GetLogger e, em seguida, considere o teste de estresse proposto com logging concorrente por 100.000 iterações. A tarefa estará concluída quando a condição de corrida não causar mais um deadlock ou travamento na finalização, sem introduzir um crash.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Describe your environment.
This issue is reproducible in one popular app on older models of iOS devices with slower processor.
Steps to reproduce.
Steps:
- application exiting.
- main thread is calling
FlushAndTeardown. - at about the same time another thread is scheduled to perform logging on
ILogger. - both clash with a deadlock in zombie logger protection code in
Logger::RecordShutdown()method.
What is the expected behavior?
Well, it is expected that applications do not abuse the logging API that way.. At the same time we have some protection mechanism in place, to allow the safe use-after-free. Just that protection mechanism is failing at extremely low rate, unique to the concurrent-use-during-free.
What did you expect to see?
I expect:
- the app should avoid doing what it is doing.
- the zombie logger logic MAY be improved to handle this race condition / deadlock in zombie-logger protection code in a better way.
What is the actual behavior?
Deadlock and hang on app termination, hang in the fool-proof code that is supposed to prevent a crash due to use-after-free. As of note, the code very reliably preventing the crash ... by hanging instead. Unfortunately that hang is eventually reported as a crash.
Additional context.
The crash rate right now is extremely low. It does not seem to affect newer devices.
I think we need to add the following stress test:
Initialize/FlushTeardownin a tight loop onLogManagerinstance.- rogue thread(s) attempting to obtain loggers via
GetLoggerand log massive volumes of data
Basic expectation here that the app should not crash after a 100,000 iterations like this. I am not sure if we can use some other fuzzy testing tools to artificially cause the deadlock.
Solution could be to perform timed-wait on mutex here:
https://github.com/microsoft/cpp_client_telemetry/blob/a924650883ecfd44f12dba131ca117f502f372b9/lib/api/Logger.cpp#L948
And when we see that the timeout happened, we return status back, and we avoid doing anything on that ILogger instance - discarding events that are timing out on that path.
- Linguagem predominante
- C
- Estrelas
- 102
- Forks
- 67
- Merge médio
- 5d 5h
- PRs com merge (30d)
- 8
Preparar o ambiente
Este projeto não oferece contêiner de desenvolvimento, Dockerfile nem guia de contribuição, então a configuração fica por sua conta: comece pelo README e veja nosso guia da primeira contribuição para os passos gerais.
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de microsoft/cpp_client_telemetry
-
C API enhancement
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 62/100
microsoft/cpp_client_telemetry#628 ·
Mantenedores costumam responder em até 5 dias
-
OneDS C++ SDK retries already-ingested iOS events, causing duplicate telemetry recordsTalvez já em andamento @bmehta001 assumiu há 6 dias. Abertabug
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 35/100
microsoft/cpp_client_telemetry#1542 · 1 comentário · 1 responsável ·
Mantenedores costumam responder em até 5 dias
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 35/100
microsoft/cpp_client_telemetry#1504 ·
Mantenedores costumam responder em até 5 dias
-
Dropping of telemetry eventsAbertabug
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 25/100
microsoft/cpp_client_telemetry#1413 · 2 comentários ·
Mantenedores costumam responder em até 5 dias
-
bug
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 38/100
microsoft/cpp_client_telemetry#1387 ·
Mantenedores costumam responder em até 5 dias
Todas as issues de microsoft/cpp_client_telemetry
Issues semelhantes
-
Discover carries headerEdges that nothing reads since #1914 moved E0507/E0517 to the compiler graphAbertatech-debt
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 84/100
Mantenedores costumam responder em até 1 dia
-
2个显示的问题Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 62/100
coolsnowwolf/lede#14209 ·
-
Cannot run pico-hsm-tool.pyAberta
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 75/100
polhenarejos/pico-hsm#147 ·
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 85/100
OpenPrinting/cups#1746 ·
Mantenedores costumam responder em até 1 dia
-
Status: Opened
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 75/100
Mantenedores costumam responder em até 1 dia