Make zombie loggers logic more robust
维护者通常 5 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 25/100
- Issue 类型
- 缺陷
- 描述清晰度
- 需要澄清
- 活跃度
- 停滞
- 技术栈
- cpp, ios
调研方向
从 lib/api/Logger.cpp 第 948 行附近的 Logger::RecordShutdown 开始,跟踪 FlushAndTeardown 期间使用的僵尸 logger 保护机制。检查 LogManager Initialize/FlushTeardown 和 GetLogger 路径,然后考虑提议的压力测试:在 100,000 次迭代中进行并发日志记录。完成的标准是该竞态不再导致死锁或终止挂起,同时不引入崩溃。
由索引模型根据 Issue 内容生成。
描述
Describe your environment.
This issue is reproducible in one popular app on older models of iOS devices with slower processor.
Steps to reproduce.
Steps:
- application exiting.
- main thread is calling
FlushAndTeardown. - at about the same time another thread is scheduled to perform logging on
ILogger. - both clash with a deadlock in zombie logger protection code in
Logger::RecordShutdown()method.
What is the expected behavior?
Well, it is expected that applications do not abuse the logging API that way.. At the same time we have some protection mechanism in place, to allow the safe use-after-free. Just that protection mechanism is failing at extremely low rate, unique to the concurrent-use-during-free.
What did you expect to see?
I expect:
- the app should avoid doing what it is doing.
- the zombie logger logic MAY be improved to handle this race condition / deadlock in zombie-logger protection code in a better way.
What is the actual behavior?
Deadlock and hang on app termination, hang in the fool-proof code that is supposed to prevent a crash due to use-after-free. As of note, the code very reliably preventing the crash ... by hanging instead. Unfortunately that hang is eventually reported as a crash.
Additional context.
The crash rate right now is extremely low. It does not seem to affect newer devices.
I think we need to add the following stress test:
Initialize/FlushTeardownin a tight loop onLogManagerinstance.- rogue thread(s) attempting to obtain loggers via
GetLoggerand log massive volumes of data
Basic expectation here that the app should not crash after a 100,000 iterations like this. I am not sure if we can use some other fuzzy testing tools to artificially cause the deadlock.
Solution could be to perform timed-wait on mutex here:
https://github.com/microsoft/cpp_client_telemetry/blob/a924650883ecfd44f12dba131ca117f502f372b9/lib/api/Logger.cpp#L948
And when we see that the timeout happened, we return status back, and we avoid doing anything on that ILogger instance - discarding events that are timing out on that path.
- 主要语言
- C
- 星标
- 102
- 派生
- 67
- 平均合并
- 5 天 5 小时
- 30 天内合并 PR
- 8
环境准备
这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
microsoft/cpp_client_telemetry 的其他 Issue
-
C API enhancement
难度 2/5 1-3 小时 新手友好度 62/100
microsoft/cpp_client_telemetry#628 ·
维护者通常 5 天内回复
-
OneDS C++ SDK retries already-ingested iOS events, causing duplicate telemetry records可能已有人在做 @bmehta001 于 6 天前认领。 未关闭bug
难度 5/5 一周以上 新手友好度 35/100
microsoft/cpp_client_telemetry#1542 · 1 条评论 · 已指派 1 人 ·
维护者通常 5 天内回复
-
难度 5/5 一周以上 新手友好度 35/100
microsoft/cpp_client_telemetry#1504 ·
维护者通常 5 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 25/100
microsoft/cpp_client_telemetry#1413 · 2 条评论 ·
维护者通常 5 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 38/100
microsoft/cpp_client_telemetry#1387 ·
维护者通常 5 天内回复
查看 microsoft/cpp_client_telemetry 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 82/100
openwrt/firmware-utils#82 ·
-
area:http-gateway good first issue priority:low type:bug
难度 2/5 1-3 小时 新手友好度 78/100
crazy-goat/php-fpm-ng#870 ·
维护者通常 1 天内回复
-
Discover carries headerEdges that nothing reads since #1914 moved E0507/E0517 to the compiler graph未关闭tech-debt
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 1 天内回复
-
2个显示的问题未关闭
难度 2/5 1-3 小时 新手友好度 62/100
coolsnowwolf/lede#14209 ·
-
难度 1/5 1 小时以内 新手友好度 75/100
polhenarejos/pico-hsm#147 ·