Make zombie loggers logic more robust
Maintainer thường phản hồi trong vòng 5 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- cpp, ios
- Lĩnh vực
- mobile-dev, observability-sre
Hướng nghiên cứu
Bắt đầu với Logger::RecordShutdown trong lib/api/Logger.cpp quanh dòng 948 và lần theo cơ chế bảo vệ logger zombie được sử dụng trong FlushAndTeardown. Xem xét các đường đi LogManager Initialize/FlushTeardown và GetLogger, sau đó cân nhắc bài kiểm tra stress được đề xuất với việc ghi log đồng thời qua 100.000 lần lặp. Được xem là hoàn tất khi race không còn gây ra deadlock hoặc treo khi kết thúc mà không đưa vào crash.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe your environment.
This issue is reproducible in one popular app on older models of iOS devices with slower processor.
Steps to reproduce.
Steps:
- application exiting.
- main thread is calling
FlushAndTeardown. - at about the same time another thread is scheduled to perform logging on
ILogger. - both clash with a deadlock in zombie logger protection code in
Logger::RecordShutdown()method.
What is the expected behavior?
Well, it is expected that applications do not abuse the logging API that way.. At the same time we have some protection mechanism in place, to allow the safe use-after-free. Just that protection mechanism is failing at extremely low rate, unique to the concurrent-use-during-free.
What did you expect to see?
I expect:
- the app should avoid doing what it is doing.
- the zombie logger logic MAY be improved to handle this race condition / deadlock in zombie-logger protection code in a better way.
What is the actual behavior?
Deadlock and hang on app termination, hang in the fool-proof code that is supposed to prevent a crash due to use-after-free. As of note, the code very reliably preventing the crash ... by hanging instead. Unfortunately that hang is eventually reported as a crash.
Additional context.
The crash rate right now is extremely low. It does not seem to affect newer devices.
I think we need to add the following stress test:
Initialize/FlushTeardownin a tight loop onLogManagerinstance.- rogue thread(s) attempting to obtain loggers via
GetLoggerand log massive volumes of data
Basic expectation here that the app should not crash after a 100,000 iterations like this. I am not sure if we can use some other fuzzy testing tools to artificially cause the deadlock.
Solution could be to perform timed-wait on mutex here:
https://github.com/microsoft/cpp_client_telemetry/blob/a924650883ecfd44f12dba131ca117f502f372b9/lib/api/Logger.cpp#L948
And when we see that the timeout happened, we return status back, and we avoid doing anything on that ILogger instance - discarding events that are timing out on that path.
- Ngôn ngữ chính
- C
- Star
- 102
- Fork
- 67
- Merge trung bình
- 5 ngày 5 giờ
- Pull request đã merge (30 ngày)
- 8
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của microsoft/cpp_client_telemetry
-
C API enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
microsoft/cpp_client_telemetry#628 ·
Maintainer thường phản hồi trong vòng 5 ngày
-
OneDS C++ SDK retries already-ingested iOS events, causing duplicate telemetry recordsCó thể đã có người làm @bmehta001 đã nhận 4 ngày trước. Đang mởbug
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
microsoft/cpp_client_telemetry#1542 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 5 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
microsoft/cpp_client_telemetry#1504 ·
Maintainer thường phản hồi trong vòng 5 ngày
-
Dropping of telemetry eventsĐang mởbug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 25/100
microsoft/cpp_client_telemetry#1413 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 5 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 38/100
microsoft/cpp_client_telemetry#1387 ·
Maintainer thường phản hồi trong vòng 5 ngày
Tất cả issue của microsoft/cpp_client_telemetry
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
libsdl-org/SDL#16444 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug Component component: net
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
RT-Thread/rt-thread#11852 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
MiSTer-devel/ao486_MiSTer#243 ·
-
Dropped last row with parallel scan of attached SQLite tables if the rowid range is a multiple of 122,880Có thể đã có người làm @staticlibs đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
duckdb/duckdb-sqlite#240 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 66/100
siderolabs/pkgs#1710 ·
Maintainer thường phản hồi trong vòng 1 ngày