Instrumented probe stores through an unguarded raw pointer into a file-backed mapping — any loss of page backing kills the process with an AccessViolation attributed to innocent user code
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 28/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
- 技术栈
- csharp
- 领域
- testing-qa, tooling
调研方向
首先跟踪 Tracker::Begin 如何从 _view.SafeMemoryMappedViewHandle 初始化,以及在注入的 probe store 执行期间如何保持 MemoryMappedFile 视图存活。运行提供的基于 truncate 的 reproducer,并检查 mapping 的生命周期和文件 backing 是否仍然有效。完成的标准是:即使 backing 丢失,coverage 失败也无法终止或错误归属 test host。
由索引模型根据 Issue 内容生成。
描述
Package: Microsoft.Testing.Extensions.CodeCoverage 18.8.0 (also reproduced on 18.9.0)
Repo: https://github.com/microsoft/codecoverage
Runtime: .NET 8 (8.0.28) and .NET 10 hosts, Linux x64 (Ubuntu container), xunit.v3 3.2.2 on Microsoft.Testing.Platform
Invocation: dotnet test --project <proj> -c Release --coverage --coverage-output-format cobertura
Summary
Static instrumentation prefixes every basic block of every instrumented assembly with an unconditional raw byte store through a static pointer:
ldsfld uint8* Tracker::Begin
ldc.i4 N
add
ldc.i4.1
stind.i1
which the JIT compiles to a single mov byte ptr [reg+idx], 1.
Begin is obtained as
Begin = (byte*)_view.SafeMemoryMappedViewHandle.DangerousGetHandle();
with no DangerousAddRef, over a small file-backed mapping (MemoryMappedFile.CreateFromFile("/tmp/CodeCoverage.<sessionGuid>.<moduleGuid>", FileMode.Open, null, <bufferSize>, MemoryMappedFileAccess.ReadWrite)).
Because the store is unguarded and executes on every basic block, any condition that makes that mapped page unbacked turns the next executed instrumented method into a process-fatal AccessViolationException. The reported stack is whichever user method happened to run at that instant, so the failure is systematically misattributed to innocent user code — including code that provably cannot fault, such as a two-field constructor.
Impact
- A test host dies with
Fatal error. System.AccessViolationException/ SIGABRT (exit 134). - The blamed frame is arbitrary and varies run to run, so users chase phantom concurrency and memory-safety bugs in their own code. We spent a substantial investigation before establishing that the faulting instruction was the injected probe rather than our code.
- On CI this reds unrelated pipelines.
Reproducer (self-contained, no proprietary code)
Reproduced deterministically, 3/3, on a synthetic solution containing only ordinary safe C# (no unsafe, no pointers, no P/Invoke, no interop):
- Create a solution with a class library
LibAand an MTP test project referencing it (xunit.v3 +Microsoft.Testing.Extensions.CodeCoverage,global.jsonwith"test": { "runner": "Microsoft.Testing.Platform" }). - Run
dotnet test --project Tests -c Release --coverage --coverage-output-format coberturaonce and terminate it while the rewritten assemblies are on disk (the run leavesLibA.dllrewritten plusLibA.dll.orig). This step just makes the instrumented artifact available for direct execution. - Read the buffer path baked into the rewritten library:
strings -el LibA.dll | grep '^/tmp/CodeCoverage\.' - Create that file at the buffer size and confirm the probes are live:
dd if=/dev/zero of=$BUF bs=1 count=<bufferSize>then run the test host directly
(dotnet exec Tests.dll) and observe non-zero bytes appear in$BUF. - Run the test host again and, while it is executing, truncate the buffer:
truncate -s 0 $BUF.
Result (3/3):
Fatal error. System.AccessViolationException: Attempted to read or write protected memory.
at BigRepro.LibA.Widget4.Touch(Int32, Int32, Int32, Int64)
at BigRepro.Tests.Suite34.Calc10Works(Int32)
Fatal error. System.AccessViolationException: ...
at BigRepro.LibA.Widget12..ctor(Double, Double, Int64, Boolean)
at BigRepro.Tests.Suite12+<>c.<RejectsBadConfig>b__1_0()
Fatal error. System.AccessViolationException: ...
at BigRepro.LibA.Widget6.Evict(Int32, Int32, Int32)
at BigRepro.Tests.Suite46+<ConcurrentTouchVsEvict>d__0.MoveNext()
Note the second one: the faulting "user code" is a constructor that only validates arguments and assigns fields.
Specificity control: unlinking and recreating the buffer file (new inode) does not crash the host — the existing mapping keeps the old inode alive. Only losing backing for the mapped page faults. So this is specifically about page backing, not about the file being touched.
Evidence from a naturally-occurring crash (not induced)
We independently captured core dumps of the same failure occurring spontaneously in CI-shaped runs (~0.5-1% of invocations). In those dumps:
- SOS
clrstack -fshows aFaultingExceptionFramedirectly above a trivial property getter. - Disassembly at the faulting offset is the probe store, before any user IL executes:
513e: 48 8b 3d cb 4d aa ff mov rdi, [rip-0x55b235] ; Tracker::Begin
5145: b8 c9 00 00 00 mov eax, 0xc9 ; probe index 201
514a: 48 98 cdqe
514c: c6 04 07 01 mov byte ptr [rdi+rax], 1 ; <-- faulting instruction
5150: ... ; user code starts only here
Begin = 0x7824061b5000(non-null), buffer size0x99b(2459), probe index 201 (in range), and the core'sNT_FILEmaps/tmp/CodeCoverage.<session>.<module>at exactly0x7824061b5000. So: valid pointer, in-range index, mapping present — the page was simply not backed at the instant of the store.- The managed heap verifies clean (
verifyheap: 0 errors in one dump; the only "errors" in
others areInvalidMethodTableat a thread'salloc_limit, i.e. allocation-boundary artifacts). So this is not heap corruption by user code. - The process then FailFasts; dumps carry
ExecutionEngineExceptionHResult 0x80131506.
We also observed, with a 5 ms poller during a normal (non-crashing) run, that buffer files are rewritten in place while mapped (size regressions such as 2459 -> 2458 on 2 of 158 buffer files), which suggests these files are not stable for the lifetime of the mapping.
Plausible real-world causes of lost page backing include truncation/rewrite of the buffer file by another participant, and filesystem exhaustion (a dirty page of a file-backed mapping that cannot be allocated yields SIGBUS; CoreCLR's PAL reports both SIGBUS and SIGSEGV as
EXCEPTION_ACCESS_VIOLATION). Our failures cluster on a small, disk-pressured CI runner.
Suggested remedies
- Hold a reference for the lifetime of the pointer — use
DangerousAddRef/DangerousRelease(or keep and use theSafeMemoryMappedViewHandledirectly) so the mapping cannot be released while probes may still execute. - Do not use a shared, world-writable, guessable path.
/tmp/CodeCoverage.<guid>.<guid>is writable by any process on the machine; a truncation by anything else is fatal to unrelated test hosts. Consider an unlinked/anonymous mapping,memfd_create, or a private directory. - Do not let instrumentation faults kill the host. A coverage probe failing should at worst lose coverage data, never terminate the process — and never surface as an
AccessViolationExceptionattributed to user code. Even a one-time validity check with a graceful disable would prevent the misattribution. - If the buffer must be file-backed, preallocate and fsync it, and handle SIGBUS.
Why this is easy to misdiagnose (please consider the diagnostics angle)
Because the injected store is attributed to the enclosing user method, the crash looks exactly like a memory-safety bug in the user's own code. In our case it repeatedly pointed at a lock-protected ConcurrentDictionary wrapper and at trivial constructors. Anything that makes the probe's provenance visible — a marker frame, a distinguishable exception, or a documented signature — would save considerable investigation time.
- 主要语言
- C#
- 星标
- 125
- 派生
- 17
- 平均合并
- 1 小时 15 分钟
- 30 天内合并 PR
- 2
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
microsoft/codecoverage 的其他 Issue
-
难度 3/5 1-2 天 新手友好度 65/100
microsoft/codecoverage#252 ·
-
难度 4/5 3-5 天 新手友好度 52/100
microsoft/codecoverage#251 ·
-
难度 4/5 3-5 天 新手友好度 35/100
microsoft/codecoverage#249 · 1 条评论 ·
-
难度 5/5 一周以上 新手友好度 35/100
microsoft/codecoverage#220 · 2 条评论 ·
-
难度 3/5 1-2 天 新手友好度 66/100
microsoft/codecoverage#218 ·
查看 microsoft/codecoverage 的全部 Issue
相似的 Issue
-
copilot documentation
难度 1/5 1-3 小时 新手友好度 88/100
维护者通常 2 天内回复
-
[Rust][Flaky Test] multiple_deadlines_fire_in_order asserts a wall-clock gap instead of firing order未关闭CI/CD ⚒️ Flaky-tests 🐦
难度 2/5 1-3 小时 新手友好度 88/100
valkey-io/valkey-glide#7255 ·
维护者通常 3 天内回复
-
bug good first issue
难度 1/5 1 小时以内 新手友好度 92/100
unoplatform/Uno.Core#99 ·
-
难度 2/5 1-3 小时 新手友好度 84/100
-
难度 2/5 1-3 小时 新手友好度 86/100
aws/aws-dotnet-ai#75 ·