Redline: concurrent calls on one Instance silently return wrong values
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
调研方向
The issue is in the Redline backend's machine state, specifically the argsBuffer, callDepth, TRAP_CODE, and pendingException fields which are shared across threads. Start by examining the machine class and the host-call argument passing mechanism. The goal is to implement a thread-owner check on the outermost call to turn silent corruption into an error, as suggested. Look for the compiled calling convention to understand the context pointer.
由索引模型根据 Issue 内容生成。
描述
Calling one native Instance from several threads returns wrong results with no error. The other backends do not do this: bytecode is correct, and the interpreter throws. Redline is the only one that answers quietly and wrongly.
Eight threads, 2000 calls each, on a shared instance of host-import-roundtrip.wat.wasm, whose callTakeI32 must always return -1:
| Backend | Failed calls | Wrong results |
|---|---|---|
| Interpreter | 16000 / 16000 | 0 |
| Bytecode | 0 / 16000 | 0 |
| Redline (jffi) | 5318 / 16000 | 783 |
Every Redline failure is TrapException: call stack exhausted. The 783 wrong results carried no error at all.
Cause
Per-call state lives in the machine, shared by every thread that enters it, rather than per thread.
argsBufferis one buffer per machine, and host-call arguments pass through it, so one thread reads another's argument.takeI32returns what it was handed, which is where the wrong values come from.callDepthis an unsynchronisedint, sooutermostCallis missed,STACK_LIMITis never re-anchored, and the stack guard fires against a stale limit.TRAP_CODEandpendingExceptionare shared too, so one thread can see and clear another's trap.
Instances are not thread-safe in general and this may simply be unsupported usage. It is filed for the failure mode rather than the usage: silent wrong answers instead of an error.
Suggested fix
Cheap, and worth doing on its own: compare-and-set an owner thread on the outermost call and throw if another thread is already inside. One atomic operation per outermost call, and it turns silent corruption into an actionable error, which is what the interpreter already does in effect.
Real: per-thread context and argument buffers, and a per-thread call depth. Bigger, because the context pointer is part of the compiled calling convention.
Note on #201
The watchdog fix does not cause this, but it does make it visible. Before it, the same probe gave 0 to 3 wrong results rather than 783: starting a thread per call serialised callers enough that most trapped before they could race. The data race is older than the watchdog work; what went away is the throttle that was hiding it.
🤖 Generated with Claude Code
- 主要语言
- Java
- 星标
- 306
- 派生
- 20
- 平均合并
- 1 天 19 小时
- 30 天内合并 PR
- 35
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
bytecodealliance/endive 的其他 Issue
-
难度 4/5 3-5 天 新手友好度 48/100
bytecodealliance/endive#203 ·
维护者通常 1 天内回复
-
难度 5/5 一周以上 新手友好度 25/100
bytecodealliance/endive#202 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 4/5 3-5 天 新手友好度 55/100
bytecodealliance/endive#201 ·
维护者通常 1 天内回复
-
难度 5/5 一周以上 新手友好度 35/100
bytecodealliance/endive#181 · 3 条评论 · 1 个 reaction ·
维护者通常 1 天内回复
-
难度 4/5 3-5 天 新手友好度 52/100
bytecodealliance/endive#175 ·
维护者通常 1 天内回复
查看 bytecodealliance/endive 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 76/100
github/copilot-sdk#2793 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
-
enhancement good first issue
难度 1/5 1-3 小时 新手友好度 88/100