Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

vmm: one-shot mode allocates CIDs outside the IdPool and is invisible to the VMM

未關閉
#997 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
5/5
預估耗時
一週以上
新手友好度
45/100
Issue 類型
缺陷
描述清晰度
基本清楚
活躍度
冷清
技術堆疊
rust

研究方向

從 vmm/src/one_shot.rs 開始,特別檢視 run_one_shot 和 CID 掃描,然後追蹤 app.rs reload_vms/reload_vms_sync 以及 crates/dstackup/src/cid.rs。比較提出的協調方案,並確定應如何在現有的配置機制中表示 one-shot CID。完成的標準是:並行的 one-shot 配置和主服務配置不能選擇相同的 CID。

由索引模型根據 Issue 內容生成。

描述

Summary

One-shot mode (vmm/src/one_shot.rs) and the main VMM service allocate vsock CIDs from the same configured range through two mechanisms that do not know about each other. Nothing prevents them from picking the same CID.

The two allocators

Main service — IdPool over [cid_start, cid_start + cid_pool_size), rebuilt on reload from the supervisor's process list (app.rs, reload_vms / reload_vms_sync).

One-shot — one_shot.rs:24-56:

// scan `ps aux` for qemu-system-x86_64 ... guest-cid=<n>
let mut one_shot_cid = config.cvm.cid_start;
while existing_cids.contains(&one_shot_cid) {
    one_shot_cid += 1;
    ...
}

It starts at cid_start, avoids collisions by scraping ps aux, and never touches the pool.

Why they can collide

One-shot launches QEMU directly (cmd.status() at the end of run_one_shot) rather than registering the process with the supervisor. occupied_cids in both reload paths is built from supervisor.list(), so a one-shot VM's CID is invisible to the main service and never gets occupied in the pool.

The blindness is one-directional:

sees the other's CIDs? via
one-shot → main service yes ps aux finds the qemu processes
main service → one-shot no one-shot never reaches the supervisor

So the main service can allocate a CID that a running one-shot VM already holds.

Two secondary issues in the same code path:

  • TOCTOU — the ps aux scan and the QEMU launch are not atomic; a concurrent allocation in the window collides regardless.
  • Parsing — CIDs are recovered by string-splitting ps aux output on guest-cid=, which is sensitive to how QEMU arguments are formatted.

Note on #907

Before #907, IdPool::allocate() had an off-by-one that made it skip cid_start entirely, while one-shot starts at cid_start. That incidentally kept the two apart. #907 fixed the off-by-one (correctly — one_shot.rs:48 and crates/dstackup/src/cid.rs both already treat the window as [start, start+size)), which removes the accidental separation.

This is not a regression introduced by #907. The protection only ever held for exactly one one-shot VM: a second one takes cid_start + 1, which was already inside the main pool's allocation range. The underlying problem is that the two allocators were never coordinated.

Possible directions

  1. Register one-shot processes with the supervisor so the existing pool machinery covers them.
  2. Reserve a dedicated range for one-shot outside [cid_start, cid_start + cid_pool_size).
  3. Have one-shot allocate through IdPool rather than ps aux.

(1) seems most consistent with how the rest of the system tracks VMs, but one-shot is deliberately a lighter path, so (2) may be the cheaper fix.

Confidence

The code paths are confirmed by reading: one-shot starts at cid_start, does not register with the supervisor, and both reload paths source occupied_cids from supervisor.list() only. Not verified on hardware — I have not observed an actual vsock CID collision, and I do not know how much one-shot mode is used in practice, which bounds how much this matters.

Found while reviewing #907.

主要語言
Rust
星號
555
分支
99
平均合併
1 天 3 小時
30 天內合併 PR
195

環境準備

  • 沒有 Dockerfile 或 Docker Compose 檔案
  • 沒有 Pull Request 範本
  • 閱讀貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

Dstack-TEE/dstack 的其他 Issue

查看 Dstack-TEE/dstack 的全部 Issue

相似的 Issue

更多 Rust Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。