Syscall/capability-scoped restriction policy for process-isolated containers (seccomp/capabilities parity)
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 25/100
- Issue 类型
- 功能
- 描述清晰度
- 需要澄清
- 活跃度
- 活跃
- 技术栈
- docker, kubernetes
调研方向
Start by tracing the HCS container configuration and the containerd runhcs shim, then inspect how CRI-containerd-Windows handles securityContext.seccompProfile and capabilities. Done means a defined policy surface that enforces equivalent restrictions on Windows process-isolated containers or explicitly fails validation instead of silently ignoring them.
由索引模型根据 Issue 内容生成。
描述
Is your feature request related to a problem? Please describe.
Process-isolated (shared-kernel) Windows Server containers have no equivalent to Linux's seccomp-bpf syscall filtering or POSIX capability splitting. Today the only per-container privilege controls are job-object limits and restricted tokens, which are coarse-grained (all-or-nothing Win32 API surface, not "allow these N syscalls / drop these M capabilities"). This creates two concrete problems:
- Kubernetes manifests that set securityContext.seccompProfile are silently ignored on Windows nodes — there's no error, no warning, the field is just a no-op. Teams running mixed-OS clusters can't apply a consistent least-privilege policy across node pools, and nothing surfaces that the policy wasn't actually enforced.
- Because the container shares the host NT kernel with no syscall-level filtering, the effective attack surface of a compromised container process is much larger than a Linux container running under a seccomp default profile, even though both are nominally "shared-kernel" containers. This is the main reason security teams push workloads to Hyper-V isolation by default, which erases most of the density/startup-time benefit of using containers in the first place.
Describe the solution you'd like
Add a policy surface at the server silo boundary that lets a container be started with:
- A syscall/API allow-list or deny-list evaluated at the NT syscall or Win32 API layer for the silo (a seccomp-bpf equivalent), configurable via HCS container config and surfaced through containerd's runhcs shim.
- A capability-style decomposition of the container process token — e.g. explicit grants like "bind to ports < 1024," "create symlinks," "adjust process priority" — instead of the current binary restricted-token/full-token split, so images can run as non-admin without losing the one or two specific privileges they actually need.
- CRI-containerd-Windows support for securityContext.seccompProfile and securityContext.capabilities so existing Kubernetes manifests written for Linux nodes either enforce equivalent policy on Windows nodes or fail validation explicitly, rather than silently doing nothing.
Describe alternatives you've considered
- Hyper-V isolation as the security boundary instead: works today, but reintroduces per-container VM overhead (memory, CPU, startup latency), which defeats the purpose for density-sensitive or high-churn workloads. This is a mitigation, not a fix for shared-kernel containers.
- WDAC/AppLocker policy at the host level: exists today, but applies host-wide or per-image, not per-silo, so it can't express "this specific running container instance gets a tighter policy than its image normally allows."
- Full Linux-style user namespaces (UID remapping/rootless containers): would also improve parity, but is a substantially larger architectural change tied to the NT token model rather than the silo/job-object layer, so we're filing it separately rather than bundling it here.
- Kernel-version decoupling between host and container base image: a related but distinct parity gap (affects portability, not privilege isolation) — also out of scope for this ticket, tracked separately.
Additional context
This is the highest-leverage item in a broader "feature/performance parity with Linux containers" effort we're scoping — namespace depth (PID/user), CimFS/overlay filesystem performance, and cgroups v2 pressure-stall telemetry are related gaps we're filing as separate, narrower tickets rather than one combined issue, since each has a different owner surface (HCS vs. containerd/runhcs vs. kernel team). Happy to file those as follow-ups and cross-link if useful.
- 主要语言
- PowerShell
- 星标
- 551
- 派生
- 76
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
microsoft/Windows-Containers 的其他 Issue
-
enhancement triage
难度 2/5 1-3 小时 新手友好度 68/100
microsoft/Windows-Containers#630 · 6 条评论 ·
-
enhancement triage
难度 5/5 一周以上 新手友好度 25/100
microsoft/Windows-Containers#653 · 1 条评论 ·
-
enhancement triage
难度 5/5 一周以上 新手友好度 30/100
microsoft/Windows-Containers#652 · 1 条评论 ·
-
enhancement triage
难度 5/5 一周以上 新手友好度 30/100
microsoft/Windows-Containers#651 · 1 条评论 ·
-
enhancement triage
难度 5/5 一周以上 新手友好度 25/100
microsoft/Windows-Containers#650 · 2 条评论 ·
查看 microsoft/Windows-Containers 的全部 Issue
相似的 Issue
-
instance instance add
难度 2/5 1-3 小时 新手友好度 68/100
searxng/searx-instances#943 · 1 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 78/100
simp/pupmod-simp-ssh#246 ·
-
难度 1/5 1 小时以内 新手友好度 88/100
simp/pupmod-simp-pupmod#261 ·
-
难度 2/5 1-3 小时 新手友好度 86/100
-
agentic-workflows
难度 2/5 1-3 小时 新手友好度 68/100
github/gh-aw-firewall#9328 ·
维护者通常 1 天内回复