[Feature Request] Support decoupled observer-model inference for model deviation in deepmd/kk
维护者通常 4 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 35/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
调研方向
从 source/lmp/pair_deepmd_kokkos.cpp 开始,重点查看 PairDeepMDKokkos::init_style() 和现有的单模型拒绝逻辑,然后对比 issue 中描述的常规 pair_style deepmd 行为。在确定如何集成 .pt2 模型和 observer 推理之前,阅读与 DeepPotModelDevi 相关的接口以及 #5758。列出的验收标准全部通过即表示完成,包括确定性一致、serial 和 MPI 覆盖,以及单模型行为不变。
由索引模型根据 Issue 内容生成。
描述
Summary
Add model-deviation support to the Kokkos-accelerated LAMMPS pair_style deepmd/kk for compatible edge/graph .pt2 models by separating the driver model from the additional observer models.
Only model 0 should run every MD step and supply the energy, force, and virial used to advance the trajectory. Models 1 through N-1 only need inference on out_freq steps to estimate committee uncertainty. This preserves normal model-deviation semantics while keeping the Kokkos device-resident path for trajectory integration.
Currently, the same configuration is rejected during initialization:
ERROR: pair style deepmd/kk does not support model deviation.
Detailed Description
Current behavior
The regular LAMMPS pair_style deepmd supports an ensemble of models:
pair_style deepmd \
graph.000.pt2 graph.001.pt2 graph.002.pt2 graph.003.pt2 \
out_freq 100 out_file model_devi.out
pair_coeff * * H C N O Cl
Its effective execution model is already decoupled:
- on ordinary MD steps, evaluate only model 0;
- on
out_freqsteps, evaluate all models; - use model 0 output for dynamics;
- use the ensemble outputs only to calculate deviation statistics.
Running the equivalent input with:
lmp -k on g 1 -sf kk -in in.lammps
selects pair_style deepmd/kk and fails in PairDeepMDKokkos::init_style() because source/lmp/pair_deepmd_kokkos.cpp explicitly rejects numb_models != 1:
if (numb_models != 1) {
error->all(
FLERR,
"pair style deepmd/kk does not support model deviation."
);
}
This restriction was introduced together with deepmd/kk in #5758, whose description states that the Kokkos path requires one model.
Why the driver and observers can be decoupled
For a committee of models M0, M1, ..., MN-1:
- M0 is the driver: it runs on every step and its force advances the trajectory.
- M1 ... MN-1 are observers: they run only when
step % out_freq == 0. - Observer outputs never modify the trajectory; they only contribute to force, energy, and virial deviation statistics.
Therefore, supporting model deviation does not require all models to participate in MD integration or to retain all model workspaces simultaneously.
A possible Kokkos execution flow is:
Every MD step:
build the device graph once
evaluate M0
scatter M0 outputs and advance dynamics
On model-deviation steps:
reuse the same device graph
evaluate M1, update online statistics
evaluate M2, update online statistics
...
evaluate MN-1, update online statistics
write model_devi.out
Sequential observer inference would avoid scaling temporary inference workspace with the committee size. Online Welford accumulation could avoid retaining an N_models x N_atoms x 3 force tensor. The persistent state can remain O(N_atoms):
- model-0 force used by dynamics;
- one observer scratch force/virial buffer;
- running mean and M2 accumulators for deviation.
For MPI/domain-decomposed execution, each observer's ghost force contributions should be reverse-communicated to owner atoms before atom-wise force-deviation statistics are updated.
Requested behavior
Allow a multi-model command such as:
pair_style deepmd \
graph.000.pt2 graph.001.pt2 graph.002.pt2 graph.003.pt2 \
out_freq 100 out_file model_devi.out
to run with:
lmp -k on g 1 -sf kk -in in.lammps
while preserving the regular pair_style deepmd semantics:
- model 0 drives the trajectory;
- observer models run only on model-deviation output steps;
model_devi.outreports compatible force/energy/virial statistics;atomic,relative, andrelative_vwork where applicable;- single-model
deepmd/kkbehavior and performance remain unchanged.
Possible implementation direction
PairDeepMDKokkos already builds the device graph and owns device-resident energy, force, and atomic-virial buffers. A possible implementation could:
- expose device-edge/canonical-graph inference for individual models held by
DeepPotModelDevi, or provide a suitable iterator/evaluation API; - validate that every committee model supports the same device graph schema, type map, cutoff, parameter dimensions, edge-vector precision, and communication contract;
- build the Kokkos graph once per timestep and reuse it across the driver and observers;
- evaluate model 0 on every step;
- evaluate models 1 through N-1 sequentially only on
out_freqsteps; - reverse-communicate observer ghost forces/virials before owner-atom statistics are accumulated;
- calculate model-deviation statistics online where possible;
- reuse the regular
pair_style deepmdoutput format and option semantics.
Acceptance criteria
- Two or more compatible graph/edge
.pt2models initialize underpair_style deepmd/kk. - Model 0 alone drives the trajectory.
- Observer models execute only at
out_freqsteps. model_devi.outagrees with regularpair_style deepmdfor a small deterministic system.- Energy, force, global virial, and atom-wise force deviation are covered.
atomic,relative, andrelative_vare supported or clearly diagnosed.- Serial and at least two-rank MPI/domain-decomposition tests pass.
- Observer evaluation does not require inference workspace proportional to the committee size.
- Existing single-model
deepmd/kkbehavior and performance remain unchanged. - Automated Kokkos regression coverage uses at least two
.pt2models.
This is relevant to active-learning workflows because model-deviation exploration conventionally uses four independently trained models. I can provide a four-model compressed DPA4C .pt2 input and help validate the implementation on an NVIDIA H20.
Further Information, Files, and Links
Related work:
deepmd/kkintroduction and explicit one-model limitation: #5758- DPA4C graph/compact
.pt2support: #5972 - DP-GEN request for
.pt2model-deviation deployment: deepmodeling/dpgen#1925 - DP-GEN implementation of
.pt2export and artifact forwarding: deepmodeling/dpgen#1926
Observed environment:
- DeePMD-kit 3.2.0 built from a clean upstream source checkout
- compressed PyTorch-exportable
.pt2DPA4C models - clean upstream LAMMPS
stable_22Jul2025_update2source build - NVIDIA H20
- four independently trained committee models
The explicit deepmd/kk multi-model rejection is also present in current upstream DeePMD-kit source (source/lmp/pair_deepmd_kokkos.cpp) and is not specific to a locally modified LAMMPS build.
- 主要语言
- Python
- 星标
- 2k
- 派生
- 651
- 平均合并
- 3 天 10 小时
- 30 天内合并 PR
- 13
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
deepmodeling/deepmd-kit 的其他 Issue
-
难度 1/5 1 小时以内 新手友好度 88/100
deepmodeling/deepmd-kit#6038 · 2 条评论 ·
维护者通常 4 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 78/100
deepmodeling/deepmd-kit#5991 ·
维护者通常 4 天内回复
-
Docs enhancement
难度 2/5 1-3 小时 新手友好度 68/100
deepmodeling/deepmd-kit#5766 · 1 条评论 ·
维护者通常 4 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 72/100
deepmodeling/deepmd-kit#5689 · 2 条评论 ·
维护者通常 4 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 78/100
deepmodeling/deepmd-kit#5686 · 1 条评论 ·
维护者通常 4 天内回复
查看 deepmodeling/deepmd-kit 的全部 Issue
相似的 Issue
-
bug status/needs-triage
难度 2/5 1-3 小时 新手友好度 86/100
prowler-cloud/prowler#12887 · 1 条评论 ·
维护者通常 1 天内回复
-
area: desktop platform: macos priority: p3 status: ready type: enhancement
难度 1/5 1 小时以内 新手友好度 92/100
use-agent-os/agent-os#3484 ·
维护者通常 2 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 86/100
open-telemetry/opentelemetry-python-contrib#5113 · 2 条评论 · 2 个 reaction ·
维护者通常 1 天内回复
-
external
难度 2/5 1-3 小时 新手友好度 68/100
langchain-ai/docs#6255 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复