sagemaker-core 2.15.0: role validation raises false-positive `RoleValidationError` under condition-based SCPs (IAM simulator can't evaluate conditional SCPs)
维护者通常 2 天内回复
评估
调研方向
从 sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py 开始,重点查看 resolve_and_validate_role 和 _evaluate_permissions,然后使用提供的 aws iam simulate-principal-policy 命令重现 IAM 模拟器响应。当基于条件的 SCP 结果不再产生错误的 RoleValidationError、真正的权限拒绝仍能得到正确处理,并且现有的 warn-and-proceed 行为得到保留时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
PySDK Version
- PySDK V2 (2.x)
- PySDK V3 (3.x)
Reported against the sagemaker-core distribution, version 2.15.0 (repo tag v3.15.0).
Describe the bug
sagemaker-core 2.15.0 added a client-side permission pre-check that runs during high-level construction (e.g. ModelTrainer(...) → TrainDefaults.get_role) before any training job is submitted: resolve_and_validate_role → _evaluate_permissions → iam:SimulatePrincipalPolicy. It raises RoleValidationError on any non-allowed simulate verdict — and that verdict includes the AWS Organizations / SCP layer (OrganizationsDecisionDetail.AllowedByOrganizations).
Per AWS docs, the IAM policy simulator does not evaluate SCPs that have any conditions. So in an account whose organization uses condition-based SCPs, SimulatePrincipalPolicy returns AllowedByOrganizations: false (with EvalDecision: implicitDeny, MatchedStatements: []) for actions that are actually permitted at run time. The pre-check treats this as a definitive denial and raises — a false positive — even though the execution role is correctly configured and the real API call would succeed.
Two observable consequences:
- Creating a brand-new, fully-permissioned role does not help — the simulate is denied at the org layer regardless of the role's own policies.
- The same role works fine from a notebook / via a direct
create_training_jobcall, because those paths don't run this client-side pre-check.
2.15.0 is currently the latest published release, so there is no fixed version to upgrade to.
To reproduce
Prerequisites: an AWS account under an organization with at least one condition-based SCP; a training execution role that trusts sagemaker.amazonaws.com and grants the training smoke-test actions at Resource: *; a calling identity that can call iam:SimulatePrincipalPolicy.
pip install 'sagemaker-core==2.15.0'
from sagemaker.core.helper.iam_role_resolver import IamRoleResolver
# Also reproducible via ModelTrainer(...) construction with role_arn set to the same role.
IamRoleResolver().resolve_and_validate_role(
role_arn="arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>",
role_type="training",
)
Result:
RoleValidationError: IAM role 'arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role>' cannot be used for 'training' workloads.
Missing permissions: cloudwatch:PutMetricData, ec2:CreateNetworkInterface,
ec2:CreateNetworkInterfacePermission, ec2:DeleteNetworkInterface,
ec2:DeleteNetworkInterfacePermission, ec2:DescribeDhcpOptions, ec2:DescribeNetworkInterfaces,
ec2:DescribeSecurityGroups, ec2:DescribeSubnets, ec2:DescribeVpcs,
ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetAuthorizationToken,
ecr:GetDownloadUrlForLayer
Confirm the verdict is an org-layer artifact rather than a real permission gap:
aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::<ACCOUNT_ID>:role/<training-exec-role> \
--action-names cloudwatch:PutMetricData ec2:CreateNetworkInterface sagemaker:CreateTrainingJob
{
"EvalActionName": "cloudwatch:PutMetricData",
"EvalDecision": "implicitDeny",
"MatchedStatements": [],
"OrganizationsDecisionDetail": { "AllowedByOrganizations": "false" }
// ...same for the other actions, including sagemaker:CreateTrainingJob —
// yet CreateTrainingJob calls from this role succeed at run time (visible in CloudTrail).
}
The identical role runs the same workload successfully from a notebook / via direct API, so the real run-time evaluation permits these actions.
Expected behavior
The pre-check should not hard-fail on an Organizations/SCP-layer denial that the IAM policy simulator cannot faithfully evaluate. Because the simulator ignores condition-based SCPs, an AllowedByOrganizations: false result (with no matched explicit identity Deny) is unverifiable, not authoritative — it should be treated the same as the existing "caller can't call simulate → warn and proceed" path, letting the real API call be the source of truth. An explicit opt-out (e.g. validate_role=False or an env var) would also let users bypass the client-side check without modifying IAM.
Screenshots or logs
.../site-packages/sagemaker/core/helper/iam_role_resolver.py:573 in resolve_and_validate_role
570 # Permission check (definitive denial blocks; unverifiable warns)
571 verdict, denied = _evaluate_permissions(iam_client, role_arn, rol...
572 if verdict is False:
> 573 raise RoleValidationError(
574 _build_validation_error_message(role_arn, role_type, miss...
System information
- SageMaker Python SDK version: sagemaker-core 2.15.0 (repo tag
v3.15.0) - Framework name or algorithm: N/A — fails during role validation, before framework/job selection (framework-agnostic)
- Framework version: N/A
- Python version: 3.12
- CPU or GPU: N/A (fails before job submission)
- Custom Docker image (Y/N): N/A
Additional context
Introduced in sagemaker-core 2.15.0 — file sagemaker-core/src/sagemaker/core/helper/iam_role_resolver.py, added in commit dba1127a ("New release (#5969)"), first tag v3.15.0. Authoring PRs: #2041 (added SimulatePrincipalPolicy-based resolve_or_create_role) → #2080 (replaced it with the raising resolve_and_validate_role). #2080 notes it gates only on *-resource "smoke test" actions to avoid false denials on resource-scoped actions, but does not account for the Organizations/SCP layer the simulate call implicitly evaluates — which is the source of this false positive.
Suggested fix: in _evaluate_permissions, when an action is implicitDeny with no matched identity/SCP statement and OrganizationsDecisionDetail.AllowedByOrganizations == false, treat it as unverifiable (warn + proceed) rather than a missing permission; and/or add an explicit opt-out.
Relevant docs:
- IAM policy simulator can't test SCPs with conditions
- IAM policy evaluation logic (explicit Deny overrides Allow)
Workarounds (both verified):
- Attach an explicit Deny on
iam:SimulatePrincipalPolicyto the identity running the SDK — it then skips the pre-check, warns, and proceeds (an explicit Deny is needed to override any existing Allow). - Pin
sagemaker-core<2.15.0, which predates the pre-check.
- 主要语言
- Python
- 星标
- 2.3k
- 派生
- 1.3k
- 平均合并
- 3 天 10 小时
- 30 天内合并 PR
- 77
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
aws/sagemaker-python-sdk 的其他 Issue
-
Cannot use spark_event_logs_s3_uri in PySparkProcessor job可能已有人在做 @rsareddy0329 于 6 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 78/100
aws/sagemaker-python-sdk#6253 ·
维护者通常 2 天内回复
-
[Bug] V3 Hyperparameter Tuning Pipeline page labelled "Download Data" in navigation due to missing title cell可能已有人在做 @admivsn 于 34 天前认领。 未关闭
难度 1/5 1 小时以内 新手友好度 93/100
aws/sagemaker-python-sdk#6232 ·
维护者通常 2 天内回复
-
[Bug] ModelTrainer with no input channels emits InputDataConfig: [], which CreatePipeline rejects (min=1) — v2 omitted the key可能已有人在做 @sagemaker-bot 于 6 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 76/100
aws/sagemaker-python-sdk#6156 · 2 条评论 ·
维护者通常 2 天内回复
-
sagemaker-train should depend on mlflow-skinny, following sagemaker-mlflow 0.5.0可能已有人在做 @mohamedzeidan2021 于 7 天前认领。 未关闭
难度 2/5 半天 新手友好度 72/100
aws/sagemaker-python-sdk#6152 ·
维护者通常 2 天内回复
-
ModelTrainer generates sm_train.sh with CRLF line endings on Windows causing training job failure可能已有人在做 @MohammedAlkindi 于 25 天前认领。 未关闭
难度 1/5 1 小时以内 新手友好度 88/100
aws/sagemaker-python-sdk#5904 · 1 个 reaction ·
维护者通常 2 天内回复
查看 aws/sagemaker-python-sdk 的全部 Issue
相似的 Issue
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)未关闭
难度 2/5 1-3 小时 新手友好度 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reused可能已有人在做 @sylvesterkaczmarek 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 78/100
google-deepmind/bsuite#56 ·
-
难度 2/5 1-3 小时 新手友好度 82/100
LearningCircuit/local-deep-research#7206 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 68/100
chingu-voyages/V62-tier3-team-33#285 ·
维护者通常 1 天内回复
-
Proxy drops log notifications from backends that don't send FastMCP's msg/extra dict可能已有人在做 @asasemahmed 今天认领。 未关闭bug server
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复