Inference Operator v3.2.0 ignores l2CacheBackend: redis, hardcodes LMCACHE_REMOTE_URL=sagemaker-hyperpod://...
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 52/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- aws, kubernetes, python, redis
- 领域
- cloud, databases, infrastructure
调研方向
首先复现 InferenceEndpointConfig,并使用提供的 kubectl jsonpath 命令检查渲染后的 worker pod。跟踪 operator 如何处理 l2CacheBackend、l2CacheLocalUrl 和 worker.environmentVariables;完成的标准是 Redis 生成 LMCACHE_REMOTE_URL=redis://...,worker 报告 LMCache backend 健康,并且 tieredstorage 部署路径已有文档说明或得到澄清。
由索引模型根据 Issue 内容生成。
描述
Environment
- EKS add-on
amazon-sagemaker-hyperpod-inferencev1.3.0-eksbuild.1 (latest available inap-northeast-1) - Inference Operator image
hyperpod-inference-operator:v3.2.0 - Worker image
lmcache/vllm-openai:v0.4.7(LMCache 0.4.7, vLLM 0.23.0) - Instance type
ml.g7e.4xlarge, HyperPod EKS-orchestrated cluster
What happened
Setting kvCacheSpec.l2CacheSpec.l2CacheBackend: redis (with l2CacheLocalUrl: redis://...:6379) in the InferenceEndpointConfig has no effect. The operator instead injects into the worker pod:
LMCACHE_REMOTE_URL=sagemaker-hyperpod://$(NODE_IP):9200
LMCACHE_EXTRA_CONFIG={"sagemaker_hyperpod_shared_memory_name": "ai_toolkit_cache"}
The redis value from the CRD is silently dropped — the operator never emits LMCACHE_REMOTE_URL=redis://.... Attempting to override LMCACHE_REMOTE_URL via worker.environmentVariables does not work either: those LMCACHE_* entries do not appear in the rendered pod (the operator strips/overrides them).
Because the worker is forced onto the sagemaker-hyperpod connector, and that connector requires a host-side ai-toolkit daemon (POSIX shared memory /ai_toolkit_cache + TCP :9200) which is not present on the cluster, LMCache enters degraded mode and the L2 cache never stores anything:
LMCache ERROR: Failed to initialize shared memory: [Errno 22] Invalid argument: '/ai_toolkit_cache'
LMCache WARNING: Health check failed: RemoteBackendHealthCheck(sagemaker-hyperpod://<NODE_IP>:9200)
LMCache WARNING: HealthMonitor: System unhealthy, entering degraded mode
LMCache WARNING: LMCache is unhealthy, skipping store operation
... LMCache hit tokens: 0 (External prefix cache hit rate: 0.0%)
Expected behavior
With l2CacheBackend: redis + l2CacheLocalUrl: redis://...:6379, the worker's LMCache should be configured with LMCACHE_REMOTE_URL=redis://..., connect to the specified Redis endpoint, report healthy, and perform L2 KV-cache store/lookup. redis is listed as a supported L2 backend in the docs (KV cache & intelligent routing), so the operator overriding it contradicts the documentation.
Reproduction
- Deploy an
InferenceEndpointConfigwith:kvCacheSpec: enableL1Cache: true enableL2Cache: true l2CacheSpec: l2CacheBackend: redis l2CacheLocalUrl: redis://<redis-svc>.<ns>.svc.cluster.local:6379 - (Optionally) add
worker.environmentVariablesentries forLMCACHE_REMOTE_URLto try to override. - Inspect the rendered worker pod:
kubectl get pod <worker> -o jsonpath='{range .spec.containers[*].env[*]}{.name}={.value}{"\n"}{end}' | grep LMCACHE - Observe
LMCACHE_REMOTE_URL=sagemaker-hyperpod://$(NODE_IP):9200(notredis://...), and the worker logs showing the LMCache unhealthy/degraded loop above.
Additional question (tieredstorage)
Separately: l2CacheBackend: tieredstorage requires the host-side ai-toolkit daemon (shm /ai_toolkit_cache + :9200). On our cluster this daemon is not installed as any DaemonSet, and the add-on configuration schema (aws eks describe-addon-configuration) exposes no toggle for it (only alb, enableCustomServiceAccounts, executionRoleArn, hyperpodClusterArn, jumpstartGatedModelDownloadRoleArn, keda, tlsCertificateS3Bucket). How is tiered storage meant to be enabled — a cluster-creation flag, and can it be enabled on an existing cluster?
References
- LMCache SageMaker HyperPod connector (client expecting a pre-existing daemon): https://github.com/LMCache/LMCache/pull/1937
- LMCache backend docs: https://docs.lmcache.ai/kv_cache/storage_backends/sagemaker_hyperpod.html
- AWS KV cache & routing docs: https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-model-deployment-caching-routing.html
- 主要语言
- Python
- 星标
- 41
- 派生
- 95
- 平均合并
- 14 小时 48 分钟
- 30 天内合并 PR
- 5
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
aws/sagemaker-hyperpod-cli 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 62/100
aws/sagemaker-hyperpod-cli#348 ·
-
难度 1/5 1 小时以内 新手友好度 68/100
aws/sagemaker-hyperpod-cli#314 ·
-
难度 3/5 1-2 天 新手友好度 55/100
aws/sagemaker-hyperpod-cli#384 ·
-
难度 3/5 1-2 天 新手友好度 45/100
aws/sagemaker-hyperpod-cli#353 ·
-
难度 3/5 1-2 天 新手友好度 45/100
aws/sagemaker-hyperpod-cli#352 ·
查看 aws/sagemaker-hyperpod-cli 的全部 Issue
相似的 Issue
-
难度 1/5 1-3 小时 新手友好度 85/100
pytest-dev/pluggy#757 ·
维护者通常 1 天内回复
-
难度 1/5 1-3 小时 新手友好度 85/100
NousResearch/hermes-agent#134960 ·
维护者通常 1 天内回复
-
HTML backend: `<br>` leaks the internal sentinel U+E000 into list items, headings and captions可能已有人在做 @morten-lagabote 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 67/100
docling-project/docling#4671 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 70/100
维护者通常 1 天内回复
-
good first issue hacktoberfest infra
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复