Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Inference Operator v3.2.0 ignores l2CacheBackend: redis, hardcodes LMCACHE_REMOTE_URL=sagemaker-hyperpod://...

未关闭
#431 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
52/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
aws, kubernetes, python, redis

调研方向

首先复现 InferenceEndpointConfig,并使用提供的 kubectl jsonpath 命令检查渲染后的 worker pod。跟踪 operator 如何处理 l2CacheBackend、l2CacheLocalUrl 和 worker.environmentVariables;完成的标准是 Redis 生成 LMCACHE_REMOTE_URL=redis://...,worker 报告 LMCache backend 健康,并且 tieredstorage 部署路径已有文档说明或得到澄清。

由索引模型根据 Issue 内容生成。

描述

Environment
  • EKS add-on amazon-sagemaker-hyperpod-inference v1.3.0-eksbuild.1 (latest available in ap-northeast-1)
  • Inference Operator image hyperpod-inference-operator:v3.2.0
  • Worker image lmcache/vllm-openai:v0.4.7 (LMCache 0.4.7, vLLM 0.23.0)
  • Instance type ml.g7e.4xlarge, HyperPod EKS-orchestrated cluster
What happened

Setting kvCacheSpec.l2CacheSpec.l2CacheBackend: redis (with l2CacheLocalUrl: redis://...:6379) in the InferenceEndpointConfig has no effect. The operator instead injects into the worker pod:

LMCACHE_REMOTE_URL=sagemaker-hyperpod://$(NODE_IP):9200
LMCACHE_EXTRA_CONFIG={"sagemaker_hyperpod_shared_memory_name": "ai_toolkit_cache"}

The redis value from the CRD is silently dropped — the operator never emits LMCACHE_REMOTE_URL=redis://.... Attempting to override LMCACHE_REMOTE_URL via worker.environmentVariables does not work either: those LMCACHE_* entries do not appear in the rendered pod (the operator strips/overrides them).

Because the worker is forced onto the sagemaker-hyperpod connector, and that connector requires a host-side ai-toolkit daemon (POSIX shared memory /ai_toolkit_cache + TCP :9200) which is not present on the cluster, LMCache enters degraded mode and the L2 cache never stores anything:

LMCache ERROR: Failed to initialize shared memory: [Errno 22] Invalid argument: '/ai_toolkit_cache'
LMCache WARNING: Health check failed: RemoteBackendHealthCheck(sagemaker-hyperpod://<NODE_IP>:9200)
LMCache WARNING: HealthMonitor: System unhealthy, entering degraded mode
LMCache WARNING: LMCache is unhealthy, skipping store operation
... LMCache hit tokens: 0     (External prefix cache hit rate: 0.0%)
Expected behavior

With l2CacheBackend: redis + l2CacheLocalUrl: redis://...:6379, the worker's LMCache should be configured with LMCACHE_REMOTE_URL=redis://..., connect to the specified Redis endpoint, report healthy, and perform L2 KV-cache store/lookup. redis is listed as a supported L2 backend in the docs (KV cache & intelligent routing), so the operator overriding it contradicts the documentation.

Reproduction
  1. Deploy an InferenceEndpointConfig with:
    kvCacheSpec:
      enableL1Cache: true
      enableL2Cache: true
      l2CacheSpec:
        l2CacheBackend: redis
        l2CacheLocalUrl: redis://<redis-svc>.<ns>.svc.cluster.local:6379
    
  2. (Optionally) add worker.environmentVariables entries for LMCACHE_REMOTE_URL to try to override.
  3. Inspect the rendered worker pod:
    kubectl get pod <worker> -o jsonpath='{range .spec.containers[*].env[*]}{.name}={.value}{"\n"}{end}' | grep LMCACHE
    
  4. Observe LMCACHE_REMOTE_URL=sagemaker-hyperpod://$(NODE_IP):9200 (not redis://...), and the worker logs showing the LMCache unhealthy/degraded loop above.
Additional question (tieredstorage)

Separately: l2CacheBackend: tieredstorage requires the host-side ai-toolkit daemon (shm /ai_toolkit_cache + :9200). On our cluster this daemon is not installed as any DaemonSet, and the add-on configuration schema (aws eks describe-addon-configuration) exposes no toggle for it (only alb, enableCustomServiceAccounts, executionRoleArn, hyperpodClusterArn, jumpstartGatedModelDownloadRoleArn, keda, tlsCertificateS3Bucket). How is tiered storage meant to be enabled — a cluster-creation flag, and can it be enabled on an existing cluster?

References
主要语言
Python
星标
41
派生
95
平均合并
14 小时 48 分钟
30 天内合并 PR
5

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

aws/sagemaker-hyperpod-cli 的其他 Issue

查看 aws/sagemaker-hyperpod-cli 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。