Inference Operator v3.2.0 ignores l2CacheBackend: redis, hardcodes LMCACHE_REMOTE_URL=sagemaker-hyperpod://...
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 52/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- aws, kubernetes, python, redis
- Lĩnh vực
- cloud, databases, infrastructure
Hướng nghiên cứu
Bắt đầu bằng cách tái tạo InferenceEndpointConfig và kiểm tra worker pod đã được render bằng lệnh kubectl jsonpath được cung cấp. Theo dõi cách operator xử lý l2CacheBackend, l2CacheLocalUrl và worker.environmentVariables; được xem là hoàn tất khi Redis tạo ra LMCACHE_REMOTE_URL=redis://..., worker báo cáo một backend LMCache hoạt động tốt và đường dẫn triển khai tieredstorage được ghi lại hoặc làm rõ.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Environment
- EKS add-on
amazon-sagemaker-hyperpod-inferencev1.3.0-eksbuild.1 (latest available inap-northeast-1) - Inference Operator image
hyperpod-inference-operator:v3.2.0 - Worker image
lmcache/vllm-openai:v0.4.7(LMCache 0.4.7, vLLM 0.23.0) - Instance type
ml.g7e.4xlarge, HyperPod EKS-orchestrated cluster
What happened
Setting kvCacheSpec.l2CacheSpec.l2CacheBackend: redis (with l2CacheLocalUrl: redis://...:6379) in the InferenceEndpointConfig has no effect. The operator instead injects into the worker pod:
LMCACHE_REMOTE_URL=sagemaker-hyperpod://$(NODE_IP):9200
LMCACHE_EXTRA_CONFIG={"sagemaker_hyperpod_shared_memory_name": "ai_toolkit_cache"}
The redis value from the CRD is silently dropped — the operator never emits LMCACHE_REMOTE_URL=redis://.... Attempting to override LMCACHE_REMOTE_URL via worker.environmentVariables does not work either: those LMCACHE_* entries do not appear in the rendered pod (the operator strips/overrides them).
Because the worker is forced onto the sagemaker-hyperpod connector, and that connector requires a host-side ai-toolkit daemon (POSIX shared memory /ai_toolkit_cache + TCP :9200) which is not present on the cluster, LMCache enters degraded mode and the L2 cache never stores anything:
LMCache ERROR: Failed to initialize shared memory: [Errno 22] Invalid argument: '/ai_toolkit_cache'
LMCache WARNING: Health check failed: RemoteBackendHealthCheck(sagemaker-hyperpod://<NODE_IP>:9200)
LMCache WARNING: HealthMonitor: System unhealthy, entering degraded mode
LMCache WARNING: LMCache is unhealthy, skipping store operation
... LMCache hit tokens: 0 (External prefix cache hit rate: 0.0%)
Expected behavior
With l2CacheBackend: redis + l2CacheLocalUrl: redis://...:6379, the worker's LMCache should be configured with LMCACHE_REMOTE_URL=redis://..., connect to the specified Redis endpoint, report healthy, and perform L2 KV-cache store/lookup. redis is listed as a supported L2 backend in the docs (KV cache & intelligent routing), so the operator overriding it contradicts the documentation.
Reproduction
- Deploy an
InferenceEndpointConfigwith:kvCacheSpec: enableL1Cache: true enableL2Cache: true l2CacheSpec: l2CacheBackend: redis l2CacheLocalUrl: redis://<redis-svc>.<ns>.svc.cluster.local:6379 - (Optionally) add
worker.environmentVariablesentries forLMCACHE_REMOTE_URLto try to override. - Inspect the rendered worker pod:
kubectl get pod <worker> -o jsonpath='{range .spec.containers[*].env[*]}{.name}={.value}{"\n"}{end}' | grep LMCACHE - Observe
LMCACHE_REMOTE_URL=sagemaker-hyperpod://$(NODE_IP):9200(notredis://...), and the worker logs showing the LMCache unhealthy/degraded loop above.
Additional question (tieredstorage)
Separately: l2CacheBackend: tieredstorage requires the host-side ai-toolkit daemon (shm /ai_toolkit_cache + :9200). On our cluster this daemon is not installed as any DaemonSet, and the add-on configuration schema (aws eks describe-addon-configuration) exposes no toggle for it (only alb, enableCustomServiceAccounts, executionRoleArn, hyperpodClusterArn, jumpstartGatedModelDownloadRoleArn, keda, tlsCertificateS3Bucket). How is tiered storage meant to be enabled — a cluster-creation flag, and can it be enabled on an existing cluster?
References
- LMCache SageMaker HyperPod connector (client expecting a pre-existing daemon): https://github.com/LMCache/LMCache/pull/1937
- LMCache backend docs: https://docs.lmcache.ai/kv_cache/storage_backends/sagemaker_hyperpod.html
- AWS KV cache & routing docs: https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-model-deployment-caching-routing.html
- Ngôn ngữ chính
- Python
- Star
- 41
- Fork
- 95
- Merge trung bình
- 14 giờ 48 phút
- Pull request đã merge (30 ngày)
- 5
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của aws/sagemaker-hyperpod-cli
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
aws/sagemaker-hyperpod-cli#348 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 68/100
aws/sagemaker-hyperpod-cli#314 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 55/100
aws/sagemaker-hyperpod-cli#384 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 45/100
aws/sagemaker-hyperpod-cli#353 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 45/100
aws/sagemaker-hyperpod-cli#352 ·
Tất cả issue của aws/sagemaker-hyperpod-cli
Issue tương tự
-
changelog investigate
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
ramnes/notion-sdk-py#409 ·
-
good first issue help wanted
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
lindicaphxag-tech/kaggle#28 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
BSData/horus-heresy-3rd-edition#3211 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Unreachable-proxy mount test depends on fixed port 9999Có thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mởbug tests
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày