Load balancing at a connection level
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 25/100
- Issue 类型
- 缺陷
- 描述清晰度
- 需要澄清
- 活跃度
- 停滞
- 技术栈
- kubernetes, python
调研方向
从 Python 入口点 riva.client.ASRService(auth) 以及报告中包含的 Kubernetes Deployment 和 Service 定义开始。确定客户端是否会复用连接,以及这是否解释了观察到的 Pod 之间的不均衡;完成的标准是记录已确认的原因,或记录重现该问题所需但缺失的证据。
由索引模型根据 Issue 内容生成。
描述
I have an application with two pods and some client from inside the same cluster connecting to them by a service, as far as I know this will do a connection level multiplexing.
There is no reason for the workload to be consistently higher at one pod than another, yet I can see one of the pods receiving nearly 3 times more load than the other over a period of 3 hours.
The pod with more load was already running when the other pod started.
My first hypothesis was session stickiness, but a quick test shows that the connections are balanced
for _ in `seq 300` ;
do
curl -b cookies.txt -c cookies.txt -s riva-api.riva:8002/metrics | grep '^nv_gpu_utilization';
sleep 0.1;
done | awk '{print $1}' | sort | uniq -c
My new hypothesis is that python riva client is reusing the connections. Does that make sense or we are guaranteed to start a new connection when calling riva.client.ASRService(auth)?
Here you can find some snippets of the configuration
riva-api (pod) partial definition
apiVersion: apps/v1
kind: Deployment
metadata:
name: riva-api
namespace: riva
labels:
app: riva-api
release: riva-api
spec:
replicas: 2
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
selector:
matchLabels:
app: riva-api
release: riva-api
template:
metadata:
labels:
app: riva-api
release: riva-api
...
spec:
...
containers:
- name: riva-api
image: nvcr.io/nvidia/riva/riva-speech:2.14.0
...
riva-api-online definition
apiVersion: v1
kind: Service
metadata:
name: riva-api
namespace: riva
spec:
ports:
...
selector:
app: riva-api
release: riva-api
- 主要语言
- Python
- 星标
- 142
- 派生
- 52
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
nvidia-riva/python-clients 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 68/100
nvidia-riva/python-clients#172 · 4 条评论 ·
-
难度 3/5 1-2 天 新手友好度 45/100
nvidia-riva/python-clients#145 · 1 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 35/100
nvidia-riva/python-clients#100 · 1 条评论 ·
-
难度 4/5 3-5 天 新手友好度 15/100
-
难度 4/5 3-5 天 新手友好度 25/100
nvidia-riva/python-clients#93 · 1 个 reaction ·
查看 nvidia-riva/python-clients 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 3 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
modelcontextprotocol/python-sdk#3648 ·
维护者通常 1 天内回复
-
docs good first issue
难度 2/5 1-3 小时 新手友好度 78/100
VenetoStato/giorgio#6 ·
-
难度 1/5 1 小时以内 新手友好度 70/100
EclipseFdn/open-vsx.org#13831 ·
维护者通常 1 天内回复
-
feature request
难度 2/5 1-3 小时 新手友好度 68/100
维护者通常 2 天内回复