Load balancing at a connection level
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 25/100
- issue の種類
- バグ
- 明瞭さ
- 説明が足りない
- 活発さ
- 停滞
- 技術スタック
- kubernetes, python
調査の方向性
Python のエントリーポイント riva.client.ASRService(auth) と、レポートに含まれている Kubernetes Deployment および Service の定義から始めます。クライアントが接続を再利用するかどうか、またそれが観測された Pod 間の偏りを説明するかどうかを判断します。確認済みの原因、または再現に必要な不足している証拠を文書化できれば完了です。
索引モデルが issue の本文から書いたものです。
説明
I have an application with two pods and some client from inside the same cluster connecting to them by a service, as far as I know this will do a connection level multiplexing.
There is no reason for the workload to be consistently higher at one pod than another, yet I can see one of the pods receiving nearly 3 times more load than the other over a period of 3 hours.
The pod with more load was already running when the other pod started.
My first hypothesis was session stickiness, but a quick test shows that the connections are balanced
for _ in `seq 300` ;
do
curl -b cookies.txt -c cookies.txt -s riva-api.riva:8002/metrics | grep '^nv_gpu_utilization';
sleep 0.1;
done | awk '{print $1}' | sort | uniq -c
My new hypothesis is that python riva client is reusing the connections. Does that make sense or we are guaranteed to start a new connection when calling riva.client.ASRService(auth)?
Here you can find some snippets of the configuration
riva-api (pod) partial definition
apiVersion: apps/v1
kind: Deployment
metadata:
name: riva-api
namespace: riva
labels:
app: riva-api
release: riva-api
spec:
replicas: 2
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
selector:
matchLabels:
app: riva-api
release: riva-api
template:
metadata:
labels:
app: riva-api
release: riva-api
...
spec:
...
containers:
- name: riva-api
image: nvcr.io/nvidia/riva/riva-speech:2.14.0
...
riva-api-online definition
apiVersion: v1
kind: Service
metadata:
name: riva-api
namespace: riva
spec:
ports:
...
selector:
app: riva-api
release: riva-api
- 主要言語
- Python
- スター
- 142
- フォーク
- 52
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
nvidia-riva/python-clients のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
nvidia-riva/python-clients#172 · コメント 4 件 ·
-
難易度 3/5 1〜2日 初心者へのやさしさ 45/100
nvidia-riva/python-clients#145 · コメント 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 35/100
nvidia-riva/python-clients#100 · コメント 1 件 ·
-
Connection_Refuseオープン
難易度 4/5 3〜5日 初心者へのやさしさ 15/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
nvidia-riva/python-clients#93 · リアクション 1 件 ·
nvidia-riva/python-clients の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
Juniper/ansible-junos-stdlib#904 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
pollen-robotics/reachy_mini#1457 ·
メンテナーはふだん 1 日以内に返信
-
area:runtime good first issue
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
WATonomous/wato_f1tenth#39 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
FireDynamics/fdsreader#123 ·
-
難易度 1/5 1時間未満 初心者へのやさしさ 85/100