Running local LLMs
まだ誰も着手していません。
評価
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 初心者へのやさしさ
- 64/100
- issue の種類
- ドキュメント
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- ollama
- 領域
- ai, documentation
調査の方向性
既存のissue本文と、そこからリンクされているアーキテクチャ図を起点にし、その後、Ollama、vLLM、SGLangがどのように説明されているかを確認します。注記で求められている、より深い洞察を概要に追加し、ローカルLLMエンジンを選ぼうとしている人にも比較が理解しやすい状態を維持します。現在の大まかな概要を超えて、文書化されている違いと技術が拡充されていれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Brief overview of Ollama, vLLM and SGLang
To use open-weight models on your machine, you have three main options: Ollama, vLLM, and SGLang.
Each engine handles requests differently. The diagram below shows the differences and the main techniques behind each engine.
-
Ollama: Ollama is best for local dev, prototyping, and laptop-scale hardware. The architecture is inherently sequential. A local user calls the OpenAI-compatible API, and requests line up in a FIFO queue. Then Ollama runs a pre-quantized GGUF model, a compressed format it pulls, and the response comes back to the user.
-
vLLM: vLLM is best for high-traffic serving, max GPU utilization, and thousands of concurrent requests. Many users hit the server at once, and continuous batching slots new requests into the running batch instead of making them wait for it to finish. PagedAttention stores the KV cache, the memory a model keeps for tokens it has already processed. The PagedAttention maps the OS memory pages to vLLM memory blocks.
-
SGLang: SGLang is best for AI agents and tool loops, multi-turn chats, and JSON/regex outputs. The most commun example is when the workflow involves using repeated context, like a static system prompt or a large RAG documents, during a CI workflow. Agents and multi-turn chats send requests whose prompts overlap heavily. A prefix-aware scheduler routes them through the RadixAttention cache, a radix tree that reuses every shared prefix instead of recomputing it.
[!NOTE]
This is for the big picture. Needs to be continued with a bit more in deep insights...
- 主要言語
- JavaScript
- スター
- 291
- フォーク
- 25
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
dwyl/technology-stack のほかの issue
-
Reactorオープンdiscuss elixir technical
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
dwyl/technology-stack#168 · コメント 7 件 ·
-
chore discuss help wanted priority-1 T25m tech-debt technical
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
dwyl/technology-stack#179 ·
-
`Antigravity`オープンAi discuss enhancement starter T25m technical
難易度 3/5 1〜2日 初心者へのやさしさ 35/100
dwyl/technology-stack#178 · コメント 7 件 ·
-
`DESIGN.md`オープンAi discuss enhancement research T1h
難易度 5/5 1週間以上 初心者へのやさしさ 30/100
dwyl/technology-stack#177 · リアクション 1 件 ·
-
discuss enhancement T25m technical
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
dwyl/technology-stack#175 ·
dwyl/technology-stack の issue をすべて見る
似ている issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
QuantEcon/lecture-python-programming#642 ·
メンテナーはふだん 1 日以内に返信
-
Missing repro Platform: Android
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
software-mansion/react-native-reanimated#10816 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
[Suggestion]: Document that useFormStatus works with a preventDefault-ed onSubmit + startTransitionオープンtype: documentation
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 2 日以内に返信
-
area:docs bug triage:confirmed
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
Cotal-AI/Cotal#2875 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
keithamus/css-minify-tests#304 ·
メンテナーはふだん 1 日以内に返信