Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Running local LLMs

Đang mở Phù hợp với người mới
#176 1 bình luận 3 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
64/100
Loại issue
Tài liệu
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
ollama
Lĩnh vực
ai, documentation

Hướng nghiên cứu

Bắt đầu với nội dung issue hiện có và sơ đồ kiến trúc được liên kết trong đó, sau đó xem lại cách Ollama, vLLM và SGLang được mô tả. Mở rộng phần tổng quan bằng những hiểu biết sâu hơn được yêu cầu trong ghi chú, đồng thời giữ cho phần so sánh dễ hiểu đối với người đang lựa chọn một engine LLM cục bộ. Công việc được xem là hoàn tất khi các điểm khác biệt và kỹ thuật được tài liệu hóa đã được mở rộng vượt ra ngoài phần tóm tắt tổng quan hiện tại.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Ai discuss enhancement

Brief overview of Ollama, vLLM and SGLang

To use open-weight models on your machine, you have three main options: Ollama, vLLM, and SGLang.
Each engine handles requests differently. The diagram below shows the differences and the main techniques behind each engine.

Image
  • Ollama: Ollama is best for local dev, prototyping, and laptop-scale hardware. The architecture is inherently sequential. A local user calls the OpenAI-compatible API, and requests line up in a FIFO queue. Then Ollama runs a pre-quantized GGUF model, a compressed format it pulls, and the response comes back to the user.

  • vLLM: vLLM is best for high-traffic serving, max GPU utilization, and thousands of concurrent requests. Many users hit the server at once, and continuous batching slots new requests into the running batch instead of making them wait for it to finish. PagedAttention stores the KV cache, the memory a model keeps for tokens it has already processed. The PagedAttention maps the OS memory pages to vLLM memory blocks.

  • SGLang: SGLang is best for AI agents and tool loops, multi-turn chats, and JSON/regex outputs. The most commun example is when the workflow involves using repeated context, like a static system prompt or a large RAG documents, during a CI workflow. Agents and multi-turn chats send requests whose prompts overlap heavily. A prefix-aware scheduler routes them through the RadixAttention cache, a radix tree that reuses every shared prefix instead of recomputing it.

[!NOTE]
This is for the big picture. Needs to be continued with a bit more in deep insights...

Ngôn ngữ chính
JavaScript
Star
291
Fork
25
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của dwyl/technology-stack

Tất cả issue của dwyl/technology-stack

Issue tương tự

Thêm issue về JavaScript

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.