This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
Repositories
vllm-project Repositories
(42 Stars) (68 Forks) (0 indexierte Issues) (0 offene good first issues)
Fast and memory-efficient exact attention
(124 Stars) (148 Forks) (0 indexierte Issues) (0 offene good first issues)
vllm-project/guidellmPython
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
(1.166 Stars) (156 Forks) (0 indexierte Issues) (0 offene good first issues)
vllm-project/recipesJavaScript
Common recipes to run vLLM
(833 Stars) (292 Forks) (0 indexierte Issues) (0 offene good first issues)
System Level Intelligent Router for Mixture-of-Models at Cloud, Data Center and Edge
(4.293 Stars) (699 Forks) (129 indexierte Issues) (121 offene good first issues)
TPU inference for vLLM, with unified JAX and PyTorch support.
(400 Stars) (276 Forks) (5 indexierte Issues) (5 offene good first issues)
vllm-project/vllmPython
A high-throughput and memory-efficient inference and serving engine for LLMs
(80.034 Stars) (16.816 Forks) (61 indexierte Issues) (44 offene good first issues)
Community maintained hardware plugin for vLLM on Ascend
(2.553 Stars) (1.955 Forks) (66 indexierte Issues) (65 offene good first issues)
vllm-project/vllm-ncclPython
Manages vllm-nccl dependency
(18 Stars) (3 Forks) (0 indexierte Issues) (0 offene good first issues)
vllm-project/vllm-omniPython
A framework for efficient model inference with omni-modality models
(4.990 Stars) (1.067 Forks) (136 indexierte Issues) (128 offene good first issues)
(43 Stars) (101 Forks) (0 indexierte Issues) (0 offene good first issues)