radixark/miles

[Feature] Implement "Miles Diffusion Router" for Workload-Aware Rollouts

开放

#541 创建于 2026年2月1日

 (8 条评论) (0 个反应) (1 位负责人)Python (244 个派生)github user discovery
good first issuehelp wanted

仓库指标

星标
 (1,513 个星标)
PR 合并指标
 (平均合并 9天 17小时) (30 天内合并 200 个 PR)

描述

Motivation

To support large-scale RL rollouts and high-throughput generation, we need to implement a dedicated router for the diffusion engine, tentatively named Diffusion Router.

This router will build upon the concepts of Cache-Aware Load Balancing and Data Parallel (DP) Routing used in the SGLang LLM engine. The goal is to implement a "workload-minimal" routing strategy that ensures requests are distributed to the most available or appropriate engine instances while maintaining system health.

Goals

  1. Core Infrastructure: Create a standalone demo/implementation of a router tailored for sglang-diffusion instances.
  2. Interface Support: Implement the following three critical API interfaces:
  • health_check: Monitor the status of downstream diffusion workers.
  • generate: Route generation requests based on current workload/availability.
  • update_weights_from_disk: Interface placeholder (can be a stub for now) to support future dynamic weight updates.
  1. Minimalist Routing: Focus on low-latency, workload-aware distribution to minimize generation bottlenecks during rollouts.

Technical Tasks

  • Study the existing SGLang Router implementation for LLMs (see resources below).
  • Develop the miles-diffusion-router script/module.
  • Implement basic load balancing logic (e.g., Least-Request or Round-Robin as a baseline, moving toward workload-minimal).
  • Create a demo script showing the router coordinating multiple sglang-diffusion backends.
  • Document the setup process and API usage.

Resources

Calling community members interested in distributed systems and RL infrastructure! Help us build the backbone of the Miles rollout system.

贡献者指南