NVIDIA-NeMo/Megatron-Bridge

[model] JoyAI-LLM Flash

開放

#3,189 建立於 2026年4月7日

 (2 則留言) (0 個反應) (0 位負責人)Python (413 個分叉)auto 404
area:modelfeaturehelp wantedtracking

倉庫指標

星標
 (800 顆星)
PR 合併指標
 (PR 指標待抓取)

描述

Hugging Face repository

https://huggingface.co/joyailab

Architecture family

MoE decoder-only (e.g. DeepSeek V2/V3, OLMoE, Qwen3-MoE, MiniMax-M2)

Required deliverables

  • Model providers
  • HF conversion bridge
  • Unit tests (config and bridge)
  • Model conversion functional tests
  • Optimal pretraining recipe
  • Optimal finetuning recipe
  • Recipe unit tests
  • Recipe functional tests
  • End-to-end CI coverage

Owner if known

No response

Extra context

Request to add JoyAI-LLM Flash (48.9B total / 2.7B active) as a supported model in Megatron-Bridge. The architecture is DeepSeek-V3-style (MLA + MoE) and should map closely to the existing DeepSeek-V3 bridge implementation.

Model was trained with Muon optimizer on an extended Megatron-Core (PP=2, EP=8, ZeRO-1) over 20T tokens.

貢獻者指南