Feature Request / Proposal: "E12B" (Effective 12B) Model Class for Gemma 5 (Optimized for 12GB VRAM / Desktop Users)
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- cpp
- Lĩnh vực
- ai, machine-learning, performance
Hướng nghiên cứu
Issue không nêu tên tệp, bài kiểm thử hay entry point nào. Đề xuất này cần có kiến trúc và phạm vi triển khai do maintainer xác định trước khi một contributor có thể xác định một thay đổi cụ thể hoặc một tiêu chí hoàn thành rõ ràng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Hi Gemma Team,
First off, thank you for the incredible work on the Gemma 4 series! The E2B and E4B models have been game-changers for light/mobile workloads.
However, there is a massive gap in the desktop consumer landscape: the 12GB VRAM bracket (e.g., RTX 4070 / 3060 12GB users), which represents a huge portion of local LLM enthusiasts and independent developers.
While standard 12B Dense models run well, they often trade off reasoning capability compared to 27B–35B models. On the other hand, running 27B+ models locally requires offloading or heavy quantization that severely degrades speed.
Proposal for Gemma 5:
Introducing an E12B (Effective 12B) architecture (e.g., a sparse/MoE or active-parameter architecture that yields 27B–35B intelligence while maintaining a 12B active VRAM/RAM footprint).
This would hit the absolute "sweet spot" for desktop hardware, allowing high-speed, high-context generation without forcing users into massive VRAM upgrades during current hardware/RAM market constraints.
Bringing the "Effective" architecture scaling up to the 12B class in Gemma 5 would empower millions of local deployment users.
Thanks for considering this feedback!
- Ngôn ngữ chính
- C++
- Star
- 7k
- Fork
- 660
- Merge trung bình
- 1 ngày 3 giờ
- Pull request đã merge (30 ngày)
- 35
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của google/gemma.cpp
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 68/100
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Tất cả issue của google/gemma.cpp
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
ros2/message_filters#338 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
subsurface/subsurface#4984 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
flutter-webrtc/flutter-webrtc#2206 ·
-
litertlm-android AAR ships no consumer ProGuard rules → "mid == null" SIGABRT in minified apps Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
google-ai-edge/LiteRT-LM#3739 ·