Lightning-AI/lightning-thunder
在 GitHub 查看quantization: process tensors on meta device directly, maybe implement CPU quantization (if it is easy)
Open
#1,111 创建于 2024年9月6日
good first issuetransforms
仓库指标
- Star
- (1,460 star)
- PR 合并指标
- (PR 指标待抓取)
描述
Currently the BitsAndBytesLinearQuant4bit for submodule always calls bitsandbytes.functional.quantize_4bit. This is somewhat touchy for CPU tensors because quantize_4bit only works on GPU tensors but it is outright not so nice for meta tensors, where we only would need to get the right shapes.