是否有计划优化GPU上的推理加速

Open
#14 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Feature
Clarity
Needs clarification
Activity status
Stale
Tech stack
cpp

Research direction

Start by reviewing the existing GPTQ quantization-layer kernel implementation and GPU inference path; the issue does not name files or tests. Benchmark quantized inference against the current implementation and define done as a measurable GPU speedup rather than a slowdown.

Written by the indexing model from the issue text.

Description

目前社区LLM采用主流GPTQ量化之后,量化层的kernel实现基本是负向优化,是否有计划支持GPU上量化后的模型推理加速。

Dominant language
C++
Stars
752
Forks
94
PR merge metrics
No merged PRs in 30d

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from MegEngine/InferLLM

All issues in MegEngine/InferLLM

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.