Issue: Error Loading 8 bit quntized ONNX Model on inference- Protobuf Parsing Failed
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 25/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- python
- Domain
- machine-learning
Research direction
Start with the provided ORTModelForCausalLM.from_pretrained example and the reported protobuf parsing error. Reproduce the failure with the dynamically quantized model and inspect the model-loading path; done means identifying a confirmed cause and documenting or validating a resolution. No repository file or test is named in the issue.
Written by the indexing model from the issue text.
Description
Description
I am facing an error when attempting to load a quantized ONNX model using the ORTModelForCausalLM class from the optimum.onnxruntime library. The error message states: "Failed to load model because protobuf parsing failed."
Context
I am using dynamic quantization for 8-bit quantization on the model before loading it for inference with ONNX.
from optimum.onnxruntime import ORTModelForCausalLM
from transformers import pipeline, LlamaTokenizer
import torch
onnx_path = "./llmonnx/"
opt_model = ORTModelForCausalLM.from_pretrained(onnx_path, file_name="model.onnx").to('cuda')
tokenizer = LlamaTokenizer.from_pretrained(onnx_path)
opt_optimum_generator = pipeline("text-generation", model=opt_model, tokenizer=tokenizer, device='cuda')
prompt = "give me translation for this?"
generated_text = opt_optimum_generator(prompt, max_length=512, num_return_sequences=1, truncation=True)
print(generated_text[0]['generated_text']) please assist me if there is any resolution for this Thanks!
- Dominant language
- C++
- Stars
- 1.7k
- Forks
- 414
- Avg merge
- 15h 1m
- Merged PRs (30d)
- 7
Getting set up
We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/onnxruntime-inference-examples
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
microsoft/onnxruntime-inference-examples#359 ·
Maintainers usually reply within 1 day
-
Minor: typoOpen
Difficulty 1/5 Under an hour Newbie friendliness 72/100
microsoft/onnxruntime-inference-examples#124 · 1 reaction ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 38/100
microsoft/onnxruntime-inference-examples#559 ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 42/100
microsoft/onnxruntime-inference-examples#548 ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
microsoft/onnxruntime-inference-examples#547 ·
Maintainers usually reply within 1 day
All issues in microsoft/onnxruntime-inference-examples
Similar issues
-
area/ysql kind/bug priority/medium
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
yugabyte/yugabyte-db#34552 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 94/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day
-
ai_p2
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
ClickHouse/ClickHouse#123351 ·
Maintainers usually reply within 1 day