Feature Request: Latin/Spanish OCR model + MeloTTS Spanish for MaixCAM2
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
- 技術スタック
- python
- 領域
- ai, embedded-iot
調査の方向性
まず、既存の pp_ocr_en.mud と melotts-zh.mud のアーティファクト、および AX630 用に確立されている Pulsar2 変換パイプラインを調査します。リンクされているラテン語 PP-OCR モデルとスペイン語 MeloTTS モデルが、これらのパイプラインと互換性があるか確認します。pp_ocr_latin.mud と melotts-es.mud の動作するアーティファクトを MaixCAM2 model zoo に追加できれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Feature Request: Latin/Spanish OCR model + MeloTTS Spanish for MaixCAM2
Platform: MaixCAM2 / MaixPy v4
Context
I am building an assistive vision system for blind and visually impaired users using MaixCAM2. The device uses Depth-Anything-V2 for obstacle detection, YOLO11n for object identification, SmolVLM for scene description, and PP-OCR for reading signs and text. All audio feedback is via MeloTTS. The target users are Spanish speakers (Latin America).
Request 1: Latin/Spanish PP-OCR recognition model
The current pp_ocr_en.mud model fails to correctly recognize Spanish text. It confuses letters like Ñ, accented vowels, and common Spanish character combinations.
There is already a pre-converted ONNX model available at:
https://huggingface.co/docato/PaddleOCR_Mobile_Models — latin_PP-OCRv3_mobile_rec_infer.onnx
This model covers Latin-script languages including Spanish, French, Portuguese, etc. Since you already have the Pulsar2 pipeline set up for PP-OCR models on AX630, compiling this to .axmodel + .mud should follow the same process as the existing English model.
Request: Please provide pp_ocr_latin.mud (or similar) in the MaixCAM2 model zoo, using the Latin PP-OCRv3 recognition model compiled for AX630.
Request 2: MeloTTS Spanish model
The current melotts-zh.mud only supports Chinese and basic English. There is no Spanish TTS option available on MaixCAM2.
The upstream Spanish model exists at:
https://huggingface.co/myshell-ai/MeloTTS-Spanish
You already have the full conversion pipeline done for the Chinese model (melotts-zh.mud). The Spanish model uses the same MeloTTS architecture, so porting it should follow the same ONNX → INT8 quantization → .axmodel + .mud process.
Request: Please provide melotts-es.mud in the MaixCAM2 model zoo, following the same structure as melotts-zh.mud.
Why this matters
Spanish is spoken by 500+ million people. Latin America represents a huge potential user base for accessibility and assistive technology applications. Both of these additions would make MaixCAM2 significantly more useful for non-English speaking markets, with very little additional work given the pipelines already exist.
Thank you for the excellent work on MaixPy and MaixCAM2!
- 主要言語
- Python
- スター
- 850
- フォーク
- 128
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
sipeed/MaixPy のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 20/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 45/100
似ている issue
-
bug server
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
sportsdataverse/sportsdataverse-py#641 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
googleapis/google-cloud-python#18532 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信