Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Feature Request: Latin/Spanish OCR model + MeloTTS Spanish for MaixCAM2

オープン
#196 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
静か
技術スタック
python
領域
ai, embedded-iot

調査の方向性

まず、既存の pp_ocr_en.mud と melotts-zh.mud のアーティファクト、および AX630 用に確立されている Pulsar2 変換パイプラインを調査します。リンクされているラテン語 PP-OCR モデルとスペイン語 MeloTTS モデルが、これらのパイプラインと互換性があるか確認します。pp_ocr_latin.mud と melotts-es.mud の動作するアーティファクトを MaixCAM2 model zoo に追加できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Feature Request: Latin/Spanish OCR model + MeloTTS Spanish for MaixCAM2

Platform: MaixCAM2 / MaixPy v4


Context

I am building an assistive vision system for blind and visually impaired users using MaixCAM2. The device uses Depth-Anything-V2 for obstacle detection, YOLO11n for object identification, SmolVLM for scene description, and PP-OCR for reading signs and text. All audio feedback is via MeloTTS. The target users are Spanish speakers (Latin America).


Request 1: Latin/Spanish PP-OCR recognition model

The current pp_ocr_en.mud model fails to correctly recognize Spanish text. It confuses letters like Ñ, accented vowels, and common Spanish character combinations.

There is already a pre-converted ONNX model available at:
https://huggingface.co/docato/PaddleOCR_Mobile_Models — latin_PP-OCRv3_mobile_rec_infer.onnx

This model covers Latin-script languages including Spanish, French, Portuguese, etc. Since you already have the Pulsar2 pipeline set up for PP-OCR models on AX630, compiling this to .axmodel + .mud should follow the same process as the existing English model.

Request: Please provide pp_ocr_latin.mud (or similar) in the MaixCAM2 model zoo, using the Latin PP-OCRv3 recognition model compiled for AX630.


Request 2: MeloTTS Spanish model

The current melotts-zh.mud only supports Chinese and basic English. There is no Spanish TTS option available on MaixCAM2.

The upstream Spanish model exists at:
https://huggingface.co/myshell-ai/MeloTTS-Spanish

You already have the full conversion pipeline done for the Chinese model (melotts-zh.mud). The Spanish model uses the same MeloTTS architecture, so porting it should follow the same ONNX → INT8 quantization → .axmodel + .mud process.

Request: Please provide melotts-es.mud in the MaixCAM2 model zoo, following the same structure as melotts-zh.mud.


Why this matters

Spanish is spoken by 500+ million people. Latin America represents a huge potential user base for accessibility and assistive technology applications. Both of these additions would make MaixCAM2 significantly more useful for non-English speaking markets, with very little additional work given the pipelines already exist.

Thank you for the excellent work on MaixPy and MaixCAM2!

主要言語
Python
スター
850
フォーク
128
PR マージ指標
30日以内にマージされた PR はありません

環境構築

このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

sipeed/MaixPy のほかの issue

sipeed/MaixPy の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。