text chat: expose image input (enable multimodal M3, incl. multi-image)
まだ誰も着手していません。
評価
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 初心者へのやさしさ
- 74/100
- issue の種類
- 機能追加
- 明瞭さ
- 明確に書かれている
- 活発さ
- 静か
- 技術スタック
- typescript
調査の方向性
src/commands/text/chat.ts から始め、特に parseMessages と chat body builder を確認してから、src/utils/image.ts と既存の vision describe image の処理を調べます。ローカルパスまたは URL による画像入力を繰り返し指定できるようにし、Anthropic image blocks を生成し、複数の画像をサポートし、必要な場合は MiniMax-M3 をデフォルトにします。messages JSON ファイルなしでテキストチャットから 1 枚以上の画像を送信できれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Summary
mmx text chat uses the multimodal MiniMax-M3 model by default, but the CLI exposes no image input flag. Users cannot send an image (let alone multiple images) to M3 through text chat without hand-writing a base64 messages JSON file. Since M3 is multimodal, this is a significant hidden capability.
Current behavior
mmx text chatonly accepts text:--message,--messages-file,--system. There is no--imageflag.parseMessages(src/commands/text/chat.ts) passescontentthrough asstring | ContentBlock[], so image blocks can reach the API via--messages-file— but the CLI does no image handling (no path→base64 conversion, unlikevision describe).- For single-image description there is
mmx vision describe(which hits/v1/coding_plan/vlm, single-image only). - No CLI path exists for multi-image input (compare/diff/joint analysis of 2+ images in one call), even though M3 supports it.
Expected behavior
A first-class image input on text chat, e.g.:
# single image
mmx text chat --model MiniMax-M3 --image ./photo.jpg --message "What breed is this dog?"
# multiple images (repeatable)
mmx text chat --model MiniMax-M3 \
--image ./before.png --image ./after.png \
--message "List every visual difference between these two."
The flag should accept local paths / http(s) URLs and auto base64-encode them (reusing toDataUri from src/utils/image.ts), then inject them as image content blocks alongside the text message.
Evidence — multi-image already works via M3
I verified that M3 accepts multiple images in one call through mmx text chat --messages-file. Example (CN region, API key auth):
node -e '
const fs = require("fs");
const img = (p) => ({ type: "image", source: { type: "base64", media_type: "image/png", data: fs.readFileSync(p).toString("base64") } });
fs.writeFileSync("/tmp/m.json", JSON.stringify([{
role: "user",
content: [
{ type: "text", text: "I am giving you TWO images. Describe one detail unique to each." },
img("/tmp/a.png"), img("/tmp/b.png")
]
}]));
'
mmx text chat --model MiniMax-M3 --messages-file /tmp/m.json --non-interactive --quiet
# → M3 correctly describes both images and distinguishes them
Format gotcha worth surfacing
mmx text chat posts to the Anthropic /messages endpoint (chatEndpoint returns ${baseUrl}/anthropic/v1/messages), not the OpenAI /chat/completions endpoint. So image content blocks must use the Anthropic shape:
// ❌ OpenAI shape — rejected: "unsupported content type 'image_url'"
{ "type": "image_url", "image_url": { "url": "data:..." } }
// ✅ Anthropic shape — works
{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "<base64>" } }
The OpenAI-compatible image format documented at https://platform.minimaxi.com/docs/api-reference/text-openai-api (type: "image_url") does not work through this CLI's text chat, because the CLI routes to the Anthropic endpoint. This mismatch took a while to debug — a --image flag that auto-formats correctly (or at least a doc note) would help a lot.
Suggested implementation
- Add a repeatable
--image <path-or-url>flag totext chat. - In
parseMessages/ the chat body builder, convert each--imageviatoDataUri, then append{ type: "image", source: { type: "base64", media_type, data } }blocks to the user message'scontent(convertingcontentfrom string to array when images are present). - When
--imageis present, default--modeltoMiniMax-M3if not set. - Optionally reuse the same
--imageflag on a futurevisionsubcommand for multi-image, since M3's chat path strictly supersedes the single-image/vlmendpoint for multi-image use cases.
Happy to open a PR if this design sounds reasonable.
- 主要言語
- TypeScript
- スター
- 2.2k
- フォーク
- 181
- 平均マージ
- 9時間 4分
- マージ済み PR(30日)
- 13
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
MiniMax-AI/cli のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
MiniMax-AI/cli#259 · コメント 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
MiniMax-AI/cli#258 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 35/100
MiniMax-AI/cli#265 ·
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
MiniMax-AI/cli#260 ·
-
難易度 3/5 1〜2日 初心者へのやさしさ 76/100
MiniMax-AI/cli#257 · コメント 1 件 ·
似ている issue
-
VerificationGate: ATTRIBUTION quote guard never matches a normal quotation (\b around the quote) オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
danielmiessler/LifeOS#2234 ·
-
T: Bug
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
-
難易度 1/5 1時間未満 初心者へのやさしさ 85/100
-
Mend: dependency security vulnerability untriaged
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100