Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

text chat: expose image input (enable multimodal M3, incl. multi-image)

オープン
#224 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
74/100
issue の種類
機能追加
明瞭さ
明確に書かれている
活発さ
静か
技術スタック
typescript
領域
ai, cli

調査の方向性

src/commands/text/chat.ts から始め、特に parseMessages と chat body builder を確認してから、src/utils/image.ts と既存の vision describe image の処理を調べます。ローカルパスまたは URL による画像入力を繰り返し指定できるようにし、Anthropic image blocks を生成し、複数の画像をサポートし、必要な場合は MiniMax-M3 をデフォルトにします。messages JSON ファイルなしでテキストチャットから 1 枚以上の画像を送信できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Summary

mmx text chat uses the multimodal MiniMax-M3 model by default, but the CLI exposes no image input flag. Users cannot send an image (let alone multiple images) to M3 through text chat without hand-writing a base64 messages JSON file. Since M3 is multimodal, this is a significant hidden capability.

Current behavior

  • mmx text chat only accepts text: --message, --messages-file, --system. There is no --image flag.
  • parseMessages (src/commands/text/chat.ts) passes content through as string | ContentBlock[], so image blocks can reach the API via --messages-file — but the CLI does no image handling (no path→base64 conversion, unlike vision describe).
  • For single-image description there is mmx vision describe (which hits /v1/coding_plan/vlm, single-image only).
  • No CLI path exists for multi-image input (compare/diff/joint analysis of 2+ images in one call), even though M3 supports it.

Expected behavior

A first-class image input on text chat, e.g.:

# single image
mmx text chat --model MiniMax-M3 --image ./photo.jpg --message "What breed is this dog?"

# multiple images (repeatable)
mmx text chat --model MiniMax-M3 \
  --image ./before.png --image ./after.png \
  --message "List every visual difference between these two."

The flag should accept local paths / http(s) URLs and auto base64-encode them (reusing toDataUri from src/utils/image.ts), then inject them as image content blocks alongside the text message.

Evidence — multi-image already works via M3

I verified that M3 accepts multiple images in one call through mmx text chat --messages-file. Example (CN region, API key auth):

node -e '
  const fs = require("fs");
  const img = (p) => ({ type: "image", source: { type: "base64", media_type: "image/png", data: fs.readFileSync(p).toString("base64") } });
  fs.writeFileSync("/tmp/m.json", JSON.stringify([{
    role: "user",
    content: [
      { type: "text", text: "I am giving you TWO images. Describe one detail unique to each." },
      img("/tmp/a.png"), img("/tmp/b.png")
    ]
  }]));
'
mmx text chat --model MiniMax-M3 --messages-file /tmp/m.json --non-interactive --quiet
# → M3 correctly describes both images and distinguishes them
Format gotcha worth surfacing

mmx text chat posts to the Anthropic /messages endpoint (chatEndpoint returns ${baseUrl}/anthropic/v1/messages), not the OpenAI /chat/completions endpoint. So image content blocks must use the Anthropic shape:

// ❌ OpenAI shape — rejected: "unsupported content type 'image_url'"
{ "type": "image_url", "image_url": { "url": "data:..." } }

// ✅ Anthropic shape — works
{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "<base64>" } }

The OpenAI-compatible image format documented at https://platform.minimaxi.com/docs/api-reference/text-openai-api (type: "image_url") does not work through this CLI's text chat, because the CLI routes to the Anthropic endpoint. This mismatch took a while to debug — a --image flag that auto-formats correctly (or at least a doc note) would help a lot.

Suggested implementation

  1. Add a repeatable --image <path-or-url> flag to text chat.
  2. In parseMessages / the chat body builder, convert each --image via toDataUri, then append { type: "image", source: { type: "base64", media_type, data } } blocks to the user message's content (converting content from string to array when images are present).
  3. When --image is present, default --model to MiniMax-M3 if not set.
  4. Optionally reuse the same --image flag on a future vision subcommand for multi-image, since M3's chat path strictly supersedes the single-image /vlm endpoint for multi-image use cases.

Happy to open a PR if this design sounds reasonable.

主要言語
TypeScript
スター
2.2k
フォーク
181
平均マージ
9時間 4分
マージ済み PR(30日)
13

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

MiniMax-AI/cli のほかの issue

MiniMax-AI/cli の issue をすべて見る

似ている issue

TypeScript の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。