Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Ability to control order of multi-modal context along with prompts

オープン
#352 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
38/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
静か
技術スタック
python
領域
ai

調査の方向性

この issue では、ファイル、テスト、エントリーポイントが示されていません。まず、現在画像をテキストより前に配置しているマルチモーダルコンテキストの組み立て箇所を見つけ、次に prompt とメディアがどのように表現されているかを確認してください。完了条件には、テキストと画像のインターリーブをサポートする、順序を定義する仕組みと、ここで説明されている順序の例を対象とするカバレッジを含める必要があります。

索引モデルが issue の本文から書いたものです。

説明

enhancement needs-discussion on-roadmap
Priority Level

Medium (Nice to have)

Is your feature request related to a problem? Please describe.

Right now all images are packed before text prompts. There is no way to control the order of multi-modal context with text prompts.

Describe the solution you'd like

The desire is to have finer control over the exact order of multi-modal context:

<images> <text>
Page1 <image1> page2 <image2> <more text>

{“<image1”>: path, “<image2>”: path2}
“Page1 <image1> page2 <image2>”

“Frame 1 timestamp x” <video frame1>
<audio> <image> <video> <image> <image> <audio>

This is necessary when working with multi-modal data.

Describe alternatives you've considered

No response

Additional context

No response

主要言語
Python
スター
2.3k
フォーク
211
平均マージ
3日 20時間
マージ済み PR(30日)
47

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA-NeMo/DataDesigner のほかの issue

NVIDIA-NeMo/DataDesigner の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。