
※本ページはプロモーションが含まれています
MiniMax H3のプロンプトの書き方|最小テンプレから0秒アンカー・話者ID・文字化けの対処まで【2026年最新】
H3は自由文ではなく決まった構造で書くモデルです。公式のフォーマットに沿うと、セリフの口の動き・カメラの振れ幅・カットの時刻まで指示が通ります。まず動かすための最小テンプレ、縦型の配信風と番組風の応用テンプレ、そして実際に出た崩れ方(文字化け・ロゴの増殖・カットで別人になる)と、その対処をまとめます。
概要
MiniMax H3 は、プロンプトを自由文で書くと実力が出ないモデルです。「かわいい女の子がカフェで笑っている、シネマティックに」のような書き方だと、セリフは読まれず、カメラは勝手に動き、渡した画像からも離れていきます。
決まった構造で書くと、セリフの口の動き・カメラの振れ幅・カットの時刻・音の層まで指示が通ります。まず動かすための最小の形から、実際に出た崩れ方の直し方までをまとめます。
まずこれだけ(最小テンプレ)
画像から動かす場合の最小形です。埋めるのは5か所だけです。
TEXTFor the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. Begin from this exact angle with 〈画面に残すもの〉 intact. integrated_multimodal_description: [Shot 1] A 10-second clip begins from the exact first frame. The camera 〈静止 or 動き〉. 〈誰が〉〈何をしている最中か〉. 〈その人〉(S1) says: <d>[Japanese] 〈セリフ〉</d> 〈相手〉(S2) 〈リアクション〉: <d>[Japanese] 〈セリフ〉</d> overall_soundscape: 〈環境音・物音・声の様子〉 non_diegetic_music: 〈BGM。無しなら None.〉
埋める5か所は「残すもの / カメラ / 誰が何をしている最中か / セリフ / 音」です。これだけで、渡した画像から始まって、指定した日本語を喋る10秒になります。
埋めた例
TEXTFor the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. Begin from this exact angle with the cafe interior and the window light intact. integrated_multimodal_description: [Shot 1] A 10-second clip begins from the exact first frame. The camera holds a static medium shot. A young Japanese woman sits by the window, both hands around a coffee cup, and leans forward slightly as she starts to speak. She says in a soft, warm voice (S1): <d>[Japanese] ねえ、ちょっと聞いてほしいことがあるの</d> She lowers her eyes, smiles, and turns the cup once in her hands. Each reaction follows its visible cause; no frozen pauses, no impossible movement. overall_soundscape: Quiet cafe room tone, a distant espresso machine, the soft clink of the cup on the saucer, and her close, breathy voice. non_diegetic_music: A slow piano loop, thin and warm, kept low under the dialogue.
公式フォーマットの中身
モデルリポジトリで共有されている image-to-video 用のシステムプロンプトが基準です。出力は3つのブロックに分かれます。
| ブロック | 書く内容 |
|---|---|
integrated_multimodal_description: | 映像・動作・カメラ・セリフを時系列でまとめる。本体 |
overall_soundscape: | 環境音と物音。1〜4文 |
non_diegetic_music: | 登場人物には聞こえないBGM。1〜2文。不要なら None. |
本体の中で使う要素は4つです。
| 要素 | 書き方 |
|---|---|
| 0秒アンカー | For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. |
| ショット分割 | [Shot 1] … [Shot 2] At 00:05.000, the camera cuts to … とカットの時刻を秒で指定 |
| カメラ | 「動きの種類+振れ幅+速度」。例 The camera pushes in with small amplitude at slow speed. |
| 話者とセリフ | (S1) を話者につけ、セリフは <d>[Japanese] …</d>。タグの中は言語タグとセリフだけ |
言い方や動作はタグの外に書きます。
TEXTShe says in a breathless, playful voice (S1): <d>[Japanese] もう無理ーっ!!</d>
タグの中に「叫びながら」のような指示を入れると、それごと読み上げようとします。
仕様の範囲
API のドキュメントでは、尺は4〜15秒、解像度は768Pと2K、比率は 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16。参照画像は9枚、参照動画3本、参照音声3本まで、プロンプトは7,000文字まで入ります。
画像から動かす場合、比率の指定は要りません(入力画像に追従します)。昔のテレビの画角で作りたいときは、画像の段階から 4:3 で作ってください。
0秒アンカーが生命線
image-to-video でいちばん効くのがこの1行です。
TEXTFor the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. Begin from this exact angle with 〈残すもの〉 intact.
「参考画像として使う」程度の書き方(use the reference image as the main appearance reference)だと、1秒目から顔も服も部屋も離れていきます。「渡した画像を0秒のフレームそのものとして扱え」と宣言し、残したいものを名指しすることで、焼き込んだ文字やセットが維持されます。
セリフと声
- 話者IDは
(S1)(S2)(S3)。 順番は登場順でなくてよく、固定したい人物から振って構いません。重要なのは全ショットで同じ人に同じIDを使うことです - 1ショットにセリフを詰めない。 10秒に2文が目安です。3文入れると早口になります
- 声を固定したいときは参照音声を渡す。
Voice timbre follows reference audio 1.を足さないと、生成のたびに声質が変わります。同じキャラでシリーズを作るなら必須です
カメラとカット
カメラは3点セットで書きます。種類(push in / pan / tilt / static)+振れ幅(small / medium / large amplitude)+速度(slow / medium / fast speed)。
TEXTThe camera pushes in with small amplitude at slow speed.
カットは時刻で切ります。範囲(SHOT 1 (0:00-0:03))ではなく、「何秒で切り替わるか」です。
TEXT[Shot 2] At 00:05.000, the camera cuts to a tighter waist-up two-shot.
10秒なら多くて1回。カットのたびに顔と服が揺らぐので、人物を見せる動画は1ショットで通すほうが安定します。
応用テンプレ1:縦型の配信風(2ショット)
SNSのショート向けに、縦の画像から2カットで作る形です。
TEXTFor the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. Begin from this exact framing with the same woman, her hairstyle, her outfit, the gaming chair, the desk microphone and the RGB-lit PC behind her, intact. integrated_multimodal_description: [Shot 1] A 10-second vertical clip from a Japanese streamer's live broadcast begins from the exact first frame. Medium shot, handheld with a slight natural drift, pushing in with small amplitude at slow speed. The young Japanese woman (S1) glances down at the smartphone in her left hand, reads a comment, then looks up into the camera with a bright smile. She speaks cheerfully, gesturing once with her free hand: <d>[Japanese] みんな来てくれてありがとう!</d> Her hands stay in frame and keep a natural shape. [Shot 2] At 00:05.000, the camera cuts to a tighter medium close-up from the same side of the room, the same lighting and the same background. She is mid-laugh, leans slightly forward and speaks with rising energy (S1): <d>[Japanese] 今日もめっちゃ楽しく配信していくよ〜!</d> She gives a small wave and ends on a warm smile held toward the lens. Her appearance, outfit, hairstyle and the room stay identical across the cut; only the framing changes. Each movement follows its visible cause; no frozen pauses, no impossible movement. No on-screen text, no subtitles, no watermark. overall_soundscape: Quiet room tone of a home streaming setup, the faint hum of a PC fan, small fabric and chair sounds as she shifts, and her voice close on the desk microphone. non_diegetic_music: None.
カットをまたぐときは、「見た目と部屋は同じ、変わるのは画角だけ」を1文で縛るのが効きます。これが無いと2ショット目で別人が出ます。
応用テンプレ2:画面の文字を動かす
画像に焼き込んだテロップを、動画の中で差し替えたりカウントダウンさせたりできます。番組風の企画で効きます。
TEXT[Shot 1] … The timer box at the lower right decreases once per second in real time, counting down from 「残り8秒」 toward 「残り4秒」 across this shot, the numeral changing cleanly without warping. [Shot 2] At 00:05.000, … At that exact moment the bottom caption is replaced by a new telop 「正解!!」 in huge white type with a thick red outline, and the timer box flashes twice and freezes at 「残り3秒」. The show logo at the upper left and the telops at the upper right keep exactly the same characters as in the first frame; do not change, re-spell or re-draw any character.
実際に出た結果がこれです。カウントダウンが1秒ずつ減り、決めゼリフが「正解!!」に差し替わっています。
動かす文字と固定する文字を分けて書くのがコツです。
| 動かす | 固定する |
|---|---|
| カウントダウンの数字 | 番組ロゴ |
| 途中で差し替わる決めゼリフ | コーナータイトル |
| 価格・在庫・タイムコード | ワイプの枠 |
画像の作り方そのものはAI美女でバラエティ番組風の動画を作るにまとめています。
つまずく癖と対処
焼き込んだ文字が別の字になる
10秒のあいだに文字が化けます。実例では「息が続くまで答えろ!」が「息が続くまて答える!」に、タイムコードの「14:33」が「14:4ろ」になりました。長い文字列ほど、そして動かす対象にした文字ほど化けます。
対処は2段構えです。文字数を削る(画面上部のテロップは8文字前後まで)、そして不変を明示する。
TEXTThe upper-right telops 「息が続くまで!」 and 「賞金10万円!?」 must keep exactly the same characters as in the first frame. Do not change, re-spell or re-draw any character.
タイムコードのように動かす文字が化けるときは、動かすのをやめて固定するのがいちばん確実です。変化はワイプの中や被写体の動きで作れます。
背景に文字やロゴが増える
セットの空いている場所に、番組ロゴの複製がぼやけて浮くことがあります。モデルが「ロゴはセットの一部」と解釈して増やすためです。
TEXTNo additional logos, signs, banners or text anywhere in the set or on the backdrop. The only graphics in the frame are the ones already present in the first frame.
カットの後に別人になる
同一性を1文で縛ります。書かないと、2ショット目で髪型・服・部屋が変わります。
TEXTHer appearance, outfit, hairstyle and the room stay identical across the cut; only the framing changes.
否定形が効かない
awkward hand anatomy, distorted fingers, extra limbs, low-quality blur のような禁止リストはほとんど効きません。効くのは No on-screen text No watermark のように、画面に足されるものを断る短い指定だけです。
手の破綻を避けたいなら、否定で書くより肯定で状態を決めるほうが通ります。
TEXTHer hands stay in frame and keep a natural shape.
変化を2つ以上同時に動かすと溶ける
カウントダウンと価格を同時に動かす、カメラを大きく振りながらテロップを差し替える、といった指定をすると文字が溶けます。動かす軸は1本に絞るのが原則です。
途中で画面の下が無地になる
テロップの差し替えは、消える時刻と出る時刻を縛らないと空白が生まれます。
TEXTThe bottom caption 「もう無理ーっ!!」 stays visible from 0.00 until 00:08.000, at which moment it is replaced by 「正解!!」. The lower area is never left without a caption.
どこで動かすか
H3はオープンウェイトなので、ComfyUIのテンプレートを選ぶだけで手に入ります。メニューの「テンプレート」から「ビデオ」を開くと MiniMax H3:画像から動画へ が並んでいます。
ComfyUIのテンプレート一覧。ビデオの先頭にMiniMax H3が並んでいる
カードの Int8 は量子化版で、VRAMに余裕がないPCでもこちらが選べます。導入手順はMiniMax H3をローカルで動かす、必要なスペックはAI美女を作るPCのスペックにまとめています。
ローカルで回せることがプロンプトの詰めでは効きます。 文字の化け方を見ながら2〜3回は作り直すことになるので、1本ごとに課金を気にしていると、そこで妥協した画を使うことになります。
AI美女のイメージよくある質問(FAQ)
Q. 自由文で書いてはいけませんか?
出ることは出ますが、セリフの読み上げ・カットの時刻・カメラの振れ幅は通りません。VISUAL STYLE CAMERA AUDIO のような独自の見出しもほぼ無視されます。構造を合わせるだけで同じ素材から別物が出ます。
Q. セリフが読み上げられません。
<d>[Japanese] …</d> のタグになっているか確認してください。「次のセリフを言う」のような地の文だと読まれないことがあります。タグの中は言語タグとセリフだけにして、言い方は外に書きます。
Q. 同じキャラで何本も作りたいです。
参照画像を9枚まで渡せるので、人物・衣装・背景を別々に固定できます。声は参照音声を渡して Voice timbre follows reference audio 1. を足してください。これを入れないと声が毎回変わります。
Q. 何秒まで作れますか?
4〜15秒です。10秒を超えるとカットを割りたくなりますが、人物の同一性が崩れやすいので、長くするより10秒を複数本つくって並べるほうが破綻しません。
Q. 日本語のテロップは何文字まで入りますか?
画像の段階では1行12文字程度まで崩れずに出ますが、動画にすると長い行ほど化けます。画面に残す文字は8文字前後に抑えるのが安全です。
まとめ
- H3は構造で書くモデル。
integrated_multimodal_description/overall_soundscape/non_diegetic_musicの3ブロックが基本形 - 画像から動かすなら0秒アンカーが生命線。「残すもの」を名指しする
- セリフは
(S1)と<d>[Japanese] …</d>。言い方はタグの外。声を固定するなら参照音声 - カメラは「種類+振れ幅+速度」、カットは時刻で指定。10秒なら多くて1回
- 文字は化ける。短くする・不変を明示する・動かす軸は1本に絞る
- 禁止リストは効かない。肯定で状態を決めるほうが通る











