H3 核心认知:它不是视频模型,是通用多模态生成器
MiniMax H3 不是传统意义上的"文生视频模型"。官方将其定义为 general-purpose omni-modal generation model——它把 T2V、I2V、首尾帧、参考视频、参考音频、视频编辑全部折叠进同一个预训练范式里。
关键后果:提示词不是"描述",而是"剧本"。H3 的提示词是一份结构化的视听简报(audiovisual brief),包含时间轴、镜头、对话、音效、音乐。自由格式写法会被模型误解。
| 参数 | 规格 | 来源 |
|---|---|---|
| 输出分辨率 | 2K(短边 1440px,宽屏约 2976×1248) | API 文档 |
| 帧率 | 24 FPS(电影级,无需转帧率) | API 文档 |
| 时长 | 5–15 秒(整数秒) | API 文档 |
| 音频 | 原生 32kHz 立体声,与画面同步生成 | API 文档 |
| 提示词上限 | 约 7000 字符(社区实测) | 社区实测 |
| 支持语言 | 对话支持 11 种语言(含中文);提示词结构字段建议英文 | 官方 FAQ |
| 比例 | 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 / adaptive | API 文档 |
H3 的提示词遵循"写你能看到的和听到的,不写你想感受到的"。避免"史诗感""电影级"这类抽象词,用具体的镜头、动作、声音来描述。
提示词三字段结构(Base 模式强制格式)
T2VA / I2VA / FL2VA / L2VA 共享以下三字段。R2V(Ref2VA)在此基础上扩展为六字段(见第七节)。
integrated_multimodal_description: [Shot 1] STYLE + INITIAL FRAMING + SUBJECT + ENVIRONMENT. REFERENCE OR OPENING-STATE CONTINUITY. SMALL NATURAL MOVEMENT. ACTION ONSET. SPEAKER DESCRIPTION + SPEAKER ID + DELIVERY: <d>[Language] Exact dialogue.</d> VISIBLE PHYSICAL RESPONSE. CAMERA MOTION TYPE + AMPLITUDE + SPEED. INTERMEDIATE COMPOSITIONAL CHANGES. SYNCHRONIZED DIEGETIC SOUND. FINAL CHARACTER STATE + FINAL COMPOSITION. overall_soundscape: The full environmental and physical sound environment. non_diegetic_music: The audience-only background score, or N/A.
| 字段 | 内容 | 禁止 |
|---|---|---|
integrated_multimodal_description |
主视听时间线:画面、动作、镜头、说话人、台词、同步音效 | 不要把环境音写在这里(会重复) |
overall_soundscape |
环境物理音:风声、脚步、碰撞、布料、呼吸 | 不要写台词;不要和 integrated 里的音效重复 |
non_diegetic_music |
背景音乐:乐器、速度、节奏、动态变化 | 不要写"史诗感音乐";角色能听到的音乐要写在 integrated 里 |
这些模式在 integrated_multimodal_description 之前增加一个 IMAGE-ALIGNMENT INSTRUCTION,用于锁定首帧/尾帧的内容(身份、服装、位置、光线、背景)。详见第十二节案例。
Integrated Multimodal Description 逐段拆解
3.1 Shot 声明
第一个 Shot 不需要时间戳。后续 Shot 必须使用严格递增的时间戳。
[Shot 1] Live-action, cinematic...
[Shot 2] At 00:04.500, the camera cuts to a close-up.
[Shot 1] At 00:00.000... ← 不需要
3.2 视觉风格(必须在开头声明)
在 Shot 1 开头立即声明整体风格,使用直接类别而非模糊赞美:
Live-action Cinematic 2D-animated 3D CG Claymation Watercolor Vintage film
3.3 参考图锚定(I2V / R2V)
立即识别参考图中必须保持一致的元素:
- Character identity(面部身份)
- Appearance(外貌)
- Clothing(服装)
- Body proportions(身体比例)
- Position(位置)
- Lighting(光线)
- Background(背景)
- Composition(构图)
The young female hiker shown in <Picture 1> remains on the narrow mountain ridge, preserving her exact facial identity, natural skin texture, brown hair...
3.4 说话人与对话(最容易翻车的部分)
说话人身份(Speaker ID)
- 首次出现时指定:年龄、性别、音高、音色、口音、语速
- 使用稳定 ID:
(S1)、(S2) - 多人同时说话用复合 ID:
(S1,S2) - 从未发声的角色不分配 ID
对话标签 <d>(绝对规则)
She shouts with joyful energy: <d>[English] Helloooooooo!</d>
Her lips, jaw, cheeks, throat and chest move naturally...
<d>[English, shouting loudly and happily] Helloooooooo!</d>
← 语气描述塞进了标签内
<d> 内部只能有:语言标记 + 确切的台词文字。所有表演指令(语气、情绪、语速)必须写在标签外面。
3.5 画外音(Voiceover)
使用精确短语 says in an off-screen voiceover,并在标签后声明角色嘴唇完全闭合:
(S1) says in an off-screen voiceover: <d>[English] The city never sleeps.</d> Her lips remain completely closed throughout.
3.6 跨切对话
如果对话跨越镜头切换,在连接点两侧都放置对话,并声明音频连续:
[Shot 1] ... <d>[Chinese] 你必须在日落前</d>
[Shot 2] At 00:03.500, the camera cuts to a close-up. <d>[Chinese] 赶到山顶。</d> The audio continues across the cut.
3.7 画面文字(On-Screen Text)
所有可见的文字(横幅、招牌、标签、字幕、霓虹灯)必须用英文双引号包裹,逐字保留原文和标点,不翻译:
A red neon sign reading "营业中" glows above the doorway.
The title card displays "EPISODE 1" in bold sans-serif.
生成后仍需视觉审查,不要假设拼写、品牌处理或法律许可是自动正确的。
镜头语言规范:写成自然动作,不是标签堆叠
H3 期望镜头语言读起来像自然动作描述,而不是 disconnected tags。每个短片段只选一个主导镜头运动。
Motion Type + Amplitude + Speed(运动类型 + 幅度 + 速度)
中等幅度和正常速度在添加无意义时可以省略。
| 需求 | 清晰写法 | 为什么有效 |
|---|---|---|
| 缓慢推进 | The camera pushes in with small amplitude at slow speed. | 定义了物理运动和节奏 |
| 水平揭示 | The camera pans right with large amplitude at fast speed. | 区分旋转与平移 |
| 跟随动作 | A tracking shot follows the cyclist through the gate. | 将镜头与主体运动连接 |
| 无运动 | The camera holds a static shot as the steam clears. | 阻止模型发明不想要的运动 |
| 新镜头 | [Shot 2] At 00:04.500, the camera cuts to a close-up. | 在有效时间引入真实剪辑 |
| 环绕 | The camera makes a slow clockwise orbit around the product. | 明确方向、速度和主体 |
| 手持跟随 | A handheld follow shot tracks behind the subject. | 控制纪录片式抖动的程度 |
| 焦点转移 | A rack focus shifts from the foreground watch to the background city. | 指定前后景关系 |
不要把多个不兼容的镜头运动堆到一个 15 秒片段里。"Orbit + tracking + push-in" 会让模型困惑。每个片段只选一个主导运动。
音频生成:对话、环境音、音乐的三层分离
5.1 三层音频架构
| 层级 | 位置 | 内容 | 示例 |
|---|---|---|---|
| 对话/歌唱 | integrated_multimodal_description |
角色说出的台词、同步音效 | <d>[Chinese] 你好</d> |
| 环境音景 | overall_soundscape |
ambience、脚步、碰撞、布料、天气、呼吸 | Steady high-altitude wind, subtle jacket-fabric movement... |
| 背景音乐 | non_diegetic_music |
观众能听到、角色听不到的音乐 | Sparse piano notes at a slow tempo... |
5.2 对话支持的语言
MiniMax 官方报告稳定对话支持 11 种语言:阿拉伯语、中文、英语、法语、德语、意大利语、日语、韩语、葡萄牙语、俄语、西班牙语。其他语言支持程度不一。
语言标记写法:<d>[Chinese] 台词</d>、<d>[English] Dialogue</d>、<d>[Japanese] 台詞</d>
5.3 音乐描述规范
用 1-3 句英文描述乐器、速度、节奏和动态变化。不要用抽象情绪词:
Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.
Epic cinematic music that feels emotional and inspiring.
角色能听到的音乐(收音机、手机、现场演奏)属于 diegetic 事件,必须写在 integrated_multimodal_description 里,而不是 non_diegetic_music。
5.4 物理发声表演(提升真实感)
不要写"完美的口型同步和真实表演",而是描述可观察的解剖学动作:
Her lips, jaw, cheeks, throat and chest move naturally. The jaw opens naturally, cheeks respond to pressure, throat engages, chest contracts during breath, lips shape the sustained vowel.
参考文件 R2V:给每个文件分配一个明确的任务
| 类型 | 数量 | 限制 |
|---|---|---|
| 参考图片 | 最多 9 张 | JPEG/PNG/WEBP,单张 ≤30MB |
| 参考视频 | 最多 3 段 | MP4/MOV,H.264/H.265 + AAC/MP3,单段 ≤50MB,2-15s,总计 ≤15s |
| 参考音频 | 最多 3 段 | WAV/MP3,单段 ≤15MB,2-15s,总计 ≤15s |
| 总计 | 最多 12 个文件 | 音频不能单独使用,必须至少配一个图片或视频 |
6.1 参考分配模式(官方推荐)
<Picture 1> defines the fictional adult subject's identity and clothing only.
<Video 1> defines body timing and the tracking camera only.
<Audio 1> defines the authorised voice timbre only.
黄金法则:给每个参考一个狭窄的任务。不要写"Use Image 1 for everything"。
6.2 音频参考的严格限制
音频参考只传递音色(timbre),不传递口音、情绪、语速、节奏、旋律。这些必须在提示词里用文字指定。参考音频必须是干净的人声,不能有背景音乐、重叠说话声、环境噪音。
Prompt: Character dialogue: "Follow the wind, live free. Leave worries behind, enjoy the moment." Use Audio 1 as the voice-timbre reference.
你需要提供:
- Video 1: 将要说话的角色视频
- Audio 1: 目标声音的干净样本,2-15 秒, ideally 无背景音乐
- 确切的台词文字(逐字引用,不转述)
为什么有效:台词被引用而非转述, timing 和口型有具体锚点;Audio 1 只分配了一个狭窄任务——音色。
6.3 可复用主体 vs 源素材
官方区分了两类引用:
- Subject(可复用主体):人、物体、场景、动作、效果——在目标中复用的内容
- Video label(源素材标签):时间或剪辑来源
一致的标签让复杂提示词更容易审查。
Ref2VA 六字段结构(R2V 专属)
Ref2VA(Reference to Video with Audio)是 H3 最复杂的模式,在 Base 三字段基础上扩展为六字段。这是身份锁定、风格迁移、声音克隆、视频编辑的核心入口。
全参考模式的第四字段官方名称为 detailed_description,不是 integrated_multimodal_description。以下结构严格对应官方 ref-en.txt 原文。
subject_definitions: Define what each reference file controls. Assign a narrow, explicit task to every Picture, Video, and Audio. summary: One-sentence overview of the creative direction. retention_analysis: For each reference, state what is preserved and what is changed. detailed_description: [Shot 1] STYLE + FRAMING + SUBJECT + ENVIRONMENT... overall_soundscape: Environmental and physical sound. non_diegetic_music: Audience-only background score, or N/A.
7.1 六字段顺序与作用
| 字段 | 作用 | 写法要点 |
|---|---|---|
subject_definitions |
给每个参考文件分配角色 | "<Picture 1> defines identity and clothing only." |
summary |
一句话概括创意方向 | 让审查者 3 秒理解你要做什么 |
retention_analysis |
声明每个参考的保留策略 | 使用官方 retention 标签(见下表) |
detailed_description |
主视听时间线(同 Base 的 integrated) | 但可引用已定义的 subject |
overall_soundscape |
环境物理音 | 不写台词 |
non_diegetic_music |
背景音乐 | N/A 或具体乐器描述 |
7.2 Retention 标签(官方词汇)
| 标签 | 含义 | 使用场景 |
|---|---|---|
fully_preserved |
参考中的元素在输出中完全保留,不做任何改变 | 身份克隆、品牌 logo、特定服装 |
partially_preserved |
参考中的部分元素保留,其余允许变化 | 保留面部但改变发型/服装颜色 |
attribute_transfer |
提取参考的某个属性(颜色、材质、风格),应用到新主体 | 把 A 的衣服风格穿到 B 身上 |
weak_reference |
参考仅作为弱引导,输出可以大幅偏离 | 氛围参考、大致风格参考 |
| 标签 | 含义 | 使用场景 |
|---|---|---|
fully_copy |
完整复制参考音频的内容和特征 | 精确克隆某段声音 |
partially_copy |
部分复制参考音频,其余由模型生成 | 提取音色但改变语调 |
reference |
参考音频作为生成目标,但允许合理变化 | 风格相似的背景音乐 |
weak_reference |
参考音频仅作为弱引导 | 大致情绪参考 |
loosely_inspired、style_only、do_not_use 不是官方 retention 标签。H3 可能无法识别这些词汇,导致保留策略失效。请严格使用上表中的官方标签。
7.3 Ref2VA 完整示例(官方标签)
subject_definitions: <Picture 1> defines the adult woman's exact facial identity, skin texture, and brown hair only. <Picture 2> defines the red silk evening dress and its drape behaviour only. <Video 1> defines the slow walking gait and hip sway timing only. <Audio 1> defines the authorised voice timbre only. summary: A woman in a red silk dress walks through a rainy Shanghai alley at night, speaking one line in Chinese. retention_analysis: - <Picture 1>: fully_preserved (facial identity, skin, hair) - <Picture 2>: fully_preserved (dress design, colour, material) - <Video 1>: attribute_transfer (gait timing applied to new environment) - <Audio 1>: fully_copy (voice timbre) - Background from <Picture 1>: weak_reference detailed_description: [Shot 1] Live-action, cinematic. The woman defined in <Picture 1> wears the red silk dress from <Picture 2>. She walks through a narrow Shanghai alley at night, neon signs reflecting on wet pavement. The camera tracks behind her with medium amplitude at slow speed. She stops, turns to face the camera. The woman with a calm, clear voice (S1) says softly: <d>[Chinese] 雨停了,该走了。</d> Her lips, jaw, and throat move naturally with each word. overall_soundscape: Light rain on wet stone, distant car horns, one passing bicycle, subtle silk fabric movement, footsteps on wet ground. No other speech. non_diegetic_music: N/A
Base 模式(T2V/I2V/FL2V)的 integrated_multimodal_description 里直接描述一切;Ref2VA 先通过 subject_definitions 和 retention_analysis 建立"合同",再在 detailed_description 里引用这些定义。这让复杂多参考任务的可控性大幅提升。
五大生成模式与选择逻辑
| 模式 | 何时使用 | 输入 | 输出比例 |
|---|---|---|---|
| T2VA Text to Video |
只有文字创意简报,没有任何素材 | Prompt | 可选 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 |
| I2VA Image to Video |
有一张起始图,需要让它动起来 | Prompt + 首帧图 | 跟随输入图片比例 |
| FL2VA First & Last Frame |
有首帧和尾帧,需要填充中间运动 | Prompt + 首帧图 + 尾帧图 | 跟随输入图片比例 |
| L2VA Last Frame |
只有尾帧,需要倒推生成 | Prompt + 尾帧图 | 跟随输入图片比例 |
| R2VA Reference to Video |
需要身份锁定、动作迁移、风格匹配、声音克隆、视频编辑 | Prompt + 最多 12 个参考文件 | 可选 adaptive 或固定比例 |
没有媒体输入 → T2V。图片是字面意义上的首帧/尾帧 → First & Last Frame。图片/视频/音频是作为参考而非字面帧 → R2V(Ref2VA)。
8.1 各模式提示词差异
T2V:需要完整描述视觉世界(主体、环境、动作、镜头、结尾)。
I2V:上传图已传达外观和构图,Prompt 聚焦运动、镜头行为、保留规则、环境变化、结尾。
FL2VA:描述过渡而非重复两张图。解释什么在动、什么渐变、什么保持对齐、如何收尾。选择透视、比例、风格兼容的帧。
R2V(Ref2VA):使用六字段结构,给参考排序并分配角色,通过 retention 标签控制保留程度。
中文使用专项:防念咒指南
官方明确确认 H3 支持中文提示词。但如果你把中文画面描述写得像对话,模型就会尝试生成语音。中文提示词必须严格按字段拆分。
9.1 中文提示词正确结构
integrated_multimodal_description: [Shot 1] Live-action, cinematic. A young woman stands in a Shanghai alley at dusk, neon signs reflecting on wet pavement. She adjusts her coat and looks toward the camera. The camera holds a static medium shot. No dialogue, no voiceover. Her lips remain completely closed. Synchronized diegetic sound: distant traffic, light rain on fabric. overall_soundscape: Steady rain on wet stone, distant car horns, one passing bicycle with a bell, subtle fabric movement as she adjusts her coat. No speech, no vocalization. non_diegetic_music: Sparse solo piano, slow tempo, minor key, fading out at the end.
9.2 如果必须用中文写画面描述
API 层面支持中文 prompt,但官方指南建议结构化字段用英文(镜头、风格、约束)。如果坚持用中文写主体描述,建议追加防误解声明(社区实践,待实测):
No dialogue, no voiceover, no spoken words. Characters' lips remain completely closed throughout. No on-screen text, no subtitles, no captions.
除非用户在"台词"区域明确填写了内容。
9.3 中文台词的正确写法
The young woman with a clear, gentle voice (S1) speaks softly:
<d>[Chinese] 今天的晚霞真美。</d>
Her lips shape each word naturally, jaw moves with relaxed rhythm.
关键:语气、情绪、语速写在 <d> 外面;标签内只有 [Chinese] 台词。
9.4 中文画面文字的写法
如果画面需要出现中文文字(招牌、字幕),用英文句子描述,文字内容用英文双引号包裹中文:
A red neon sign reading "营业中" glows above the doorway.
The subtitle displays "第二集" in white sans-serif.
然后必须追加约束:
Do not add Chinese text elsewhere. Do not generate garbled characters. Do not misspell.
负面约束清单:免费的质量提升
官方 45 个案例的共性:文字拼写和负面清单不增加生成成本,但品质大多来自这里。以下约束为社区实践,非官方三字段格式,建议追加在 integrated/detailed_description 末尾或作为独立备注。
Do not add any on-screen text. Do not add subtitles. Do not add captions. Do not generate garbled characters or misspellings. Do not introduce Chinese text unless explicitly requested. No soft dissolves or fluid morphs. No tearing, black frames, hard cuts, obvious VFX, or compositing seams. No giant eyes, split mouths, fangs, threatening behavior, or jump scares. No camera movement unless explicitly directed. No repetitive titles or names. Each credit and title appears only once.
| 场景 | 追加约束 |
|---|---|
| 产品广告 | Do not deform the product. Do not change label position or logo geometry. |
| 人物特写 | Do not alter facial proportions. Keep hands anatomically correct. |
| 品牌内容 | Do not introduce competing brand marks. Replace real logos with fictional ones. |
| 多镜头 | Do not repeat the same transition type. Vary wipe, mask, and cut styles. |
| 无音乐 | No music. No background score. Ambience only. |
| 无对话 | No dialogue. No voiceover. Lips remain closed. |
21 点写作自检清单(官方审计)
来自官方指南的 21 点 audit,写完 prompt 后逐条确认。
案例模板(官方规范 + 社区优化)
以下案例基于官方 Video Prompt Writing Guide 结构。Constraints 段为社区实践(非官方三字段格式),建议按需追加。
12.1 产品广告(T2V)
integrated_multimodal_description: [Shot 1] A brushed-steel automatic watch rests on black volcanic stone. Fine rain lands on the glass while the second hand moves naturally. The camera begins with an extreme macro, then makes a slow 30-degree orbit, ending on the watch face. High-contrast luxury commercial, deep black background. Keep the logo and dial geometry unchanged. Subtle rain and room-tone audio. overall_soundscape: Soft mechanical ticking, light rain on glass, one distant low thunder rumble. non_diegetic_music: One restrained low cinematic pulse at 0s, fading by 3s. No melody. Constraints: No dialogue. No on-screen text. No soft dissolves.
Constraints 段为社区实践,非官方三字段格式。
12.2 人物对话(R2V + 音频参考)
integrated_multimodal_description: [Shot 1] Live-action, cinematic. The young woman shown in <Picture 1> sits in a hospital corridor at night, preserving her exact facial identity and dark hair. 0 to 5s she returns with two paper cups and sits. 5 to 11s she speaks. 11 to 15s neither answers and one takes the other's hand. Medium-close coverage, restrained shot-reverse-shot. The young woman with a calm, low, natural voice (S1) says with controlled sadness: <d>[Chinese] 有些话,我们都不想说出口。</d> Her lips, jaw, and throat move naturally with each word. Her eyes remain fixed forward. overall_soundscape: A distant trolley, the vending machine's hum, fabric movement as she sits. non_diegetic_music: N/A Constraints: No music. No subtitles. No on-screen text. No soft dissolves.
Constraints 段为社区实践,非官方三字段格式。
12.3 风格迁移(R2V · 简化写法)
以下示例使用单字段简化写法(Base 模式风格)。进阶多参考任务请使用第七节 Ref2VA 六字段结构(subject_definitions / retention_analysis / detailed_description)。
integrated_multimodal_description: [Shot 1] Image 1 is the character and art-style reference: keep this exact hand-painted 2D webcomic look, the same watercolour texture, the same ink outline weight, the same dog, the same little hat, the same mug. Do not redraw him in 3D, do not change his proportions, do not clean up the linework. Put him in a different room and keep his composure. A cramped open-plan office at night, fluorescent tubes flickering, a wall of monitors all showing a red error state, printer paper drifting down through the frame, a small electrical fire in the corner. He sits in an office chair at the centre of the frame, mug in paw, looking straight at the camera, completely relaxed. Camera: locked off, static wide shot. No push in, no cuts. overall_soundscape: Fluorescent hum, printer grinding, soft alarm chirp every 2s, distant electrical crackle, one calm sip at 4s. non_diegetic_music: N/A Constraints: No on-screen text. No subtitles. No other characters.
Constraints 段为社区实践,非官方三字段格式。
12.4 首尾帧过渡(FL2VA)
How the reference pictures align with the generated video: - Picture 1 aligns with 0.00s (the first frame). - Picture 2 aligns with 15.00s (the last frame). integrated_multimodal_description: Hold the opening composition briefly. Daylight gradually fades as shop windows illuminate, traffic begins moving, and reflections appear across the wet street. Keep the buildings and camera position aligned. Resolve smoothly into the supplied nighttime frame without a sudden cut. No camera movement. No dialogue. overall_soundscape: Day ambience fading into evening: distant traffic increasing, one car horn, wet tire sounds, shop door chimes. non_diegetic_music: N/A
12.5 纯动作无台词(T2V · 运镜示范)
integrated_multimodal_description: [Shot 1] A 21:9 ultra-wide of a single-track causeway running dead straight across a tidal flat, the incoming sea already crossing it in thin sheets. A small pale car slows, hesitates, and begins a three-point turn as water climbs the tyres. Flat overcast light, silver water against dark wet sand, the horizon a low hard line across the middle of the frame. Slow lateral drift of the camera, the emptiness on either side doing the work. overall_soundscape: Shallow water moving over tarmac, wind, one distant gull. non_diegetic_music: N/A Constraints: No music, no dialogue, no camera cuts, no push-in.
Constraints 段为社区实践,非官方三字段格式。
常见翻车与修复方案
| 翻车现象 | 根本原因 | 修复方案 |
|---|---|---|
| 中文提示词被念出来 / 乱语 | 画面描述写得像对话,或没有声明 No dialogue | 在 integrated/detailed_description 开头加 "No dialogue, no voiceover. Lips remain completely closed." |
| 给了参考音频但声音不像 | 参考音频里有背景音乐/噪音 | 换一段干净人声,无伴奏,2-15 秒 |
| 角色没按参考图生成 | 没有写 "<Picture 1> defines identity only" | 明确分配参考任务,不要写 "use all images" |
| 画面出现乱码 / 错误英文 | 没有约束文字生成,或没有逐字拼写 | 需要出现的文字用英文双引号逐字写出;追加 "Do not generate garbled characters" |
| 镜头运动不自然 | 堆叠了多个不兼容的镜头运动 | 每个片段只选一个主导运动:static / push-in / tracking / orbit |
| 软溶解 / 流体变形 | 模型默认过渡方式 | 明确写 "No soft dissolves or fluid morphs" |
| 音频和画面不同步 | 台词文字长度和片段时长不匹配 | 控制 <d> 内文字量,15 秒片段不宜超过 20 个汉字 |
| 角色面部漂移 | 同时改变了姿势、服装和镜头角度 | 一次只改一个变量;用参考图锁定面部 |
| 生成结果太"通用" | 用了太多抽象形容词,缺少具体动作 | 用 "A does B" 替代 "epic cinematic dynamic" |
| 产品变形 / 标签错位 | 没有写保留约束 | "Preserve the exact bottle shape, label position, cap, glass material, colours, and proportions." |
| Ref2VA 参考混乱 | 没有写 subject_definitions 和 retention_analysis | 补六字段结构,用官方 retention 标签(fully_preserved / attribute_transfer / weak_reference 等) |
| FL2VA 首帧/尾帧对不上 | 缺少 IMAGE-ALIGNMENT INSTRUCTION | 在 integrated/detailed_description 之前写 "Picture 1 aligns with 0.00s; Picture 2 aligns with S.SSs" |
1. 生成一个简单版本 → 2. 识别最大失败点 → 3. 只改一个指令 → 4. 保持种子和设置不变 → 5. 同一 prompt 跑多次 → 6. 记录成功/失败/时间/成本 → 7. 保存精确 prompt 和参考顺序。