Prompting AI Images and Video
Markdown--- name: prompting-ai-images-and-video description: Converts scripts, product briefs, or creative concepts into structured, copy-paste-ready image and video generation prompts, including shot lists, character/style bibles, and image-to-video motion prompts. Use when preparing prompts for AI image or video generators (e.g., Midjourney, Runway, Kling, Sora, Meta AI, Happy Horse AI), breaking a script into shots, maintaining character/style consistency across a sequence, or planning an AI-generated video from voiceover to final edit. --- # Prompting AI Images and Video
Given a script line, subject, and style, produce this per shot:
SHOT 03 — "New sneaker feels like walking on clouds"
IMAGE PROMPT: Close-up of a woman's feet in white running sneakers landing softly on a cloud-like surface, studio setting, soft diffused lighting, pastel blue background, minimal shadows, product photography style, shallow depth of field --ar 9:16
MOTION PROMPT: Feet land in slow motion, subtle bounce on impact, camera holds static low angle, cloud surface ripples gently outward, background stays still, 3-second clip
MODEL: Midjourney (image) → Kling (video)
NOTES: Cut on the bounce peak; add soft "whoosh" sound design to mask any foot-warping artifacts
If the brief is missing required inputs, ask once for the gaps, then proceed using clearly stated assumptions.
Progress:
- Step 1: Collect inputs (script, subject, style, character/environment, tech specs, constraints)
- Step 2: Break script into shots (one action per shot, timed to voiceover)
- Step 3: Build the character/style bible (one reusable block)
- Step 4: Write image prompt per shot (subject, action, setting, style, camera, lighting, mood)
- Step 5: Write motion (I2V) prompt per shot (what moves, speed, camera movement, what stays still)
- Step 6: Check against model limits (clip length, known artifact risks, banned terms)
- Step 7: Review shot-to-shot continuity against the brief
- Step 8: Deliver shot list, prompts, style bible, and editing notes
Step 1: Collect Inputs
Required before generating anything:
- Script/voiceover — exact lines, split by beat
- Subject — what it is, audience, key message
- Style — art style, color palette, mood
- Character/environment — locked description or reference images
- Tech specs — target model, aspect ratio, clip length, total runtime
- Constraints — banned terms, brand rules, known model weaknesses to avoid (hands, text, physics)
Rule: if something is missing, ask once, then proceed with stated assumptions (e.g., "Assuming 16:9, cinematic realism, 5-second clips since no aspect ratio was given").
Step 2: Break Down the Script
One shot = one action = one beat of voiceover. Number shots sequentially and tie each to its script line.
Step 3: Build the Character/Style Bible
A single reusable block, pasted into every prompt for consistency:
CHARACTER: 28-year-old woman, short curly red hair, freckles, wearing a cropped yellow jacket and white sneakers
STYLE: 3D cartoon, Pixar-like rendering, warm saturated palette
ENVIRONMENT: sunlit urban rooftop, potted plants, string lights
CONSISTENCY TAGS: same character, same outfit, same lighting style, same color grade across all shots
Step 4: Write the Image Prompt
Formula: subject + action + setting + style + camera + lighting + mood
Use concrete, specific wording. Include shot type, angle, lens/DOF, and composition where relevant.
Step 5: Write the Motion (I2V) Prompt
Separate from the image prompt. Specify:
- What moves (subject/element)
- How fast (subtle, slow-motion, quick)
- Camera movement (static, push-in, pan, orbit, handheld)
- What stays still (anchor the rest of the frame)
Keep it short — overloaded motion prompts cause the model to drop or blend details.
Step 6: Check Model Limits
- Note max clip length for the target model
- Flag risky elements (hands, legible text, complex physics, fast multi-subject motion)
- Avoid banned/flagged terms per platform rules
Step 7: Review
For each shot, confirm: Does it match the brief? Does it flow visually into the next shot? Is anything a "near-miss" that should be regenerated rather than accepted?
Step 8: Deliver Output
Standard output package:
- Shot list/storyboard — numbered, tied to script lines
- Image + motion prompt pairs — one fixed format per shot
- Character/style bible — reused block
- Editing notes — transitions, pacing, sound cues, where to hide AI weaknesses
- Variants (optional) — 2–5 alternate prompts per shot when requested
Example 1: Product Ad Shot
Input: Script line: "This blender crushes ice in seconds." Style: cinematic realism, dark moody kitchen. Target: Runway, 16:9, 4-second clips.
Output:
SHOT 05 — "This blender crushes ice in seconds"
IMAGE PROMPT: Close-up of a black blender on a dark marble countertop, ice cubes mid-crush visible through the glass, dramatic side lighting, steam-like condensation, cinematic realism, shallow depth of field, low-key lighting --ar 16:9
MOTION PROMPT: Ice crushes rapidly inside the blender, blades spin fast, camera pushes in slowly, background stays static and dark, 4-second clip
MODEL: Runway Gen-3
NOTES: Cut before blade blur turns to artifact soup; layer a crunch sound effect to mask any motion glitching
Example 2: Character Consistency Across Shots
Input: A 3-shot sequence of the same animated fox character walking through a forest.
Output:
STYLE BIBLE:
CHARACTER: Orange fox, white chest fur, oversized round eyes, wearing a small green scarf
STYLE: Flat 2D storybook illustration, warm autumn palette
ENVIRONMENT: Dense forest, falling leaves, dappled sunlight
SHOT 01: Wide shot, fox walking left to right through forest path, leaves falling, gentle breeze, warm afternoon light — Motion: fox trots steadily, camera pans slowly right, leaves drift down continuously
SHOT 02: Medium shot, fox pausing to sniff a mushroom, soft golden light — Motion: head tilts down slowly, tail flicks once, camera holds static
SHOT 03: Close-up, fox looking up toward camera, eyes widen — Motion: subtle blink, ears perk up, camera holds static, background blurs slightly
- Reuse the exact character/style bible text verbatim in every prompt — don't paraphrase it shot to shot.
- Write image and motion prompts separately; never combine static description with movement in one block.
- Use concrete nouns and measurable qualifiers ("3-second slow-motion bounce") instead of vague adjectives ("nice smooth movement").
- Name the target model explicitly — prompt phrasing that works in Midjourney won't behave the same in Kling or Sora.
- Change one variable at a time when iterating (lighting, then camera, then style) to isolate what improves the output.
- Keep a running prompt library of what worked per model, for reuse across projects.
- Plan edits to hide AI weaknesses proactively: cut before glitchy hand/text frames, use sound design and speed ramps to mask seams.
- Don't stack too many descriptors in one prompt — models drop details past a certain density; prioritize the essentials.
- Don't skip the style bible step — inconsistent characters/environments across shots is the most common failure in AI sequences.
- Don't ask AI to render legible text, complex hand poses, or precise physics unless the target model is known to handle them — plan around these instead.
- Don't accept "near-miss" outputs to save time — flag and regenerate; near-misses compound into visible inconsistency across a sequence.
- Don't mix aspect ratios or color grades mid-sequence without a deliberate reason.
- Don't forget licensing/likeness checks before using AI-generated content commercially.