Engineering AI Image & Video Prompts
Don't write vague descriptions — architect prompts with explicit layers: subject → environment → lighting → color → style → mood → technical spec. Then apply platform-specific syntax.
Example (DALL-E):
A digitally illustrated sloth sitting at a sleek futuristic workstation, surrounded by
holographic data screens. Dark, near-black background with deep navy tones. Teal (#00D4AA)
neon light from the screens casts electric aqua reflections. Purple (#7B61FF) ambient glow
in background. Serene, unhurried expression. Cyberpunk aesthetic, clean design. High detail
digital illustration, cinematic lighting, 8K quality, professional marketing artwork.
Same concept for Midjourney (keyword clusters + params):
sloth at futuristic workstation, holographic data screens, deep navy background, electric
teal neon lighting, purple ambient glow, cyberpunk aesthetic, digital illustration, cinematic,
8k, professional marketing quality
--ar 16:9 --v 6 --style raw
Progress:
- Identify the target platform (DALL-E, Midjourney, SD/Flux, or video tool)
- Confirm aspect ratio / platform destination (feed, story, hero, mockup)
- Build the prompt using the layered structure for that platform
- Add color direction using brand palette tokens (descriptive + hex anchor)
- Add negative prompt / exclusions ("no text", "no watermark", etc.)
- Add quality boosters appropriate to the platform
- Output in the standard format (below) with iteration guidance
- If result misses, log iteration (Version 1 → Version 2 → Locked)
DALL-E — natural language, full sentences, layered scene description. Best for illustrations, mascots, hero images.
Midjourney — comma-separated keyword clusters + parameters (--ar, --v 6, --style raw, --no, --chaos). Best for photorealism, editorial/lifestyle shots.
Stable Diffusion / Flux — weighted tokens (term:1.2) + exhaustive negative prompt. Best for controlled stylistic output, product mockups, UI assets. Keep positive prompt ≤200 chars.
Video tools — each has distinct conventions; verified vs. conservative status matters (see below).
| Tool | Status | Structure | Length |
|---|---|---|---|
| Veo 3.1 | ✅ verified | Cinematography+Subject+Action+Context+Style | ≤500 chars |
| Runway Gen-4.5 | ⚠️ conservative | subject+action, composition, lighting, cinematic | ≤500 chars |
| Kling 3.0 | ✅ verified | Motion→Camera→Lighting→Duration→Quality tail | ≤600 chars |
| FLUX.1 | ✅ verified | weighted subject, environment, lighting, style, negatives | ≤200 chars |
| Luma Ray 2 | ⚠️ conservative | Mood prefix: subject+action+minimal context | ≤150 chars |
⚠️ tools use conservative defaults from unverifiable docs — don't "improve" them without confirming against live official documentation first. Treat every unverified adapter as a hypothesis, not a fact.
| Token | Hex | Prompt language |
|---|---|---|
| Navy/surface | #0A0F1C | "deep navy background", "near-black environment" |
| Teal | #00D4AA | "vivid teal", "electric aqua glow", "cyan-teal accent" |
| Purple | #7B61FF | "electric purple", "violet neon", "indigo with purple cast" |
| Coral | #FF6B6B | "warm coral", "salmon pink energy" |
| Amber | #FFB347 | "warm amber", "golden orange glow" |
Always pair descriptive color language WITH the hex anchor — e.g., "electric teal (#00D4AA)" — this reduces drift across regenerations.
Platform
[DALL-E / Midjourney / Stable Diffusion / Flux / Veo / Runway / Kling / Luma]
Aspect Ratio
[1:1 / 9:16 / 16:9 / 4:5] — [use case]
Positive Prompt
[ready to paste]
Negative Prompt (if applicable)
[negative prompt text]
Platform Parameters
[--ar, --v, --style, --no, --chaos, etc.]
Color Direction Notes
[how palette is referenced]
What to Watch For
[failure modes specific to this prompt type]
Iteration Path
[which parameter to adjust first if the pass misses]
Example 1 — Fitness campaign photo (Midjourney) Input: "Need a vertical Instagram Story image of an athlete training, on-brand dark/teal aesthetic." Output:
young adult athlete in dark athletic wear, urban gym environment, dramatic teal neon side
lighting, dark moody atmosphere, cyberpunk aesthetic, fitness motivation, cinematic portrait
photography, shallow depth of field, editorial quality, high contrast, deep navy background
--ar 9:16 --v 6 --style raw --q 2 --no stock photo, fake smile, oversaturated, bright background
Example 2 — UI feature illustration (DALL-E) Input: "Flat illustration for the app's sleep-tracking feature screen, no text." Output:
Flat vector illustration for a mobile app feature screen depicting sleep tracking — a calm
crescent moon and gentle wave lines representing sleep cycles. Color palette strictly: deep
navy background (#0A0F1C), teal (#00D4AA) primary accent, purple (#7B61FF) secondary. Clean,
minimal, geometric shapes. No text, no words, no letters. Modern app illustration style,
Figma-ready flat design, professional product design quality.
Example 3 — Kling video prompt (image-to-video) Input: "Animate this product photo — coffee mug on a counter, subtle motion, 5 seconds." Output:
Steam rises gently from the ceramic mug as morning light shifts subtly across the wood counter.
Camera holds a steady slow push-in. Warm soft natural lighting. Duration: 5 seconds. Natural
physics, detailed motion, cinematic lighting.
- Always specify background color/tone explicitly — models default to neutral/white otherwise
- Lead with subject → environment → lighting → style, in that order
- Include a negative prompt for every Midjourney/SD/Flux generation
- Add "no text, no words, no letters" on any asset destined for a layout with overlaid copy
- For photography: use "authentic", "editorial", "real moment" — never let it read as stock photo
- For mascots/characters: include "professional character design" + "marketing illustration quality"
- Quality boosters that consistently help: "8K", "high detail", "cinematic", "professional quality", "editorial"
- For video tools, respect character-length ceilings — verbose prompts confuse most models except Kling
- Veo: always append "No subtitles. No text overlays." — it hallucinates text otherwise
- Don't mix DALL-E's sentence style into Midjourney (it tolerates prose but performs worse than keyword clusters)
- Don't skip aspect ratio — mismatched ratios force ugly crops later
- Don't treat ⚠️-marked video tool conventions as verified fact — confirm against live docs before trusting them for high-stakes production runs
- Don't omit hex anchors when brand consistency across multiple generations matters — descriptive color words alone drift
- Don't write negative prompts as vague vibes ("nothing weird") — list concrete exclusions
- Don't over-detail Runway or Luma prompts — both underperform with long, layered descriptions relative to Veo/Kling