Crafting Image Generation Prompts
YAML--- name: crafting-image-generation-prompts description: Transforms vague or weak image ideas into detailed, realistic scene prompts optimized for AI image generation models. Use when a user has a rough concept, a short prompt, or an unclear image idea and needs it expanded into a rich, production-quality prompt with proper subject, environment, lighting, composition, and style details. --- # Crafting Image Generation Prompts
Given a weak prompt, don't just add adjectives — reconstruct the entire scene: subject, action, environment, lighting, camera/lens, mood, and style.
Weak prompt: "a man in a city"
Reconstructed prompt: "A weary businessman in his 40s walking through a rain-soaked Tokyo street at night, neon signs reflecting in puddles, shot on a 35mm lens with shallow depth of field, cinematic lighting, moody blue-and-pink color grade, hyperrealistic, 8k detail."
Progress:
- Step 1: Extract the core idea — what subject, action, or feeling is the user actually after?
- Step 2: Identify missing elements (setting, time of day, mood, camera angle, style, lighting)
- Step 3: Ask 1-3 targeted clarifying questions ONLY if the idea is ambiguous enough that guessing would waste the user's time (e.g., realistic photo vs. illustration, indoor vs. outdoor)
- Step 4: Draft a full scene description covering: Subject + Action, Environment, Lighting/Time of Day, Camera/Composition, Mood/Atmosphere, Art Style/Rendering quality
- Step 5: Present the enhanced prompt back to the user and ask if it matches their vision
- Step 6: Refine based on feedback — adjust one element at a time (lighting, angle, style) rather than rewriting from scratch
- Step 7: Deliver the final polished prompt, ready to paste into an image model
Example 1: Input: "a dog on a beach" Output: "A golden retriever sprinting joyfully across a sunset beach, ocean waves crashing behind, wet sand reflecting orange-pink sky, shot with a wide-angle lens at golden hour, dynamic motion blur on the paws, warm cinematic color grade, ultra-realistic photography style."
Example 2: Input: "make a fantasy castle" Output: "A colossal gothic castle perched on a cliffside above a misty valley, illuminated by a crescent moon, torches flickering along stone battlements, dragons circling distant peaks, dramatic low-angle shot, painterly fantasy art style with rich contrast and epic scale, inspired by matte painting concept art."
Example 3: Input: "cat sitting" Output: (Ask first) "Do you want this photorealistic or stylized/illustrated? Indoors or outdoors? Any particular mood — cozy, mysterious, funny?" → then build the full prompt from the answer.
- Always preserve the user's original intent — enhance it, don't replace it with your own idea.
- Layer details in this order: subject → environment → lighting → composition/camera → mood → art style.
- Use concrete sensory language (wet pavement, golden hour glow, shallow depth of field) over vague adjectives (nice, beautiful, cool).
- Mention camera/lens or rendering style (35mm, wide-angle, cinematic, oil painting, 3D render) — this anchors the AI model's output style.
- Iterate in small steps: if the user says "too dark," adjust only lighting, not the whole scene.
- When the request is already detailed and strong, refine and tighten it rather than over-expanding.
- Don't pile on generic hype words ("stunning, amazing, epic, masterpiece") without concrete visual substance.
- Don't ignore ambiguity that changes the whole scene (realistic vs. cartoon) — a quick clarifying question saves rework.
- Don't discard the user's original concept in favor of a completely different scene.
- Don't overload the prompt with contradictory styles (e.g., "photorealistic anime watercolor 3D render").
- Don't skip mood/atmosphere — it's what makes a scene feel alive, not just descriptive.