AI Skill Report Card

Producing AI Generated Singing Character Ads

A90·Oct 7, 2026·Source: Web
YAML
--- name: producing-ai-generated-singing-character-ads description: Guides end-to-end production of AI-generated animated singing character advertisements (Pixar-style 3D look), covering creative brief, original song/lyrics, character and style bible design, keyframe generation, image-to-video animation, audio-driven lip-sync, editing, legal/IP compliance, and platform delivery specs. Use when producing short-form (15-30s) branded video ads featuring an animated mascot that sings, for TikTok, Reels, Shorts, YouTube, or Meta feed. --- Quick Start To produce a 20-second AI-generated singing mascot ad: 1. Lock a one-page brief: product + key benefit, single CTA, audience, platform/aspect ratio, tone, mandatory brand elements (logo, hex colors, disclaimers). 2. Write original lyrics (30-60 words): hook in first 3 seconds, benefit line(s), brand name repeated 2x, CTA line. 3. Generate the song (Suno/Udio/ElevenLabs Music or a composer), then separate stems (full mix, vocal, instrumental). **Lock audio before animating.** 4. Build a character reference sheet + style bible (front/3/4/side + 5 expressions) — paste the style bible verbatim into every image prompt. 5. Build a timing sheet mapping lyric lines to 8-12 shots (2-4s each) from the audio waveform. 6. Generate keyframes per shot (subject + action + setting + style + camera + lighting + mood), animate with image-to-video, then run audio-driven lip-sync on close-ups using the isolated vocal stem. 7. QA every shot for lip-sync drift, hand/face warping, label morphing, and flicker. Regenerate failures. 8. Edit to the beat, composite real product/logo, add burned-in captions, color-match, export per platform specs. Workflow Progress: - [ ] Step 1: Lock creative brief (product, benefit, CTA, audience, platform, tone, mandatory elements, constraints) - [ ] Step 2: Clear legal/IP basics (no Pixar/Disney IP, original character, original song, licensed voice, AI-disclosure rules) - [ ] Step 3: Write and produce song (lyrics, tempo/key, vocal, stems, loudness ~-14 LUFS) - [ ] Step 4: Build character reference sheet and style bible - [ ] Step 5: Write script and shot list (one action per shot, mapped to lyric lines) - [ ] Step 6: Generate keyframes (3-4 variants per shot, check hands/eyes/text/product legibility) - [ ] Step 7: Animate (image-to-video for motion) + lip-sync (audio-driven, using vocal stem) separately - [ ] Step 8: QA gate — mouth drift, warping, flicker, product accuracy, continuity - [ ] Step 9: Edit — cut to beat, captions, branding, color grade, SFX - [ ] Step 10: Export to delivery specs and run final legal/platform checklist **Step details:** 1. **Brief** — one sentence product/benefit, single CTA, audience, platform+ratio+length, tone, mandatory elements (logo/tagline/price/disclaimers/hex codes), constraints (words to avoid, regional needs). 2. **Legal/IP** — no references to Pixar/Disney (describe style instead: "soft 3D animated feature-film style"); original character design; original song/lyrics (no soundalikes); licensed/consented voice; no real-person likeness without release; apply AI-content labels per platform rules; confirm commercial-use licensing on every tool; substantiate any health/finance/claims; keep records of prompts, tool versions, licenses, consents. 3. **Audio** — structure: hook → benefit line(s) → brand name + CTA, brand name 2x minimum; 30-60 words for ~20s; tempo 100-130 BPM typical; clean dry lead vocal (no heavy reverb, for accurate lip-sync); deliver full mix + vocal stem + instrumental as WAV 44.1/48kHz, -14 LUFS. Lock before animating. 4. **Character/style bible** — species/type, proportions, face/eyes/hair/outfit, one quirk; 3 personality adjectives; visual style keywords (stylized 3D, soft rounded forms, subsurface shading, oversized eyes); lighting (warm key, soft rim); 3-5 color palette incl. brand colors; 1-2 environments; negative list (realistic skin, extra fingers, text artifacts, other brand logos, gritty tones); reference sheet (front/3/4/side + 5 expressions: neutral, happy, singing "ah", singing "oo", surprised). 5. **Shot list** — one action per shot tied to a lyric line; 2-4s per shot; mix close-up/medium/wide/product-hero; singing shots use front/3/4 face, mouth visible, minimal head turn; 2-3 product-visible shots + end card; timing sheet columns: shot #, timecode, lyric line, shot type, action, camera, transition; respect 9:16 safe zones. 6. **Keyframes** — tool with character-reference support (Midjourney, Flux/Flux Kontext, GPT-image, Gemini image, Seedream); prompt formula: subject + action + setting + style + camera + lighting + mood + style bible block; min 1080x1920; fixed seeds/character refs for consistency; 3-4 variants per shot; check fingers, eye symmetry, stray text, product legibility, mouth position. 7. **Animation + lip-sync** — split body-motion pass (image-to-video: Kling, Runway, Veo, Sora, Seedance, Luma, Hailuo) from lip-sync pass (audio-driven: Hedra, Sync, Kling lip-sync, Runway Act-Two, LatentSync, Wan) using the isolated vocal stem on close-ups; 5-10s clips; simple motion (head bobs, brow raises, blinks, gestures); avoid fast spins/complex hand-object interaction; budget 2-4 regen attempts per shot, 20-30% of shots needing fixes; consistent 24/30fps. 8. **QA gate** — check lip-sync drift, warping hands, face drift mid-clip, flickering textures, morphing labels before moving to edit. 9. **Edit** — cut to beat, lyric-aligned transitions, burned-in safe-zone-aware word-synced captions, logo/CTA end card, color-match all shots, light grain/glow for cohesion, sparse SFX, upscale if needed. Tools: CapCut, Premiere Pro, DaVinci Resolve. 10. **Delivery** — 1080x1920 (9:16) minimum, also 4:5/16:9 if needed; H.264 MP4 + AAC; ~10-20 Mbps at 1080p; 24-30fps; 15-30s; -14 LUFS no clipping; burned-in captions + separate SRT; thumbnail. Examples **Example 1:** Input: Energy drink brand wants a 20s TikTok ad (9:16), comedic tone, mascot is an anthropomorphic spark-creature, CTA "Try it now," brand colors neon blue/orange. Output: Locked brief → 45-word lyric with hook "Feel the spark!" in first 2s, brand name sung twice → Suno-generated song + stems → character sheet (spark-creature, 3/4 + 5 expressions) + style bible → 10-shot timing sheet → Midjourney keyframes with character reference → Kling image-to-video for body motion → Hedra lip-sync on 3 close-up shots using vocal stem → CapCut edit with burned captions, product composited in, end card with CTA → final 1080x1920 MP4, -14 LUFS, SRT + thumbnail delivered alongside 3 alternate hook cuts for A/B testing. **Example 2:** Input: Client asks to "make it look like Pixar" and wants to use a real celebrity's singing voice. Output: Redirect — do not reference Pixar by name or replicate recognizable character designs; use descriptive style language instead ("soft 3D animated feature-film look"). Do not clone the celebrity's voice without a signed license/consent; substitute a licensed voice actor or compliant AI voice with commercial terms, and document rights in the compliance file. Best Practices - Lock audio before any animation work — re-timing visuals afterward is expensive. - Always paste the style bible verbatim into every image prompt for character consistency. - Separate body-motion generation from lip-sync generation; use the isolated vocal stem (not full mix) for lip-sync tools. - Use the last frame of one clip as the first frame of the next for continuity; lock seeds, aspect ratio, and frame rate across all clips. - Generate products/logos as blank placeholders in AI shots, then composite the real packshot/logo in post — AI garbles labels and logos. - Budget for 30-60 total generations and 2-4 regen attempts per shot on a typical 20s/8-12 shot ad. - Build 3-5 hook variants (first 3 seconds) for A/B testing while reusing the same body footage and CTA. - Test 2-3 lip-sync tools on a single close-up before committing to one for the whole project — stylized faces behave differently than realistic ones. - Keep records of prompts, seeds, tool versions, licenses, and consent forms for every project. Common Pitfalls - Referencing Pixar/Disney by name or replicating recognizable character designs in prompts or copy. - Using soundalike melodies/lyrics or AI-cloning a real person's voice without consent. - Animating before the song is finalized, forcing costly re-timing. - Skipping the character reference sheet, causing visible inconsistency shot-to-shot. - Feeding the full music mix (instead of the isolated vocal stem) into lip-sync tools. - Overloading singing shots with hand props or fast head turns, which breaks lip-sync believability. - Trusting AI-generated product labels/logos without compositing the real asset in post. - Skipping the AI-disclosure label required by platform policy. - Forgetting safe-zone/caption placement for 9:16, resulting in UI overlap. - Not testing the final cut muted — most viewers watch without sound.
0
Grade AAI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
14/15
Workflow
15/15
Examples
17/20
Completeness
19/20
Format
14/15
Conciseness
13/15