ai-avatar-video
🎯Skillfrom runcomfy-com/skills
Create AI avatar, talking-head, and lip-sync videos on RunComfy, routing across OmniHuman for full-body audio-driven avatars, Wan 2-7 for mouth sync, HappyHorse for native in-pass audio, and Seedance v2 Pro for multi-modal cinematic generation.
Overview
AI Avatar & Talking Head Video is a Claude Code skill that creates audio-driven avatar, talking-head, and lip-sync videos through the RunComfy CLI. It intelligently routes across five models based on user intent: OmniHuman for full-body audio-driven avatars from a portrait plus audio file, Wan 2-7 for open-weights lip-sync with full scene control, Wan 2-2 Animate for stylized/illustrated character animation, HappyHorse 1.0 for script-to-video without a pre-recorded audio file, and Seedance v2 Pro for multi-modal cinematic compositions combining reference images, videos, and audio.
Key Features
- Five specialized routes: OmniHuman (default, portrait + audio file), Wan 2-7 (scene control + custom audio), Wan 2-2 Animate (stylized/illustrated characters), HappyHorse 1.0 (script-only, no audio file needed), and Seedance v2 Pro (up to 9 reference images, 3 reference videos, 3 reference audio tracks)
- Intent-based model selection: Automatically classifies whether the user has a pre-recorded audio file or only a script, whether the subject is photoreal or stylized, and whether the output needs single-shot simplicity or cinematic composition
- Multi-language dubbing support: Use the same portrait with different audio files per language to create multi-language brand videos with consistent identity across all variants
- Chaining with other skills: Combine with image generation skills to first create a portrait, then animate it as a talking head with OmniHuman or extend the result using video-extend
Who is this for?
- Marketing teams creating UGC-style product ads and virtual presenter videos with custom voiceovers
- Content creators building multi-language dubbed videos from a single portrait and swapped audio tracks
- Developers and agencies seeking a programmable alternative to HeyGen or Synthesia for automated avatar video generation
- Animators working with illustrated or stylized characters who need audio-synchronized full-body motion
Same repository
runcomfy-com/skills(30 items)
Installation
npx vibeindex add runcomfy-com/skills --skill ai-avatar-videonpx skills add runcomfy-com/skills --skill ai-avatar-video~/.claude/skills/ai-avatar-video/SKILL.mdSKILL.md
More from this repository10
A Claude Code skill for generating and editing images with OpenAI GPT Image 2 (ChatGPT Images 2.0) via the RunComfy API, excelling at embedded text, multilingual typography, logos, and directive precision for layout-critical imagery.
A Claude Code skill that acts as a smart router for image editing on RunComfy, automatically selecting the best model (Nano Banana Edit, GPT Image 2 Edit, Flux Kontext Pro, or Z-Image Turbo Inpaint) based on the user's editing intent.
Extend an image beyond its original canvas on RunComfy — uncrop, change aspect ratio, or fill in what the camera missed. Routes across Nano Banana 2 Edit, GPT Image 2 Edit, FLUX Kontext Pro, and brand edit endpoints based on whether the outpaint is prose-driven, reference-driven, or brand-locked.
A Claude Code skill for generating images with Google Nano Banana 2, a flash-tier Gemini-family text-to-image model on RunComfy, optimized for rapid ideation, social thumbnail batches, and strong in-image typography rendering.
Swap a face into a still image or video on RunComfy, routing across multiple models including Wan 2-2 Animate for character animation, GPT Image 2 Edit for precise still face swap, Flux Kontext for high-fidelity local edits, and Kling Motion Control for transferring motion onto a target character.
A Claude Code skill for fast image generation with Black Forest Labs' Flux 2 Klein on RunComfy, offering sub-second latency for rapid creative iteration, multi-reference brand styling, and declarative subject-first prompts in 9B and 4B variants.
A Claude Code skill for generating cinematic short-form video with ByteDance Seedance 2.0 Pro via the RunComfy CLI, supporting multi-modal references (up to 9 images, 3 videos, 3 audio) with native lip-synced audio.
A Claude Code skill for mask-driven image inpainting on RunComfy, routing to Z-Image Turbo Inpainting for mask-based edits or to identity-preserving edit models when regions must be described in prose, supporting object removal, watermark removal, and region replacement.
A Claude Code skill that generates AI music on RunComfy, routing between ElevenLabs Music Generation for premium 44.1 kHz stereo vocal tracks and ACE Step for budget-friendly tag-driven composition, plus audio inpainting to fix sections and outpainting to extend tracks.
A Claude Code skill that lip-syncs faces to audio tracks on RunComfy, routing across OmniHuman, Sync Labs, Kling lipsync, and Creatify models based on whether the input is a portrait still, existing video, or a text script to generate and sync.