Skill detail

ai-avatar-video

Creates AI avatar, talking-head, and lip-sync videos on RunComfy, routing across OmniHuman, Wan 2-7, HappyHorse, and Seedance.

MatchDirectReviewed for Video Editing
Sourceprime-skills/​runcomfy-agent-skillsExternal source
Reported installs361,182Popularity signal only

Inspect before use

Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.

Saved source preview

SKILL.md

The saved excerpt is a snapshot from review. The external source remains the complete and most current version.

---
name: ai-avatar-video
displayName: "AI Avatar & Talking Head Video"
allowed-tools: Bash(runcomfy *)
description: >
  Create AI avatar, talking-head, and lip-sync videos on RunComfy via
  the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven
  full-body avatar), Wan-AI Wan 2-7 (audio-driven mouth sync via
  `audio_url` on a portrait), HappyHorse 1.0 (Arena #1 t2v / i2v with
  in-pass audio), and Seedance v2 Pro (multi-modal cinematic with
  reference audio + reference subject). Picks the right model for the
  user's actual intent — UGC voiceover, virtual presenter, dubbed
  product demo, lip-synced character, dialog scene — and ships each
  model's documented prompting patterns plus the minimal `runcomfy run`
  invoke. Triggers on "talking head", "lip sync", "avatar video",
  "make X speak", "audio to video", "audio driven avatar", "virtual
  presenter", "AI spokesperson", "dubbed video", "UGC avatar",
  "HeyGen alternative", "Synthesia alternative", "digital human",
  "make this portrait talk", "video from voiceover", or any explicit
  ask to put words in a face.
homepage: https://www.runcomfy.com
license: MIT
---

# AI Avatar & Talking Head Video

Put words in a face. This skill routes across RunComfy's audio-driven avatar models — OmniHuman, Wan 2-7 with audio_url, HappyHorse, Seedance v2 — picking the right path for the user's intent and shipping the documented prompts + the exact `runcomfy run` invoke for each.

[runcomfy.com](https://www.runcomfy.com/?utm_source=skills.sh&utm_medium=skill&utm_campaign=ai-avatar-video) · [Lip-sync feature](https://www.runcomfy.com/models/feature/lip-sync?utm_source=skills.sh&utm_medium=skill&utm_campaign=ai-avatar-video) · [CLI docs](https://docs.runcomfy.com/cli/introduction?utm_source=skills.sh&utm_medium=skill&utm_campaign=ai-avatar-video)

## Powered by the RunComfy CLI

```bash
# 1. Install (see runcomfy-cli skill for details)
npm i -g @runcomfy/cli      # or:  npx -y @runcomfy/cli --version

# 2. Sign in
runcomfy login              # or in CI: export RUNCOMFY_TOKEN=<token>

# 3. Generate an avatar video
runcomfy run <vendor>/<model>/<endpoint> \
  --input '{"prompt": "...", "audio_url": "https://...", "image_url": "https://..."}' \
  --output-dir ./out
```

CLI deep dive: [`runcomfy-cli`](https://www.skills.sh/agentspace-so/runcomfy-agent-skills/runcomfy-cli) skill.

## Install this skill

```bash
npx skills add agentspace-so/runcomfy-agent-skills --skill ai-avatar-video -g
```

---

## Pick the right model for the user's intent

Listed newest first. The agent classifies user intent — pre-recorded audio file or just a script? Photoreal portrait or stylized character? Single shot or cinematic composition? — and picks one route below.

**OmniHuman** — `bytedance/omnihuman/api` *(default)*
> ByteDance audio-driven full-body avatar. Feed one portrait + one audio file, get back a video where the subject speaks / sings / gestures naturally. Listed on RunComfy's `/feature/lip-sync` as the curated default.
> Pick for: UGC voiceover, virtual presenter, dubbed product demo, multi-language clips from same portrait.
> Avoid for: no audio file available (need to generate speech from a script) — use **HappyHorse 1.0**.

**HappyHorse 1.0** — `happyhorse/happyhorse-1-0/text-to-video` (t2v) · `happyhorse/happyhorse-1-0/image-to-video` (i2v)
> Arena #1 t2v / i2v with in-pass audio generated from prompt. No external audio file required — quote the spoken line inside the prompt.
> Pick for: written script with no audio file, "write a script → get a video", concept clips, i2v talking-head from an existing portrait.
> Avoid for: precise lip-sync to a specific MP3 — audio is regenerated each call, not locked.

**Seedance v2 Pro** — `bytedance/seedance-v2/pro`
> ByteDance multi-modal flagship — up to 9 reference images, 3 reference videos, 3 reference audio tracks composed in one pass with cinematic motion / lens / lighting control.
> Pick for: cinematic monologue with reference subject + 
Read the full source on GitHub (opens external page)
Context

Related work