OmniHuman 1.5 API
Drives a talking avatar from one photo + an audio clip. Drop on any shot in a brand agent's storyboard.
$0.084/s on Varosity credits ($0.080/s with your own key). Clips up to 60s.
no fal account needed · no subscription · failed renders never billed
Best for
- talking avatars
- from single photo
- lip-sync from audio
- drop-on-any-shot avatar layer
Trade-offs
- needs photo + audio
- fixed framing
- no scene control
Call it from your code
One Varosity key works for every model. Video renders are async — the response has a jobId; poll GET /api/v1/jobs/{jobId} until status is "succeeded", then read outputUrl.
curl -X POST https://varosity.ai/api/v1/video/generate \
-H "Authorization: Bearer $VAROSITY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "cinematic aerial of a coastal road at golden hour, slow push-in",
"modelId": "omnihuman",
"aspectRatio": "16:9",
"durationSec": 5
}'Or ask your agent
Connect Claude, Cursor, or any MCP client to https://varosity.ai/api/mcp with your key, then say:
"Use Varosity with omnihuman to make a 5-second clip of …"Agent setup guide
Alternatives to OmniHuman 1.5
- HeyGen Avatar (Digital Twin)studio-grade lip sync · Avatar IV / Avatar V motion engines$0.084/s
- HeyGen Photo Avatar (Avatar IV)animate ANY photo as the speaker · Avatar IV motion engine$0.084/s
- HeyGen Video Translatemultilingual lip-sync dubbing · preserves original speaker appearance$0.105/s
- D-ID AI Presentertalking avatars from any photo · text-to-presenter$0.053/s
- HeyGen Video Agentprompt → finished video · agent writes script, picks avatar & scenes$0.126/s
- LatentSync Lip-Syncsmooth temporal consistency · fast inference$0.053/s