Cinematic Avatar: Direct the Shot, Not Just the Script
The Cinematic Avatar skill in Varosity's public skills library lets you direct a staged avatar scene from a single prompt — framing, motion, setting, mood — using HeyGen's v3 cinematic engine, routed through one Varosity API key. It's the tier above a talking head: instead of an avatar reading a script to camera, you describe a shot and the engine stages the avatar in it. That's the whole premise.
I kept running into the same wall. I'd want a spokesperson walking through a lab, or standing on a rooftop at dusk with a city behind them, and what I had was a tool that put an avatar in front of a white background and had them read my script. That's fine for some things. It wasn't what I actually wanted.
What I wanted was to direct. Subject, action, setting, lens, mood, time of day. A proper shot brief, not a teleprompter.
So that's what this skill does.
Two engines, one decision
There are two engines under the hood and knowing which to use matters.
heygen-cinematic takes your scene prompt and renders a directed 4–15s shot of a chosen avatar look. You pick the look from your library, describe the shot like you're talking to a cinematographer, and it stages the avatar in that scene. Good for when you know exactly what you want — a specific person, a specific framing, a specific moment.
heygen-video-agent takes a single brief and produces a finished short video. The agent storyboards and builds it. You're not directing a shot; you're handing over a concept and getting something produced. It can run without an avatar at all if that's what the brief calls for.
Pick by intent. If you need a specific look staged in a scene, use cinematic. If you need the system to produce a whole short from one brief, use video-agent. Same key, same Credits, different jobs.
What you actually get out of it
The output is an MP4 at the outputUrl that comes back from the job after polling. That URL lives on Varosity's asset infrastructure and is durable — save it to your manifest, download it if you need a local copy. Don't re-render when you can reuse.
The rendered shot is a real scene. Not a green screen composite you have to finish in post. Not a background image pasted behind a floating head. The avatar is staged in the environment you described.
For a multi-shot sequence — say a 45-second piece — you write 3–4 shots, render them sequentially, and concat them. Cinematic caps at 15 seconds per shot. That's not a limitation once you start thinking in shots instead of continuous takes. A 45s piece is three shots. Write three briefs, chain the renders, concat with ffmpeg. The key is keeping the same avatarIds across the sequence so the character stays consistent.
How the flow actually works
Variosity is the gateway. You write the prompt, pick the engine, and POST to /api/v1/video/generate with a Varosity key (vsk_…) scoped to generate:video. The job comes back immediately with a jobId and status running. Then you poll GET /api/v1/jobs/{jobId} every 20 seconds until you get succeeded or failed. On success, outputUrl is the MP4 and durationSec is the billed length.
For a cinematic shot the request looks like this. You set modelId to heygen-cinematic, write your scene prompt, set aspectRatio (16:9, 9:16, or 1:1), set durationSec between 4 and 15, and in providerOptions you pass avatarIds — an array of 1–3 look IDs from your account's avatar library. Resolution is 720p or 1080p. That's the shape.
Here's a concrete example. Say I'm making a 10-second opener for a product launch. The brief: mid-40s founder type in a charcoal blazer walks slowly through a sunlit modern lab, camera tracking left, shallow depth of field, golden hour light, confident and measured. I pull the right look ID from GET /api/v1/heygen/library?only=avatars — that's the authoritative call, don't guess IDs. I set durationSec to 10, resolution to 1080p, aspectRatio to 16:9, and fire the POST. I get a jobId back. I poll every 20 seconds. Typically within a few minutes I get succeeded and an outputUrl pointing to the rendered MP4. I download it, drop it into my timeline, and it's the first shot.
For video-agent the request is similar but the prompt is a production brief — what the finished video should be — rather than a shot description. You can optionally pass avatarId, voiceId, styleId. The agent handles the rest. Poll the same way.
Getting look and voice IDs
GET /api/v1/heygen/library?only=avatars returns your available avatar looks with id, name, gender, previewImageUrl, and defaultVoiceId. Use the id value as avatarIds in cinematic or avatarId in video-agent. For voices, call with ?only=voices or pull the full library. Do this at the start of a session — don't hardcode IDs you grabbed six months ago and haven't checked since.
The gotchas worth knowing
The prompt is the driver. A vague prompt gets generic output. "A person talking about our product" tells the engine nothing about framing, nothing about setting, nothing about mood. A shot brief — subject, action, setting, lens, mood, time of day — is what makes the difference. Write it like you're talking to a DP on set.
Duration clamps at 15 seconds for cinematic. Ask for more and it renders 15. If your story needs 45 seconds, segment it. Three shots. Chain them.
If you request 4k resolution it silently downgrades to 1080p. Ask for 720p or 1080p and get what you asked for.
Rate limit is roughly 3 generations per rolling 60 seconds, shared across your account. Render shots sequentially. On a 429, back off 60 then 90 then 120 seconds. Polling is free and doesn't count against the rate limit.
If a render stalls past 20 minutes of polling, re-submit the shot once. Transient provider stalls happen. If it stalls again, surface the jobId and stop — don't loop indefinitely.
Don't bake brand text or titles into the scene prompt. Overlay them in post. Let the render be clean video; add words with ffmpeg or HTML/CSS after.
And if what you actually need is an avatar reading a script to camera — that's the talking-avatar skill, not this one. Wrong engine for that job.
Cost
Cinematic shots run $0.08 per second in Varosity Credits. A 10-second shot is $0.80. The full 4–15 second range puts a single shot at roughly $0.40 to $1.20. Video-agent bills on the length of the produced output. Failed renders aren't billed on either engine or either payment path.
You can BYOK your HeyGen key — go to varosity.ai → Settings → Providers → HeyGen and add your X-Api-Key. With BYOK you pay zero markup. Without it the platform key is used and Credits carry a 5% surcharge. You can also force the path per-call with authPreference:"byok" or authPreference:"credits" in the request. If a BYOK key is invalid you'll get a 409 provider_not_configured — fix the key. If Credits are empty you'll get a quota error — top up.
Installing and running the skill
The skill lives in Varosity's public skills library at https://varosity.ai/api/v1/skills/cinematic-avatar. To pull it locally: curl -s https://varosity.ai/api/v1/skills/cinematic-avatar > ~/.varosity/skills/cinematic-avatar.md. It's set to auto-update — call refresh_skills at the start of a session to make sure you're running the current version.
From there it's available to any MCP-connected agent or callable directly via REST or the CLI. The trigger phrases in the skill spec are what any agent needs to route to it — things like "put the avatar in a scene," "prompt-to-finished avatar video," "a high-end avatar shot, not a talking head." The skill handles the routing logic; you handle the brief.
One key. One scope. The rest is direction.