One API Key, Every Video Platform: How This Skill Works
The Varosity Video Orchestration skill is a reusable agent playbook in the Varosity public skills library that coordinates multi-shot video production across every Varosity-connected platform — Kling, Veo, Runway, Hailuo, Pika, Luma for video, Suno for music, ElevenLabs for voice — through a single API key and one pool of Varosity Credits. It handles provider fallback automatically, so if one model's unavailable, the run doesn't die. You install it once and any agent you run can call it.
Here's the problem I kept running into before I built this.
Every video platform has its own API. Its own key. Its own billing dashboard. Its own way of naming models, its own polling logic for async jobs, its own error shapes. If you wanted to produce a 30-second brand clip with cinematic video, a custom music track, and a voiceover, you were wiring together three separate SDKs, managing three separate credit balances, and writing fallback logic for each one independently. And if Kling went down mid-run, your agent just... failed. You lost the whole job.
What I actually wanted was a single endpoint I could point an agent at and say: make this video. One key, one billing pool, one pipeline.
That's what this skill is.
What you actually get out of it
When an agent runs this skill, it walks through a deliberate sequence. It doesn't just fire off API calls. The first thing it does is show you a format selector — short-form or long-form, lifestyle or brand film, 9:16 or 16:9. That single answer drives every downstream decision: which models get used, what aspect ratio the shots render at, how long the music needs to be.
After you pick a format, the agent drafts a full creative brief. Concept, shot list, visual mood, which model it's planning to use, estimated cost. Then it generates three reference image candidates using Varosity's image generation endpoint. You pick one. That image becomes the visual anchor for every shot in the project — same colors, same lighting, same aesthetic across the whole piece.
That approval gate is the part most people skip when they're building this stuff themselves. You approve the brief and the visual direction before a single expensive video render gets submitted. The image generation to get there costs around three cents. A bad video render that doesn't match what you had in mind costs two to ten dollars and can't be undone. The gate exists for a reason.
After you approve, the agent runs the rest without stopping: create_project, generate_video for every shot using your reference image URL, get_job to poll all of them, then render_project to composite everything together. Music and voice get generated in parallel. The output is a single production-ready MP4 — video, audio, branding — ready to post.
A concrete run-through
Say you're making a 30-second Instagram Reel for a coffee brand. You want three shots: beans being poured, a pour-over close-up, someone's hands wrapping around a mug. Cinematic, warm, 9:16.
You tell your agent to run the Varosity Video Orchestration skill. It shows you the format selector. You reply: A, 2, TikTok/Reels. The agent drafts a brief — three shots, each 5-8 seconds, warm golden-hour palette, Hailuo for fluid motion on the pour shots, Luma for the lifestyle close-up, Suno generating a low-key ambient track at 35 seconds, ElevenLabs for a three-word end-card voiceover.
It generates three reference image candidates via POST /api/v1/images. You look at them, pick the second one — it has the right warmth — and reply "approved, maybe slightly cooler highlights."
From there the agent doesn't pause again. It calls POST /api/v1/video/generate for each shot, passing your reference image URL and the consistency parameter so every shot inherits the same visual signature. It polls get_job across all three concurrently. When they're done it calls render_project, which pulls in the Suno track and the ElevenLabs line and composites everything.
You get back one MP4.
The whole thing — the brief, the approval, the renders, the audio, the composite — runs through one Varosity API key. The Credits come out of one pool. You don't touch a Suno dashboard or an ElevenLabs account or a Runway billing page.
Model selection and fallback
The skill distinguishes between platform-funded models and BYOK models. Platform-funded models — ws-hailuo-02, ws-runway-gen4, ws-pika-2.2, ws-luma-ray-2 — don't require you to have your own provider account. These are the ones to use first. If you have funded accounts with fal.ai, Replicate, or Google, BYOK models like kling-3.0, veo-3.1-direct, and seedance-4.5 unlock additional options.
If a provider goes down during a run, the skill routes to the documented fallback chain rather than failing. muapi-wan-t2v is there as a text-to-video fallback when the ws-* models aren't available. The run keeps going.
For music, Suno handles original composition — any genre, any mood. For voice, ElevenLabs runs through Varosity, so you're not managing a separate ElevenLabs key.
What to watch out for
Cost per second varies by model. Kling runs around $0.105 per second, Veo around $0.158 per second. A five-second shot on Veo is about eighty cents. Three shots plus music plus voice on a mid-tier model selection will typically run two to six dollars for a finished 30-second piece. Failed renders aren't billed, which matters when a provider hiccups.
The Stage 0 approval gate isn't optional. If an agent skips it and goes straight to video renders without an approved brief, you'll spend real money on output that probably doesn't match what you wanted. The skill's rules are explicit about this: the brief must be approved first. Don't override that.
Aspect ratio is set once at brief approval — 9:16 or 16:9 — and applies to every shot. Make sure you've confirmed the platform before you hit approve.
One more thing: the skill is self-updating. Varosity AI maintains it. At the start of any video session, your agent should call refresh_skills to pull the latest version. You can also update it manually with a curl call to the source URL — the skill file itself has the command.
How to install and run it
The skill lives in the Varosity public skills library. Source URL is https://varosity.ai/api/v1/skills/varosity-video-orchestration. It installs to ~/.varosity/skills/varosity-video-orchestration.md and has auto_update set to true by default.
You need a vsk_ API token from your Varosity config. Your agent needs three permission scopes: generate:video, generate:voice, and generate:image. That's it. One key covers all three.
Once it's installed, any agent in your stack that picks up the skills directory can call it. Point it at a content brief, let it run the format selector, approve the reference image, and let it execute. The pipeline handles the rest.
I built this because I got tired of managing five billing dashboards to produce one video. If you're building anything that involves video output — ads, product demos, brand content, social clips — this is the skill to reach for first.