One Skill, Every Video Model — Varosity Agent Video SDK
The Varosity Agent Video SDK is a reusable agent skill in Varosity's public skill library that gives any MCP-compatible agent direct access to every major video generation model through a single API key. It's version 3.2, maintained by Varosity.ai, and it runs standalone — no external workflow tools, no third-party dependencies, just your agent and one key.
I built this because I kept hitting the same wall. You want to add video generation to an agent, and suddenly you're juggling five different API keys, five different SDKs, five different failure modes. Kling wants its own auth. Runway has its own job polling pattern. Veo has its own endpoint structure. And if you want to swap models mid-project — say, iterate on Wan 2.5 because it's cheap, then finish on Veo 3.1 for the hero shot — you're wiring that routing yourself. Every time. For every project.
What I actually wanted was one thing my agent could call, that already knew all the models, already handled the polling, already knew the cost tradeoffs. So I built it.
What you get
A single Varosity API key (starts with vsk_) gives your agent access to the full model roster: Veo 3.1, Kling 3.0 Pro, Seedance 4.5, Runway Gen-4.5, Luma Ray 3, MiniMax Hailuo 02, Wan 2.5, Pika 2.2 and 2.5, Sora 2, Hunyuan Video, LTX Video, D-ID AI Presenter, WAN Effects, LatentSync, several Wavespeed-routed variants, and OmniHuman 1.5 for avatar generation. That's not a marketing list — that's what list_models kind=video returns right now.
Beyond video, the same key covers images (FLUX, Imagen 4, Ideogram V3, Recraft V3, DALL-E 3, and more), voice synthesis (ElevenLabs v2 and v3, OpenAI TTS-1 and TTS-1 HD, Deepgram Aura 2), and music generation (Suno v4, Google Lyria 2, ElevenLabs Music). Everything bills to one Varosity Credits pool. No per-vendor billing accounts. No surprises at the end of the month from a model you forgot was on a different card.
The skill ships with 34 MCP tools organized into six functional areas: discovery and setup, project management, generation, job and shot management, planning and storyboard, and rendering. Your agent doesn't have to know the difference between a fal-routed job and a Runway direct job. That's already handled.
How to actually use it
Every session starts the same way. Call refresh_skills first, then list_models. Do not skip refresh_skills. The model roster changes — new models get added, prices shift, parameters update — and a stale skill will produce errors that are genuinely annoying to debug. This is the one thing I'd tell you to tattoo on your agent system prompt.
Here's a concrete example. Say you're building a product demo video for a physical gadget — a portable coffee grinder. Your agent needs a 15-second cinematic clip showing the product in use.
First: refresh_skills, then list_models kind=video to confirm what's available. Then create_project — call this once per user request, not once per shot.
Before generating any video, run pick_reference_images. This is mandatory. It generates three candidate reference images so you or the user can confirm the visual direction before spending credits on a full render. Pick one.
Now call generate_video. You're passing the modelId (say, wan-2.5 for iteration — it's $0.07/second, the cheapest realistic motion option), your prompt, durationSec, aspectRatio, your project_id, the shot_index, and the reference_image_url from the step above. The call returns a job ID.
Poll that job with get_job until status is complete. You get back an output URL.
Not happy with the motion? Swap to Kling 3.0 Pro via Replicate at $0.09/second and re-run. Still iterating. Once you're locked, re-render the hero shot on Veo 3.1 at $0.15/second for final delivery quality.
Need narration over it? generate_voice — OpenAI TTS-1 HD at $0.008/second if you want natural prosody, Deepgram Aura 2 at $0.002/second if you need to keep costs down. Need a music bed? generate_music with Suno v4 or Lyria 2.
When all shots are done, render_project stitches everything into a final MP4 and returns the URL. That's the whole loop.
For longer-form work, plan_storyboard and generate_storyboard_keyframes let your agent build out a full shot list before any video credits get spent. You can also use suggest_model — pass it a shot description and it returns ranked model recommendations. Useful when you're not sure whether a particular scene is better handled by Luma Ray 3 (fluid camera movement, good for B-roll) or Runway Gen-4.5 (effects and transitions).
A few things to watch
OmniHuman 1.5 is not in the main model picker. It's a kind: avatar model — it needs a reference photo plus an audio clip, and you access it specifically when the use case is a talking-head avatar video. Don't try to call it like a standard text-to-video model.
Fish Audio and Cartesia are not on Varosity. If your agent tries to route to those, it'll fail. Don't suggest them.
MiniMax Music is gated — it's a continuation model that needs a reference song, not a text-to-music model. If someone asks for music from a text prompt, use Suno v4 or Lyria 2 instead.
On Suno v4: review licensing before any commercial use. It's in beta. The quality is excellent but the licensing situation is something to verify before you ship a client deliverable.
Cost-wise: LTX Video at $0.04/second is the cheapest option if you're doing high-volume iteration and don't need photorealism. Wan 2.5 at $0.07/second is the sweet spot for realistic product motion. Veo 3.1 at $0.15/second is what you use when it actually ships to a client. Don't use Veo 3.1 for iteration — you'll burn through credits fast.
Wavespeed-routed variants are worth knowing about. Runway Gen4 via Wavespeed is $0.01/second — that's not a typo. Pika 2.2 via Wavespeed is $0.04/second. If you're building a pipeline that generates a lot of draft clips, these matter.
Installing and running it
The skill lives in Varosity's public skill library. The source URL is https://varosity.ai/api/v1/skills/varosity-agent-video-sdk and it installs locally to ~/.varosity/skills/varosity-agent-video-sdk.md. It has auto_update: true, so once it's installed, refresh_skills at session start keeps it current without manual intervention.
You need one thing to run it: a Varosity API key. One key. That's it. No per-model credentials, no vendor accounts, no external workflow dependencies. Your agent calls the MCP tools, Varosity handles the routing to fal, Replicate, Runway, Google, OpenAI, whatever vendor backs the model you picked.
If you're building agents that need to produce video — product demos, ad creative, training content, social clips, anything — this is the fastest path from agent to finished MP4 I know of. I built it because I needed it. It works.