Connect Any MCP Agent to Varosity in 10 Minutes
The Varosity MCP Agent Integration skill is a complete reference playbook for connecting any MCP-compatible agent — Claude Desktop, Cursor, Claude Code, Hermes, or a custom Python or Node.js agent — to Varosity.ai's video, image, voice, and music generation platform. It covers auth, all 35 MCP tools, the skill self-update mechanism, and ready-to-paste connection configs. If you're building an agent that needs to generate media, this is the starting point.
---
The problem I kept hitting
Before I built this, every time I wanted an agent to generate a video clip or voice line, I had to hard-code credentials for whatever provider I was using that week. Runway config here, ElevenLabs config there, fal keys somewhere else. Then the model IDs would change, a new provider would come out, and I'd be digging through docs again.
What I actually wanted was one key, one endpoint, and a single place where an agent could discover what's available and just call it. That's what Varosity's MCP server is. One vsk_ token. One endpoint at https://varosity.ai/api/mcp. Thirty-five tools covering video, image, voice, and music — fal, Replicate, Runway, Luma, OpenAI, xAI, Google, ElevenLabs, Deepgram, and more — all billed through Varosity Credits.
This skill is the map for all of it.
---
What you actually get
When an agent loads this skill, it has everything it needs to:
- Authenticate to the Varosity MCP server
- Discover available models by media type
- Generate video, images, voice, and music through the right tool for each
- Manage projects and shots
- Poll async jobs correctly
- Keep itself up to date automatically
That last part matters. Model IDs change. New models ship. Tool schemas evolve. The auto_update: true flag in this skill means your agent pulls the latest version from https://varosity.ai/api/v1/skills/varosity-mcp-agent-integration at session start. You don't have to manually update anything.
---
Connecting your agent
First, generate a vsk_ token at https://varosity.ai/app/keys/api-keys. That's your single credential for everything.
Then drop the config into whatever host you're using.
Claude Desktop on macOS: Open ~/Library/Application Support/Claude/claude_desktop_config.json and merge this into the mcpServers block:
{
"mcpServers": {
"varosity": {
"url": "https://varosity.ai/api/mcp",
"transport": "streamable-http",
"headers": {
"Authorization": "Bearer vsk_<YOUR_TOKEN>"
}
}
}
}Restart Claude Desktop. Same structure works for Cursor — drop it in ~/.cursor/mcp_config.json and reload settings.
For Claude Code, it goes in ~/.claude.json or a project-level .claude/settings.json. Hermes uses hermes.config.yaml with the equivalent YAML structure, or you can pass VAROSITY_API_KEY as an environment variable.
Custom Python agent? The skill includes a working httpx example. One thing to know: MCP responses from Varosity are wrapped in a content array. You have to decode result["content"][0]["text"] as JSON to get the actual payload. Miss that and you'll be confused about why your responses look wrong.
---
How an agent actually runs a generation
Here's a concrete flow. Say you want a 5-second cinematic video clip.
1. Session start — always.
``
refresh_skills
list_models kind=video
`
Run refresh_skills first, every time. Stale skills cause hard-to-diagnose errors when tool schemas have changed. Then list_models` confirms what's actually available right now.
2. Pick a reference image.
Before generating video, call pick_reference_images with a prompt describing the visual style. It returns three candidate images for the user to choose from. This is the mandatory pre-flight unless the user is supplying their own reference image.
3. Create a project.
Call create_project once per request. Give it a title. You get back a projectId.
4. Submit the video job.
Call generate_video with the modelId from step 1, your prompt, durationSec, aspectRatio, project_id, shot_index, and reference_image_url from step 2.
generate_video modelId: kling-v3 prompt: "Slow push into a neon-lit Tokyo alley, rain-slicked pavement, bokeh lights" durationSec: 5 aspectRatio: 16:9 project_id: <projectId> shot_index: 0 reference_image_url: <url from pick_reference_images>
This returns a jobId. Not a video. A job ID.
5. Poll get_job until done.
This is where most people mess up the first time. Video rendering takes 30 seconds to 3 minutes depending on the model. Fast models like LTX are around 30 seconds. Kling 3.0, Veo 3.1, Seedance — 60 to 120 seconds is normal. Under load, up to 3 minutes.
Poll get_job every 5–8 seconds. A status of running past 60 seconds is not a hang. Don't report it as stuck before 3–4 minutes have passed. When the job succeeded, you get outputUrl — that's your video.
Images are different. Call generate_image and you get imageUrl back directly. No job, no polling. If you pass an image model to generate_video, it fails. The tool you use has to match the model's kind — list_models returns that field for every model.
Voice: Call list_voices before generate_voice. Always. It returns valid voiceId values for ElevenLabs, OpenAI TTS, Deepgram, and Cartesia. Guessing a voice ID gets you an error.
---
The gotchas worth knowing
The number-one mistake I see: passing an image model to generate_video. Models like flux-1-schnell, recraft-v3, aurora, dall-e-3, imagen-4 — those are image models. They go to generate_image. Sending them to generate_video gets rejected immediately.
The render_project tool — which stitches all shots into a single MP4 — currently returns an Unauthorized error in some configurations. It's documented that way in the skill. If it fails, fall back to client-side ffmpeg concatenation. Don't build a workflow that depends on it until that's resolved.
Costs run through Varosity Credits. The skill doesn't publish per-model credit costs inline — check your dashboard or run list_models to see current pricing. Long video renders on premium models (Kling, Veo, Seedance) cost more than fast models. Factor that in when you're building agent loops that might call generate_video multiple times.
One more thing: create_project should be called exactly once per user request. I've seen agents call it in a loop and end up with a dozen orphaned projects. The skill is explicit about this — once per request.
---
Installing and running the skill
The skill lives in Varosity's public skill library. It's stored locally at ~/.varosity/skills/varosity-mcp-agent-integration.md once installed, and the source URL is https://varosity.ai/api/v1/skills/varosity-mcp-agent-integration.
Because auto_update is set to true, the refresh_skills tool at session start will pull the latest version automatically. You install it once and the agent keeps it current.
To install: go to the Varosity skills library at https://varosity.ai, find "Varosity MCP Agent Integration" under the integration category, and add it to your agent. Then follow the connection config for your host above.
After connecting, verify it's working by asking your agent: "What video models are available in Varosity?" It should call list_models kind=video and return the current roster. If it does, you're live.
The skill is at version 2.3, last updated 2026-05-08. Everything in it — tool names, parameter shapes, connection examples, the polling flow — reflects what's actually in the platform right now. That's the point of keeping it in the library with auto-update on. When something changes, the skill changes with it.