Turn a Script Into a Hosted Voiceover URL — One API Call
The "Add Voiceover" skill turns a text script into a hosted MP3 URL via Varosity AI. You give it text and a voice choice; it returns an audio URL you drop straight into an app, video, demo, or IVR. One synchronous call — no job polling, no second API key.
I built this because I kept running into the same wall. You're putting together a product demo or an onboarding flow, you need a voiceover line, and suddenly you're staring at a new provider dashboard, a new API key, a new SDK, and a docs page that assumes you have three hours. The audio part of a project kept costing more setup time than it deserved.
What I actually wanted was: give an agent a script, get back a URL. Done. That's what this skill does.
---
What you get out of it
A hosted audioUrl — something like https://cdn.varosity.ai/audio/vo-xyz.mp3 — that you can put in an <audio> tag, feed to a video editor, wire into a phone tree, or just hand to a user as a download. The file is there immediately. No polling a job ID, no webhook to set up. The response from POST /api/tts carries the URL directly.
You also get three paths to a voice, depending on what you're working with.
---
Three ways to pick a voice
The fastest path is just omitting a voice entirely. Skip the voice setup step and POST /api/tts uses a good default. For a quick demo line or an internal tool, that's often all you need.
If you want something specific, hit GET /api/v1/voices to list the premade library plus any voices you've already created. Copy a voiceId, drop it into your synthesis call.
Want a voice that doesn't exist yet? Two options:
Design one from a description. POST /api/v1/voices/design with something like "warm, calm female narrator, mid-30s, American" and an optional audition line. You get back preview URLs and a generatedVoiceId. Give it a name and it's saved for reuse.
Clone one from a sample. If you have a recording — your own voice, a client's voice, a character voice — POST /api/v1/voices with a sampleUrls array and removeBackgroundNoise: true. You get a voiceId back immediately, usable in the next call.
The clone path is what makes this useful for anything brand-specific. A startup that wants their founder's voice on every demo video. A SaaS product that needs a consistent narrator across onboarding. You do the clone once; you reuse the voiceId forever.
---
The actual synthesis step
Once you have a voice (or you're using the default), the call is simple:
POST https://varosity.ai/api/tts Authorization: Bearer vsk_... Content-Type: application/json
{ "text": "Meet Acme — ship features in minutes, not weeks.", "format": "mp3", "speed": 1.0 } ```
That's the minimum. voiceId, format, and speed are all optional. You get back:
{ "audioUrl": "https://cdn.varosity.ai/audio/vo-acme.mp3" }Drop it in:
<audio controls src="https://cdn.varosity.ai/audio/vo-acme.mp3"></audio>
Text can be up to 5000 characters. Format options are mp3, wav, ogg, opus, aac, flac, pcm. Speed runs 0.25 to 4.0 — though honestly, anything outside 0.8–1.2 starts sounding wrong for narration.
---
Multi-line scripts need a different call
Here's the part most people miss. If you push a ten-line script as one POST /api/tts call, the pacing suffers. The model treats it as one continuous block. Pauses between lines go flat. It sounds like a wall of text read aloud.
For anything with multiple lines — or multiple speakers — use POST /api/v1/audio/assemble instead. You break the script into segments, one object per line, each with its own text and optional voiceId:
POST https://varosity.ai/api/v1/audio/assemble
{
"segments": [
{ "text": "Welcome to Acme.", "voiceId": "voiceA" },
{ "text": "Let's take a tour.", "voiceId": "voiceA" }
],
"gapSec": 0.35
}You get back one stitched audioUrl. Up to 60 segments, gap between 0 and 3 seconds. If you're doing a product video with a host and a character, give each segment the right voiceId and the assembly call handles the rest.
This is the same auth, same API key. Just a different endpoint for a different job.
---
Cost and a few things to watch
A one-line voiceover runs roughly $0.01–$0.03 on Varosity Credits. A short designed or cloned voice adds a few cents, once. Failed generations aren't billed — if the call errors, nothing is charged.
The key needs the generate:voice scope. If you get a 403 missing scope, re-issue the key with that scope from the API Keys page at varosity.ai.
If your cloned voice sounds off, the most common cause is a noisy sample or one that's too short. Use a clean recording, at least 30 seconds, mono if possible, and pass removeBackgroundNoise: true. A bad sample produces a mediocre clone no matter what.
If a voiceId comes back not found, run GET /api/v1/voices to see what's actually available. Premade voices and your owned voices both show up there.
One more thing: TTS is effectively deterministic for the same text, voice, model, and speed. If you need to reproduce an exact take, save the audioUrl from the first synthesis and reuse it. Re-synthesizing the same inputs will generally sound the same, but persisting the URL is the safe move.
---
When not to use this skill
If you're producing a full multi-voice podcast episode with cloned guests, that's a different skill — varosity-podcast-production. If you need a narrated video end to end, pair this skill with custom-video-generation. This skill does one thing: text in, audio URL out. It doesn't edit video, it doesn't add music beds beyond the optional intro field in assemble, it doesn't handle transcript alignment.
---
How to install and run it
The skill lives in the Varosity public skills library. To pull it:
curl -s https://varosity.ai/api/v1/skills/add-voiceover > ~/.varosity/skills/add-voiceover.md
It's set to auto_update: true, so if you're using the MCP server, call refresh_skills at the start of a session to make sure you're on the latest version.
From there, any agent connected to the Varosity MCP server at https://varosity.ai/api/mcp can run it. The key is set once on the connection — your vsk_ key from the API Keys page. REST calls go in the Authorization: Bearer vsk_... header.
The trigger is simple: a developer wants a voiceover from a script. The agent reads the skill, runs the steps, returns the URL. You don't have to prompt-engineer it; the skill file tells the agent exactly what to do.
That's the whole thing. Script in, MP3 URL out.