Clone any voice and produce a real podcast in minutes
The Varosity Podcast Production skill lets you produce a multi-voice podcast episode end-to-end through a single API: clone each speaker's voice from a real recording, then generate a scripted conversation delivered in their own voice and speaking style. It runs as a reusable agent playbook in the Varosity public skills library, using the Varosity gateway verbs — transcribe, audio/clip, voices, audio/assemble — while your agent stays the orchestrator.
Here's what actually pushed me to build this.
I kept running into the same wall. I had interviews, recordings, source material — and I wanted to repurpose them into a clean back-and-forth conversation. Not a transcript. Not a summary. An actual listenable episode where the people sound like themselves. Every tool I tried either forced me to stitch together five different APIs, each with its own key and billing, or handed me some "magic" single-call endpoint that did too much and let me control nothing. I didn't want magic. I wanted a pipeline I could reason about.
That's the real problem this skill solves. Not just cloning a voice — lots of tools do that. It's the whole chain: get clean audio, identify the speaker, cut the right segments, clone, mine how that person actually talks, write lines that sound like them, and assemble it in one shot. And do all of that with one API key and one credit balance.
What you actually get out of it
At the end of a run, you get an audioUrl — a hosted mp3 of the finished episode. Two (or more) voices, scripted conversation, proper gaps between turns, optional intro sting. The voices aren't generic TTS. They're clones built from real recordings of the actual people. And the lines are written in each person's style — their fillers, their openers, their cadence — not generic narration dropped into a synthetic voice.
For returning speakers (say, regular hosts), you reuse the voiceId you already have. No re-cloning, no extra spend. For a new guest, you pull in any public audio of them — a prior podcast appearance, an interview, a talk — and the skill walks you through extracting clean samples and cloning from those.
Default target is about five minutes. You can go longer.
How it actually works — the real steps
The pipeline has six stages. Here's what your agent does at each one.
Stage 0: Brief. Before spending anything, the agent collects the subject and angle, the people involved and their roles (host vs. guest), a public audio URL for each new person who needs to be cloned, and the target length. If you already have a voiceId for someone — like a regular host — you pass that in and skip cloning entirely. The agent presents the brief and waits for your approval. This is also where you confirm consent. Cloning a real person's voice requires it. That's a human go/no-go before the skill touches the API.
Stage 1: Transcribe. For each new person, the agent calls POST /api/v1/transcribe with diarization on. It comes back with the full transcript plus a segments array — contiguous same-speaker runs, each with a start time, end time, and text. The first speaker in the segments is usually the host of that source recording, not your guest. You confirm identity from the content of the first few utterances before trusting the speaker order.
Stage 2: Pick clean windows. This is pure agent judgment — Varosity doesn't make this call. From the guest's segments, you choose five or six clean solo runs, each at least seven seconds, totalling roughly 150 to 280 seconds. You skip anything with overlapping voices, background music, or crosstalk. Diarization degrades on overlapping audio, so if you feed it bad segments, the clone suffers. Pick carefully.
Stage 3: Clip and clone. The agent calls POST /api/v1/audio/clip with the chosen time ranges, gets back public clip URLs, then calls POST /api/v1/voices with those URLs to create the clone. That returns a voiceId you'll use for every line this person speaks.
Concrete example — say you've got a guest named Dana Reed and her interview is at a public URL. After transcribing and picking segments, it looks like this:
clip_audio { "audioUrl": "https://.../dana-interview.mp3",
"segments": [{ "start": 12.0, "end": 19.5 }, ...] }
→ { "clips": [{ "url": "..." }] }clone_voice { "name": "Dana Reed", "sampleUrls": ["<clip urls>"] } → { "voiceId": "k9Pm..." } ```
Those clip URLs are public — feed them straight into clone_voice promptly.
Stage 4: Mine style, write the script. This is the part most people miss. A voice clone reproduces how someone sounds. But if you write generic lines and drop them into the clone, it still doesn't sound like that person. You have to read the transcript and pull their verbal fingerprints: what fillers they use ("you know," "honestly," "look"), how they open a thought, what phrases they repeat, whether they're punchy or they ramble and self-correct, what their recurring thesis is. Then you write each line in that voice.
For the TTS to read correctly: space out acronyms ("E O S" not "EOS"), spell year-numbers as words, and write laughs as actual text ("Ha,") — never stage directions like "(laughs)" because the model will read those aloud.
The script becomes a segments.json — an array of objects, each with text, voiceId, and speaker. Opens and closes with the host, alternates like a real conversation.
Example turns from a real run:
[
{ "speaker": "Jon", "voiceId": "9TIB7DTUsbGuCfMUpWAW",
"text": "Welcome back. Today we're getting into E O S — and why most teams run it wrong." },
{ "speaker": "Greg", "voiceId": "7dIzHGjvwkv0uBmuwwmA",
"text": "Honestly? They treat the L ten like a status meeting. At the end of the day, that's where it breaks." },
{ "speaker": "Jon", "voiceId": "9TIB7DTUsbGuCfMUpWAW",
"text": "Right. So how do you fix it without blowing up the cadence?" }
]Stage 5: Assemble. One call to POST /api/v1/audio/assemble with your segments array, a gapSec of 0.4, and optionally an intro sting. Returns the final audioUrl. If you don't want a sting, just omit the intro field.
POST /api/v1/audio/assemble
{ "segments": <segments.json>, "gapSec": 0.4, "title": "Modern Founder & CEO — EOS",
"intro": { "audioUrl": "https://.../sting.mp3", "durationSec": 4,
"crossfadeSec": 1.2, "volume": 0.8 } }
→ { "audioUrl": "https://cdn.varosity.ai/aud/ep.mp3" }Stage 6: Deliver. Hand back the URL or push it wherever you publish.
What to watch out for
Diarization is good but not perfect. Overlapping speech and background music confuse it. The skill flags this as a known failure mode — you verify clean solo segments before cloning. Don't skip that step and then wonder why the clone sounds off.
TTS isn't seed-locked. If you re-run the assembly, the output will differ slightly. For reproducibility, keep your segments.json and re-run from Stage 5 if you need a fresh render.
Consent isn't an API field. It's on you to confirm it before you clone a real person. The skill is explicit about this — the agent is supposed to stop and get a human answer before Stage 1 begins.
Cost runs low. Transcription is priced per second plus a small margin. Cloning and assembly run at standard voice rates. A full five-minute episode with one new guest clone comes in at cents to low single-digit dollars. Returning speakers with existing voiceIds are cheaper — no clone step.
How to install and run it
The skill lives in the Varosity public skills library. To pull it:
curl -s https://varosity.ai/api/v1/skills/varosity-podcast-production \ > ~/.varosity/skills/varosity-podcast-production.md
It's set to auto_update: true, so calling refresh_skills at the start of a session pulls the latest version. You don't need to manage versioning yourself.
You need one Varosity API key with two scopes active: transcribe for the STT step, and generate:voice for cloning, TTS, and assembly. Every call goes to https://varosity.ai with Authorization: Bearer vsk_... in the header.
If you have returning hosts with house clones already set up, enumerate them with GET /api/v1/voices — returns each voice's voiceId and name. Pass those IDs directly into Stage 4 and skip Stages 1 through 3 for those speakers.
That's the whole thing. One key, one credit balance, six stages, an episode at the end. The agent does the judgment work — which segments are clean, what the person sounds like on the page, what the conversation should actually say. Varosity handles the compute. It's the split that makes sense to me, and it's why I built it this way.