Clone Any Voice, Script a Real Podcast, Ship It Today
The Varosity Podcast Production skill lets you produce a multi-voice podcast episode end-to-end through Varosity.ai — clone each speaker's voice from a real recording, then generate a scripted conversation in their own voice and speaking style. It runs on four Varosity gateway verbs: transcribe, audio/clip, voices, and audio/assemble. One API key, Varosity Credits.
I built this because I kept running into the same wall. We had guests, we had recordings, we had things worth saying — and the production overhead was the bottleneck. Not the ideas. Not the content. The logistics of getting clean audio out of a multi-person conversation and into something that sounds like a real show. That's what this skill handles.
---
The actual problem it solves
Most podcast tools want you to record fresh. That's fine if everyone's available, if the recording setup is consistent, if you have time to schedule, record, edit, mix, and export. Most builders don't have all of that lined up at once.
What I actually wanted was: take a recording that already exists — an interview, a panel clip, a past conversation — pull the voice out of it, and produce a new episode in that person's actual voice, not a generic TTS voice. The difference matters. A clone that sounds like Greg Grand sounds like Greg Grand. A generic voice sounds like a robot reading a script.
That's the gap this fills. You bring the recording. The skill handles the diarization, the clipping, the cloning, and the final stitch. You stay the orchestrator — you pick the clean segments, you mine the speaking style, you write the script.
---
What you actually get out of it
A hosted mp3 at a Varosity CDN URL. A real multi-voice episode where each speaker sounds like themselves. And because you wrote the script with their verbal fingerprints baked in — their fillers, their openers, their pet phrases, their cadence — it doesn't just sound like their voice. It sounds like them talking.
For Jon and Greg, those voiceIds are already in the reference run. For a new guest, you're running through the full clone path. Either way, the output is the same: an audioUrl you can publish directly or hand off.
---
How it actually works — the real steps
The pipeline has six stages. Here's exactly what happens.
Stage 0 — Brief and approval. Before you spend a single credit, stop and collect: the subject and angle, who's hosting and who's guesting, a public audio URL for each person you need to clone fresh, and the target length (default is around five minutes). If a speaker already has a voiceId from a previous run, use it — don't re-clone. Then confirm consent. Cloning a real person's voice requires their permission. That's not an API field, it's a human go/no-go you confirm before anything else runs.
Stage 1 — Transcribe. Call POST /api/v1/transcribe with diarize: true against the source recording. You get back segments — contiguous same-speaker runs with start and end timestamps. The opening speaker is usually the host of that source show, which means the person you actually want is typically further in. Confirm identity from the first few utterances before trusting speaker order.
Stage 2 — Pick clean windows. This is yours. Varosity gives you the segments; you decide which ones are usable. You're looking for roughly five or six clean solo runs, each at least seven seconds long, totalling somewhere between 150 and 280 seconds. Skip anything with overlap, background music, or crosstalk. The quality of the clone depends directly on the quality of what you feed it here.
Stage 3 — Clip and clone. Two calls. First, POST /api/v1/audio/clip with the source URL and the time ranges you selected — you get back public clip URLs. Then POST /api/v1/voices with a name and those clip URLs — you get back a voiceId. Use the clips promptly; they're public URLs not stored indefinitely.
Concrete example from the reference run: ``` transcribe_audio { "audioUrl": "https://.../guest-interview.mp3", "diarize": true } → segments with speaker, start, end, text
clip_audio { "audioUrl": "...", "segments": [{ "start": 12.0, "end": 19.5 }, ...] } → { "clips": [{ "url": "..." }] }
clone_voice { "name": "Dana Reed", "sampleUrls": ["<clip urls>"] } → { "voiceId": "k9Pm..." } ```
Stage 4 — Style-mining and scripting. This is the part most people skip, and it's why a lot of cloned audio sounds hollow even when the voice is technically accurate. Read the transcript. Find the verbal fingerprints: how they open sentences, what filler words they use, whether they ramble or punch, whether they restart themselves mid-thought, their recurring thesis or metaphors. Write every line in that voice.
For the TTS to read cleanly, a few things matter. Space your acronyms — "E O S" not "EOS", "C R M" not "CRM", "V T O" not "VTO". Spell out year-numbers. And write laughter as real text — "Ha," — never as a stage direction like "(laughs)". Stage directions get read aloud. That's not a recoverable error once the audio's assembled.
Your segments.json ends up looking like this:
``json
[
{ "speaker": "Jon", "voiceId": "9TIB7DTUsbGuCfMUpWAW", "text": "Welcome back. Today we're getting into E O S — and why most teams run it wrong." },
{ "speaker": "Greg", "voiceId": "7dIzHGjvwkv0uBmuwwmA", "text": "Honestly? They treat the L ten like a status meeting. At the end of the day, that's where it breaks." },
{ "speaker": "Jon", "voiceId": "9TIB7DTUsbGuCfMUpWAW", "text": "Right. So how do you fix it without blowing up the cadence?" }
]
``
Open with the host. Close with the host. Alternate like a real conversation.
Stage 5 — Assemble. Call POST /api/v1/audio/assemble with your segments, a gap between turns (0.4 seconds is the default, feels natural), the episode title, and optionally an intro sting:
``
{ "segments": <segments.json>, "gapSec": 0.4, "title": "Modern Founder & CEO — EOS",
"intro": { "audioUrl": "https://.../sting.mp3", "durationSec": 4, "crossfadeSec": 1.2, "volume": 0.8 } }
→ { "audioUrl": "https://cdn.varosity.ai/aud/ep.mp3" }
``
No sting? Omit the intro field. That's it.
Stage 6 — Deliver. Hand back the audioUrl. Publish it wherever you keep episodes.
---
What to watch out for
Diarization isn't perfect on messy audio. If the source recording has overlapping speech, music beds, or crosstalk, the segment boundaries get fuzzy. You have to check. Don't feed the model ambiguous clips and expect a clean clone — the garbage-in problem is real here, and it shows up in the voice quality more than anywhere else.
TTS is not seed-locked. That means two runs of the same script won't be bit-identical. If you need to reproduce an episode exactly, keep the segments.json and the voiceIds — you can re-assemble, but it won't be the same take.
On cost: transcribe is per-second plus five percent. Cloning and assembling a standard five-minute episode runs cents to low single dollars. It's not a budget concern for most use cases, but if you're running batches or experimenting with long-form, keep an eye on it. The best practice is to reuse existing voiceIds for returning speakers — don't re-clone someone you've already cloned.
Consent is non-negotiable. The skill surfaces this as a hard stop before Stage 1. It's not something you can batch past.
---
Installing and running it
The skill lives in Varosity's public skills library. To pull it locally:
``
curl -s https://varosity.ai/api/v1/skills/varosity-podcast-production > ~/.varosity/skills/varosity-podcast-production.md
`
It auto-updates. At the start of any session, run refresh_skills` to make sure you're on the latest version — the skill is maintained by Varosity and the reference spec can change.
You need one API key with transcribe scope for STT and generate:voice scope for cloning, TTS, and assembly. Every call goes to https://varosity.ai with Authorization: Bearer vsk_... in the header. To see what voices you already have on your account, GET /api/v1/voices returns the full list with voiceId and name.
For returning speakers like Jon (9TIB7DTUsbGuCfMUpWAW) and Greg (7dIzHGjvwkv0uBmuwwmA), skip straight to Stage 4. The clone work is done.
The skill is in the Varosity public library under category: media. Search "Varosity Podcast Production" or find it tagged: podcast, voice-cloning, multi-voice, elevenlabs, audio. Version 1.1, updated 2026-06-20. Source URL: https://varosity.ai/api/v1/skills/varosity-podcast-production.
If you've been putting off building a show because the production overhead felt like too much — this is the thing that removes that excuse.