Generate a full UI sound pack from text — one API key
The App Sound Effects skill on Varosity lets you generate UI sounds, game cues, notification pings, and transition effects from plain text descriptions. Each call to POST /api/v1/sfx returns a hosted audio URL synchronously — no polling, no job IDs. You can generate one effect or loop through a full pack in a single agent session.
Here's the problem I kept running into. I'd be building something — an app, a small game prototype, an internal tool — and I'd need a handful of sound effects. A button click. An error buzz. A success chime. Simple stuff. But sourcing them was always a pain. Free libraries have licensing gray areas. Paid libraries mean another account, another subscription, a bunch of files I mostly don't need. And custom audio work for three UI sounds is absurd overkill.
What I actually wanted was to type "short crisp button click, soft and modern" and get an MP3 back in one second. That's it. No new account. No separate API key. Just describe the sound and get the file.
That's what this skill does.
What you actually get
You describe a sound in plain English. The skill writes a tight prompt for it, calls POST /api/v1/sfx on Varosity, and hands back an audioUrl — something like https://cdn.varosity.ai/audio/sfx-click.mp3. You drop that URL into your game engine, an <audio> tag, or your asset pipeline. Done.
For a pack, it loops that process over your full list and returns a table: effect name on the left, hosted URL on the right. Everything's ready to use immediately.
The call is synchronous. That part matters more than it sounds. There's no job to poll, no webhook to set up, no "check back in a few seconds" logic to write. You send the request, you get the URL. One round-trip.
How to actually run it
First, you need a Varosity API key with the generate:voice scope. Sound effects share the audio scope with voice generation — that's just how the scopes are structured on Varosity. You manage keys at varosity.ai under API Keys. One key, one scope, works across both REST and MCP.
To install the skill in your agent runtime, pull it from the Varosity skills library:
curl -s https://varosity.ai/api/v1/skills/app-sound-effects > ~/.varosity/skills/app-sound-effects.md
The skill is self-updating. If you're running a Varosity MCP session, call refresh_skills at the start to pull the latest version automatically.
Once it's loaded, you trigger it with something like: "UI pack: button click, success chime, error buzz, soft notification ping." The skill takes that, writes a prompt per effect, and starts generating.
The real steps an agent runs
Here's exactly what happens under the hood, so you know what you're working with.
First, it sharpens the prompt for each effect. "Button click" becomes "short, crisp UI button click, soft and modern." "Error buzz" becomes "low muted error buzz, short and dry." The prompt describes the sound — material, action, character, length — not the scene it belongs to. That distinction matters. Vague prompts produce generic results.
Then it calls the API once per effect:
POST https://varosity.ai/api/v1/sfx Authorization: Bearer vsk_... Content-Type: application/json
{"text": "Short crisp UI button click, soft and modern", "durationSeconds": 0.6, "promptInfluence": 0.5} ```
The response comes back with audioUrl and bytes. That's the file, hosted, ready to use.
For a full pack, it runs those calls sequentially — not in parallel. That's intentional. The global generation rate limit is roughly 3 requests per window, and parallel calls will hit a 429 faster than you'd expect. Sequential with backoff on 429 is the right pattern here.
Let me walk through the actual reference run. Three effects: button click, success chime, error buzz. Duration set to ~0.6–0.8 seconds each.
POST /api/v1/sfx {"text": "Short crisp UI button click, soft and modern", "durationSeconds": 0.6}
→ {"audioUrl": "https://cdn.varosity.ai/audio/sfx-click.mp3"}POST /api/v1/sfx {"text": "Bright positive success chime, two quick ascending notes", "durationSeconds": 0.8} → {"audioUrl": "https://cdn.varosity.ai/audio/sfx-success.mp3"}
POST /api/v1/sfx {"text": "Low muted error buzz, short and dry", "durationSeconds": 0.5} → {"audioUrl": "https://cdn.varosity.ai/audio/sfx-error.mp3"} ```
Three calls, three URLs, one table delivered back. Cost: roughly $0.01–0.05 per effect on Varosity Credits, so this pack runs about $0.06. Failed generations aren't billed.
The parameters worth knowing
text is the only required field. durationSeconds takes 0.5 to 30 — omit it and the model picks the length automatically, which works fine for most UI effects. promptInfluence runs 0 to 1. Higher means the output stays closer to your description. Lower gives the model more room to interpret. If a generated sound doesn't match what you asked for, bump promptInfluence toward 1 before re-running.
One thing to know about determinism: there's no seed parameter. Re-running the same prompt gives a similar but not identical result. If a generated sound is exactly right, save that audioUrl and reuse it. Don't regenerate and expect the same file.
What can go wrong
A few failure modes worth knowing before you hit them.
If effects are getting cut off, set durationSeconds explicitly. The auto-length works well for short effects but can underestimate longer ones. Max is 30 seconds.
If sound output feels generic or off-prompt, raise promptInfluence. And tighten the prompt — describe the sound, not the context. "Medieval tavern ambience" is a scene. "Low warm background murmur, crowd noise, distant clink of glasses" is a sound description. The second one generates better.
If you hit a 429, back off and retry sequentially. Don't retry the whole pack from the top — just resume from the effect that failed.
Scope errors (403 missing scope) mean your key needs generate:voice re-issued. Provider errors (409 provider_not_configured) mean you need credits or a BYOK key configured. Wrong endpoint returns a 404 — sound effects live at POST /api/v1/sfx, not under the voice or music paths.
This skill is specifically for effects: UI sounds, game cues, notification tones, transitions, ambient one-shots. If you need spoken narration, that's the add-voiceover skill. Background music tracks use Varosity's music endpoint. Different tools, different paths.
Installing and running it
Pull it once with the curl command above, or let refresh_skills handle it in an MCP session. The source URL is https://varosity.ai/api/v1/skills/app-sound-effects. Authentication at varosity.ai → API Keys, scope generate:voice. Then just describe your effects list and let it run.
If you're building something with any kind of UI, game mechanics, or user feedback loops, you probably need sound effects. This is the fastest path I've found from "I need a button click sound" to an actual MP3 in your project.