All articles

Generate Any Custom Video On Demand With a Reference Lock

The Custom Video Generation On Demand skill is a reusable agent playbook in Varosity's public skill library. It generates custom videos through Varosity AI, and it enforces a mandatory reference-image selection before any render happens — so you're not burning credits on a clip that misses the mark visually. Install it once, and any agent you connect to Varosity can run this full pipeline on request.

I built this because I kept running into the same failure mode. Someone describes a video they want. The model interprets it one way, you interpret it another, and $1.50 later you've got footage that's technically correct but visually nothing like what you had in mind. The prompt said "cozy coffee shop" — you pictured warm tungsten light, shallow focus, a ceramic mug. The render gave you a bright overhead shot of a generic cafe. That's not a bad model. That's a missing anchor.

The reference image is that anchor.

Why the pre-flight exists

Before this skill submits a single video job, it does something most pipelines skip: it locks a visual reference frame. If you already have an image you want the video to match, you pass it in. If you don't, the skill generates three candidate frames using Imagen 4 — costs about $0.03 — shows them to you, and asks which one fits. You pick one, that URL becomes the referenceImageUrl on the video call, and now the model has a concrete visual anchor instead of just your words.

Skipping this step is the number-one reason video renders go sideways. An off-target render on Veo 3.1 costs you $0.50 to $2. The pre-flight costs $0.03. The math is obvious.

What you actually get

A hosted MP4 URL. That's it. No watermarks, no end cards, no subtitles baked in, no auto-posting anywhere. Just a clean file you can play, download, or drop into whatever comes next. The skill is intentionally minimal on the output side so you can use it as a building block — feed the URL into an editor, a storyboard pipeline, a social scheduler, whatever.

Native audio comes along with the clip when the model generates it. You can specify a vibe in your prompt — "quiet, contemplative score" or "uptempo, kinetic energy" — but there are no licensed tracks here. It's AI-composed audio, and it's baked into the render.

The two models and when to use which

Veo 3.1 is the default for most work — cinematic output, strong on character and scene, native audio generation built in. It's what I reach for on storytelling clips, product narratives, anything that needs to feel like it was shot.

Kling 3.0 handles product motion and dynamic action better. If you're rendering something with fast movement, physics, or a product that needs to look real in motion, Kling tends to hold up better.

If you're genuinely unsure, you can hit POST /api/v1/route and let Smart Route recommend one based on your shot description. Or just call GET /api/v1/models to see the full list with capabilities and current pricing.

How the pipeline actually runs

Here's a concrete example. Say you want an 8-second vertical clip of a barista pouring latte art — cozy vibe, nothing complicated.

The agent starts by generating a reference frame. It calls POST /api/v1/images with a scene prompt, gets back an image URL, and either shows you options or uses one if you've already provided your own.

Then it engineers the prompt. Your raw ask — "cozy barista clip, 8 seconds, vertical" — gets fleshed out into something the model can actually work with: close-up, slow-motion pour, espresso into steamed milk forming a rosetta in a ceramic cup, barista's hands steady in frame, warm morning window light, shallow depth of field, soft cafe bokeh, gentle steam rising, cinematic 9:16. That detail work is what separates a prompt that produces something usable from one that produces something generic.

Then it submits to the video endpoint:

POST /api/v1/video/generate with the engineered prompt, modelId: "veo-3.1", aspectRatio: "9:16", durationSec: 8, and referenceImageUrl pointing to the chosen reference frame. The response comes back immediately with a jobId and a status of rendering.

Then it polls. Every 2–5 seconds: GET /api/v1/jobs/{jobId}. Status moves from rendering to succeeded. On succeeded, the response includes outputUrl — that's your MP4. The agent returns that URL and stops.

For that barista clip, the image pre-flight runs about $0.005. The render is 8 seconds at $0.158 per second on Veo 3.1, so roughly $1.27 total. If the render fails, you're not billed for it.

Duration limits and longer pieces

A single video call caps at 10 seconds. That's the model limit — Veo tops at 8 seconds, Kling at 10. If you need something longer, you're looking at the storyboard pipeline: POST /api/v1/storyboards to create the project and keyframes, generate each shot individually, then POST /api/v1/projects/{projectId}/stitch to combine them into one MP4. This skill handles the per-shot generation step in that workflow — it's not an end-to-end long-form editor on its own.

Aspect ratios: 9:16 for TikTok, Reels, Shorts — 16:9 for YouTube and widescreen — 1:1 and 4:5 for feed posts — 21:9 for cinematic ultrawide.

A few things to watch for

Renders take minutes. Not seconds. If you're building an agent that checks job status, don't treat a 3-minute render as a hang. Keep polling. The skill sets a ceiling of about 120 polls over roughly 10 minutes — if you hit that without a result, something's stuck on the infrastructure side, not normal latency. Abandon that job ID (you weren't billed for an unfinished render), submit a fresh call.

There's no seed parameter on these models, so exact reproduction isn't possible. The practical workaround: save the referenceImageUrl you used and the verbatim engineered prompt. Use both again and you'll get something visually equivalent — same anchor, same instructions — even if it's not bit-identical.

English prompts work best. Lip-sync is approximate. Don't promise clients frame-perfect mouth movement.

If you're hitting 404s, double-check your paths. Image generation is POST /api/v1/images. Video generation is POST /api/v1/video/generate. Job polling is GET /api/v1/jobs/{jobId}. They're different endpoints and the skill routes to each one at the right step.

Installing and running it

The skill lives in Varosity's public skill library at https://varosity.ai/api/v1/skills/custom-video-generation. It's set to auto_update: true — at the start of any video session, your agent should call the refresh_skills MCP tool to pull the latest version. Or update it manually:

curl -s https://varosity.ai/api/v1/skills/custom-video-generation > ~/.varosity/skills/custom-video-generation.md

Authentication is your vsk_ API key, managed at varosity.ai under API Keys. For REST calls, pass it as Authorization: Bearer vsk_... in the header. Your key needs both generate:video and generate:image scopes — the image pre-flight requires the second one. For MCP, connect the Varosity MCP server at https://varosity.ai/api/mcp over Streamable HTTP and configure the key once on the connection.

Billing pulls from your Varosity Credits balance by default. If you've added your own provider key under BYOK, it uses that with zero markup.

The skill is version 3.1, last updated June 20, 2026. Category is generation. It'll stay current as long as you're running refresh_skills at session start.

If you want to see the current model list, pricing, or capabilities before you build: GET /api/v1/models. That's always the source of truth — don't hardcode pricing into your agent logic.