All articles

Turn Any Website Into a 30-Second Brand Video

The Website Promo Video skill produces a research-driven 30-second brand video for any company website — end to end. You give it a URL, it researches the brand, scripts a voiceover, generates keyframes, animates clips, adds a branded end card, and mixes everything into a hosted MP4. One Varosity API key. Around $4–6 in Credits.

I kept running into the same wall. A founder would say "we need a quick promo video" and the actual path was: open five browser tabs, prompt an image model, prompt a different video model, copy-paste a voiceover into some TTS tool you're paying for separately, find a music track, wrestle ffmpeg, and somehow stitch it into something that didn't look like a Frankenstein. Every model lived in its own dashboard with its own billing. Painful.

What I actually wanted was one key, one credits balance, and an agent that could just run the whole thing. So I built it. This skill is the distilled playbook from that build.

---

What the skill actually does

It's a six-stage pipeline your agent orchestrates. Varosity supplies the generation verbs — image, video, voice, music. The agent supplies the judgment: reading the brand, writing the script, framing the shots, assembling the edit.

Stage 0 is just collecting the brief. The skill asks for the company URL, target length (default 30 seconds), visual style (cinematic live-action, motion graphics, whatever fits), audio preference, and aspect ratio. It presents the brief back and waits for your approval before spending anything. A 30-second spot runs $4–6, so it's worth the two-second confirm.

Stage 1 is research, and this is the part most people miss. The agent web-searches the company, then fetches the homepage and an About or Services page with a specific extraction prompt — pull taglines, value props, slogans, brand colors, tone. The reason: a generic promo that could be for any company is useless. The script has to echo the brand's own language. When I built the reference run for TGG Accounting, the research surfaced "The TGG Way™," "From Chaos to Confidence," and a navy-and-gold palette. That's what drove every shot prompt and every line of voiceover. You can't fake that with a template.

Stage 2 is the voiceover script. Thirty seconds at roughly 150 words per minute is about 70–78 words. The structure is simple: pain, promise, tagline, name. One thing the skill enforces that saves you a bad render: write for TTS. Space out acronyms ("C F O" not "CFO"), spell year-numbers, no stage directions in the text. Target 28–29 seconds so the end card fills out the 30.

Stage 3 is the storyboard — roughly five shots, one arc. Each shot gets a vivid prompt locked to the aspect ratio. Before committing any expensive renders, the skill runs a pick_reference_images pre-flight to surface three style options. Worth doing. Catching a style mismatch before you've queued five video clips saves real money.

Stage 4 is generation, and here's the rule you cannot ignore: call sequentially, never in parallel. There's a global, account-level rate limit of roughly three generation requests per rolling window — shared across image, video, TTS, and music. If you fan out a batch of calls at once, most of them 429. The skill generates keyframes one at a time with POST /api/v1/images (imagen-4 for photoreal), then voice with POST /api/tts, then music with POST /api/music, then video clips one by one with POST /api/v1/video/generate using Kling 3.0. Polling the job status endpoint is a read call — you can poll freely, it doesn't count against the limit. On a 429, back off around 60 seconds and grow the delay from there. If a keyframe keeps rate-limiting, skip it and have Kling generate that shot text-to-video instead. That's what happened on shots 3 and 4 in the TGG build. Not a failure mode — just an alternate path.

Stage 5 is assembly, and there's another gotcha here. Do not rely on ffmpeg's drawtext filter for your end card or any on-screen text. A lot of ffmpeg builds ship without libfreetype, and you'll get "Filter not found" mid-render. Instead, rasterize the end card — brand colors, tagline — as an image using Pillow or generate_image, make it a clip, and composite it that way. The skill normalizes all clips to a common resolution and frame rate first, because video models return slightly different sizes, then concatenates, then mixes: voiceover at full level, music ducked to roughly 0.14 with a tail fade. Voiceover length drives the total timing.

Stage 6 is delivery — the hosted MP4 URL comes back and you put it wherever you keep brand assets.

---

The TGG Accounting reference run

Here's a concrete example so this isn't abstract.

Brief: tgg-accounting.com, 30 seconds, cinematic, AI voiceover plus music, 16:9.

Research pulled "The TGG Way™," "From Chaos to Confidence," "Clarity in your numbers," and a navy-and-gold color palette.

The voiceover script landed at 28.6 seconds through ElevenLabs TTS (default voice, no specific voiceId needed):

> Running a business is hard enough. Your numbers shouldn't make it harder. At T G G Accounting, we turn financial chaos into clarity and confidence. Get a full finance team. C F O, controller, and accountants. For a fraction of the cost. Real-time visibility. Smarter decisions. A clear plan for growth. This is the T G G Way. T G G Accounting. Clarity in your numbers.

Note the spaced acronyms. That's not a stylistic choice — TTS reads "CFO" as a word otherwise.

Music was generated at 32 seconds through POST /api/music with elevenlabs-music: uplifting corporate instrumental, warm piano and soft strings building to optimistic resolve, no vocals.

Five shots, each targeting 6 seconds with Kling 3.0: 1. A stressed business owner at a cluttered desk late at night — keyframed with imagen-4, animated with Kling image-to-video 2. Two professionals in conversation, confident handshake — keyframed, animated 3. Clean financial dashboard, numbers snapping into order — text-to-video (keyframe rate-limited) 4. Business owner reviewing clear reports, relieved — text-to-video 5. Branded end card — navy background, gold TGG logo, tagline — rasterized with Pillow, 3.5-second clip

Total: around $5 in Credits. One hosted MP4.

---

Cost and what to expect

Approximate rates: images are cents each, Kling video is around $0.105 per second, ElevenLabs voice is around $0.005 per second, music is around $0.021 per second. If you're paying with Varosity Credits (versus a BYOK key), there's a 5% markup. BYOK is zero markup. Failed renders are never billed — if a generation fails, you don't pay for it.

A full 30-second spot with five keyframes, five clips, voiceover, and a music bed lands between $4 and $6. That's not a rough estimate — that's what the TGG build actually cost.

The main limit to plan around is that sequential rate. Don't try to batch. If you're generating for multiple brands back to back, space them out. Each full spot takes a while because of polling, but the agent handles that — you're not babysitting it.

---

Installing and running the skill

The skill lives in Varosity's public skill library. To pull it locally:

curl -s https://varosity.ai/api/v1/skills/website-promo-video > ~/.varosity/skills/website-promo-video.md

It's set to auto_update: true, so running refresh_skills at the start of a session pulls the latest version. The skill is maintained by Varosity AI — when the underlying models or API shapes change, the playbook updates.

Authentication is a single Bearer token (vsk_...) — either in the Authorization header for REST calls, or connect the Varosity MCP server once and the agent picks it up from there. You need three scopes: generate:image, generate:video, generate:voice. Billing is automatic against your Credits balance.

To trigger it: tell your agent you want a promo video or brand video for a specific site. The skill's trigger covers anything in the range of "30-second video," "website hero video," "brand teaser," or "explainer" for a company or URL. The agent loads the skill, runs the brief stage, and waits for your confirm before touching any paid calls.

The dependencies are ffmpeg, Pillow, and those three scopes. Nothing exotic. If you're running a local agent that already uses ffmpeg for anything, you're probably already set — just remember the drawtext gotcha and rasterize text as images instead.

That's the whole thing. Research the brand, script it right, generate sequentially, build the end card as an image, mix the audio properly. Thirty seconds of video that actually sounds and looks like the company it's for.