Run a Virtual Influencer End to End for Under $4
The Virtual Influencer skill in Varosity's public library stands up a brand-owned, AI-disclosed synthetic creator account — one consistent face, a real voice, a batch of identity-locked feed posts with captions, a talking Reel, and a posting calendar — in a single agentic session. It's the content-engine layer, not just a one-off render. The persona it produces is a named, recognizable character that can post week after week without drifting.
I kept running into the same wall. A brand wants a recurring face — think Lil Miquela or Lu do Magalu, but theirs — and I'd hand them a great portrait. Then post two looks subtly off. Post five looks like a cousin. By post ten, the "character" is gone. The problem was never one good render. It was holding identity across a hundred posts without a trained model or a six-figure production budget.
What I actually wanted was a portable, ownable identity mechanic that routes each post to whatever model holds the face best, with no vendor lock-in. That's what this skill is built around.
---
The identity-locking mechanic — this is the whole thing
None of the generation models in Varosity expose a seed you can replay. So reproducibility doesn't come from seeds. It comes from a single saved reference image URL that you thread into every subsequent render.
Here's exactly how the skill runs it:
1. Canonical face render. The agent calls POST /api/v1/images via the gemini-3-pro-image model (Nano Banana Pro) with a tight portrait prompt — 1:1, photoreal, neutral pose, even light, describing every invariant: face shape, hair, eye color, signature features. You pick one from three candidates and that image URL becomes persona.json → canonical_face_url. That URL is the identity anchor from here forward. It's not a style reference. It's the face.
2. Feed posts via nano-banana. Every feed image is POST /api/v1/images with referenceImageUrl set to that canonical URL. Nano-banana is the model here — it's cheap (around $0.005 per image) and it has the best identity preservation when you're doing image-to-image off a reference. Each prompt restates the invariants: "same face, same hair, same [signature feature]" and describes the scene. The scene changes. The face doesn't.
The critical rule: never use a previous post as the reference for the next one. Always go back to the canonical. Drift compounds. One generation off a generation and you're two steps from the original. Four posts in, you have a different person.
3. Voice via ElevenLabs TTS. One voice ID, one persona, always. The agent calls POST /api/v1/tts with the voiceover script to get the audio clip. Same voice ID on every Reel. That voiceId goes into persona.json alongside the face URL.
4. Talking Reel via HeyGen photo avatar. POST /api/v1/video/generate with referenceImageUrl pointing to the canonical face and providerOptions.audioUrl pointing to the TTS clip. HeyGen animates the face to the voice. Keep it 9:16, frontal, short — face wobble gets worse with big motion. Up to 300 seconds is supported but shorter is safer for a first Reel.
5. Optional B-roll via Kling 3.0. If you want motion from a feed still — a post image that slowly pulls back, or a lifestyle scene that moves — that's POST /api/v1/video/generate with the post image as referenceImageUrl. Image-to-video, cinematic.
6. Optional music bed via ElevenLabs Music. POST /api/music. Reel needs a track, you don't want to deal with licensing — generate one that fits the persona's vibe.
7. Persona bible + captions + calendar. This is your work, not the model's. The agent uses web research (WebSearch + WebFetch) to pull brand voice and niche trends, then drafts the bible and caption templates. But you write the real voice. You set the cadence. You decide what this person actually says.
---
A worked example: "Sable" — a sustainable fashion account
Brand brief: eco fashion label, wants a persona in the 25–34 sustainable-style space. Name: Sable. Look: South Asian features, natural hair, minimalist aesthetic. Vibe: thoughtful, dry humor, never preachy.
Session runs like this:
- Agent researches sustainable fashion content pillars and the brand's own tone from their site.
- Drafts a persona bible: Sable's backstory (fictional, labeled AI), her content pillars (secondhand hauls, fabric transparency, capsule styling), her caption voice, her disclosure line ("Virtual creator · AI" in bio,
#AIcreatedin every post). - Calls Nano Banana Pro three times to generate three canonical face candidates. You pick one. That URL is saved.
- Calls nano-banana five times with that canonical URL as the reference — five feed posts in different scenes: a thrift store, a rooftop, a flat-lay. Each prompt restates her face invariants. Each comes back with her face.
- Agent drafts captions for all five posts in Sable's voice, with the disclosure tag baked in.
- Calls ElevenLabs TTS with a 30-second voiceover script Sable would deliver to camera.
- Calls HeyGen photo avatar with the canonical face and that audio clip to get the talking Reel.
- Writes a four-week posting calendar: cadence, content themes, which posts go Stories vs. feed vs. Reels.
- Publishes samples to the Varosity gallery via the showcase manifest for the client handoff link.
Total cost for that launch kit: somewhere between $1.50 and $4.00 in Varosity Credits depending on how many face candidates you render, how many posts, and whether you add Kling B-roll. The HeyGen talking Reel at around 15 seconds runs about $1.20. Nano-banana feed posts at ~$0.005 each are basically free. Failed renders aren't billed.
---
What to watch out for
Rate limits. There's a global account-level generation rate limit of roughly three requests per rolling 60 seconds. The agent handles this by running renders sequentially and backing off on a 429 — 60 seconds, then 90, then 120. Polling GET /api/v1/jobs for job status is a free read, so the agent can watch a render finish without burning requests.
Text baked into images. Don't prompt the handle, caption, or brand name into the image itself. Text generation in image models is unreliable and you'll get garbled glyphs. Captions and handles live in the platform's text layer. Overlays are HTML/CSS. The image is just the image.
The disclosure line is not optional. Meta and TikTok both require AI-content labeling. FTC requires clear disclosure for any sponsored or endorsement content. If you're building this for a client and the disclosure gets dropped in production, that's a policy violation and potentially a legal one. Build it into the caption template so it cannot be forgotten. AI label in the bio. AI tag on every post. Every time.
The out-of-scope cases. This skill is for disclosed brand creators. A covert fake-person account, a romance persona, anything designed around "is she real?" ambiguity as a monetization tactic — that's out of scope. The skill won't help you build that, and I'd argue no tool should.
No seed means no magic replay. If you lose persona.json, you lose the reference URLs, and you lose the identity. Persist every canonical URL, every post URL, every voiceId, and every prompt to that file. That file is the persona.
---
How to install and run it
The skill lives in Varosity's public library. To pull it manually:
curl -s https://varosity.ai/api/v1/skills/virtual-influencer > ~/.varosity/skills/virtual-influencer.md
It's set to auto_update: true, so if you call refresh_skills at the start of a session your agent picks up the latest version automatically. The skill requires generate:image, generate:video, and generate:voice scopes on your API key. One key. All models.
To run it, give the agent a brand brief: the persona name, niche, look invariants, target platforms, how many posts in the v1 drop, whether the persona talks, and the disclosure handle and bio language. The skill derives the persona bible, every image prompt, and every caption from that input. You review and approve at the persona-bible stage before any renders fire.
Variety covers all the model calls — Gemini image, nano-banana, ElevenLabs TTS, HeyGen, Kling, ElevenLabs Music — under one API key, billed in Varosity Credits. No juggling five separate accounts or rate limits across vendors.
The output is a folder: persona.json (the identity manifest), the canonical face image, all feed post images, all captions, the TTS clip, the talking Reel, optional B-roll, optional music, the posting calendar, and the showcase gallery link for client handoff.
That's the launch kit. A disclosed AI creator account, ready to post, built in one session.