64 models · one account
Every model, one API.
Live prices, limits, and a copy-paste API example for each. Use them in the browser or call them from your app or agent — no provider accounts needed.
Video models
- Veo 3.1Best lip syncbest lip sync · native 4K · synced audio$0.158/s
- Veo 3.1 (Quality)Highest qualitystate-of-the-art fidelity · best lip sync · native 48kHz audio$0.420/s
- Veo 3.1 LiteBudget Veocheapest Veo tier · native 48kHz audio · high-volume iteration$0.053/s
- Kling 3.0 ProBest valuemulti-shot consistency · cinematic motion · dialogue close-ups$0.105/s
- Grok Imagine VideoCheapest w/ audiocheapest frontier video with native audio · synchronized sound + dialogue · flexible 1–15s duration$0.074/s
- Seedance 4.5Audio-nativeunified audio-video generation · multi-shot from one prompt · phoneme-level lip-sync$0.147/s
- Seedance 2.0Top rankedtop-ranked motion + physics · cinematic camera control · native synced audio$0.318/s
- Seedance 2.0 FastBest valuenear-flagship quality at lower cost · faster renders · native synced audio$0.254/s
- OmniHuman 1.5Avatar layertalking avatars · from single photo · lip-sync from audio$0.084/s
- LivePortrait (Performance Transfer)Performance transferperformance transfer from a driving video · transfers head motion + facial expressions + lip movement · onto a different person's photo$0.053/s
- Wan 2.2 Animate (Full-Body Transfer)Full-body (live)FULL-BODY motion transfer from a driving video · transfers whole-body movement + face + expression · onto a different character image$0.105/s
- Kling 3.0 Pro (Replicate)multi-shot consistency · cinematic motion · cheaper than direct$0.095/s
- HappyHorse 1.0Top ranked#1-ranked motion + prompt adherence · joint audio-video generation · multilingual lip-sync$0.147/s
- Luma Ray3 (fal)native 16-bit HDR · high-fidelity motion · strong realism$0.126/s
- Runway Gen-4.5 (Replicate)camera control · motion brush · scene consistency$0.053/s
- Pika 2.5stylized motion · character consistency · fast$0.084/s
- MiniMax Hailuo 02physics realism · complex motion · long prompts$0.116/s
- WAN Video EffectsVarosity Creditsnamed effect catalog (Cakeify, Squish, VHS, Samurai…) · frame consistency · platform-funded$0.063/s
- LatentSync Lip-SyncVarosity Creditssmooth temporal consistency · fast inference · any video + audio$0.053/s
- WAN 2.1 Text-to-VideoVarosity Creditsplatform-funded (no BYOK) · up to 720p / high quality · reliable fallback$0.032/s
- HeyGen Avatar (Digital Twin)studio-grade lip sync · Avatar IV / Avatar V motion engines · voice emotion + speed control$0.084/s
- HeyGen Photo Avatar (Avatar IV)animate ANY photo as the speaker · Avatar IV motion engine · motion prompt + expressiveness control$0.084/s
- HeyGen Cinematic Avatarprompt-driven cinematic shots · blends 1–3 avatar looks into a scene · reference videos/images for style$0.105/s
- HeyGen Video Agentprompt → finished video · agent writes script, picks avatar & scenes · accepts reference files$0.126/s
- HeyGen Video Translatemultilingual lip-sync dubbing · preserves original speaker appearance · supports 40+ languages$0.105/s
- D-ID AI Presentertalking avatars from any photo · text-to-presenter · fast render$0.053/s
- Hunyuan Videoopen-source quality · long coherent motion · strong physics$0.095/s
- LTX VideoFastestfastest open video model (<5s) · image-to-video · good for iteration$0.042/s
- Wan 2.6open-weight (Apache-2.0) · strong motion · self-hostable$0.063/s
- HunyuanVideo 1.5open-weight (Apache-2.0) · long coherent motion · strong physics$0.095/s
- LTX-2Fastest openopen-weight (Apache-2.0) · very fast · image-to-video$0.053/s
- Luma Ray 2Varosity Creditsfluid motion · cinematic quality · strong prompt adherence$0.084/s
- Pika 2.2Varosity Creditsfast generation · stylized output · good character consistency$0.042/s
- Hailuo 02Varosity Creditsphysics realism · complex motion · high resolution$0.084/s
- Runway Gen 4Varosity Creditscamera control · cinematic motion brush · video-to-video$0.011/s
Image models
- FLUX.1 [schnell]fast (1–3s) · good prompt adherence · low cost$0.003/img
- FLUX.2 [pro]current-gen FLUX quality · strong prompt adherence · multi-reference editing$0.032/img
- FLUX.2 [klein]current-gen FLUX at draft price · crisper text than FLUX.1 · fast (4-step)$0.006/img
- Z-Image TurboCheapestultra-cheap (~$0.01) · ~1s render · open-weight$0.011/img
- Seedream 4.0photorealism · strong composition · low cost$0.032/img
- FLUX 1.1 Prohighest-quality Flux · strong prompt adherence · fine detail$0.042/img
- Grok Imagine Imagecheap ($0.02/image) · photorealism · fast$0.021/img
- Nano Banana Pro (Gemini 3 Pro Image)Recommendedbest-in-class legible in-image text · multilingual text · reasoning-driven composition$0.141/img
- Nano Banana 2 (Gemini 3.1 Flash Image)production-scale quality at flash speed · legible in-image text · up to 4K$0.070/img
- Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Fastultra-low latency · cheapest Google tier · high-volume iteration$0.036/img
- Ideogram V3best text-in-image · legible typography · poster / cover art$0.042/img
- Recraft V3long detailed prompts · brand-consistent style · vector-style output$0.042/img
- DALL-E 3strong text-in-image · creative compositions · photorealism$0.042/img
- GPT Image 2best-in-class instruction following · accurate in-image text · photorealistic detail$0.084/img
- Nano Bananavery fast · good identity preservation · cheap$0.005/img
- Flux DevVarosity Credits12B parameter model · strong prompt adherence · fast guided distillation$0.016/img
Voice models
- ElevenLabs Multilingual v229 languages · emotion control · voice cloning$0.315/min
- ElevenLabs v3 (Alpha)dialogue style · highest expressiveness$0.504/min
- OpenAI TTS-1fast · 6 voices · low latency$0.252/min
- OpenAI TTS-1 HDHD qualityhighest OpenAI voice quality · 6 voices · natural prosody$0.504/min
- Deepgram Aura 2Lowest latencyultra-low latency · natural prosody · cheap ($0.030/1K chars)$0.126/min
- Cartesia Sonic 2ultra-low latency (~90ms) · natural prosody · large public voice library$0.378/min
- Cartesia Sonic 3~90ms TTFA · 42 languages · AI laughter + emotion$0.378/min
- Kokoro (fal)open-weight (MIT) · runs on CPU · very cheap$0.050/min
- Chatterbox (fal)open-weight (MIT) · expressive + emotion control · instant voice cloning$0.063/min
- Fish Audiomultilingual · cheap · large community voice library$0.252/min