
VidVox × VidModel
VidVox delivers fast generation and lifelike visual realism — produce character video, blend reference images, or transfer motion across subjects. Access all VidVox models through VidModel's unified API.
Fast generation and lifelike realism, across four character video modes
VidVox is built for speed and visual realism. Pick the model by input type — portrait image, reference frames, or text prompt — and generate high-quality character video using the same API pattern across every task.
6
included models
30s
max video output
Image + Text + Ref
input modes
Included models
Each model handles a distinct generation task. Pick by input type and the output speed you need.
VidVox 2.0 Reference to Video Turbo
Same quality as VidVox 2.0, with faster generation.
VidVox 2.0 Reference to Video
30s HD per gen with lipsync — best for talking & audio scenes
VidVox 2.0 Text to Video Turbo
Same quality as SoulOmni 2.0, with faster generation.
VidVox 2.0 Text to Video
30s HD per gen with lipsync — best for talking & audio scenes
VidVox 2.0 Image to Video Turbo
Same quality as SoulOmni 2.0, with faster generation.
VidVox 2.0 Image to Video
30s HD per gen with lipsync — best for talking & audio scenes
Pick the right model
Four models, one API pattern — pick by input type and the speed you need.
| Model | Input | Speed | Best For |
|---|---|---|---|
| Portrait photo | Standard | Talking personas, product faces | |
| Reference images | Standard | Multi-character scenes, motion transfer | |
| Portrait batch | Fast | High-volume pipelines, personalization | |
| Text prompt | Standard | Script-to-character, no photo needed |
Match each VidVox model to a real character workflow
Four models, one API pattern — pick by input type and the speed you need.
Talking product personas
Animate a brand face or spokesperson portrait into lifelike talking-character video for product pitches, feature highlights, or brand introductions. VidVox I2V is built for speed and visual realism — the output character moves expressively rather than rigidly, with lip movement that reads as natural at HD resolution.
Multi-reference character scenes
VidVox R2V accepts multiple reference images and three distinct modes of operation: blend separate character photos into a shared scene, extract and transfer motion from a reference clip, or swap a subject into existing video. One API call handles all three — switch between modes by passing referenceImages or referenceVideo.
High-volume content batches
VidVox Flash cuts generation time significantly compared to the standard I2V tier while preserving the same character quality and lifelike output. Kick off tasks concurrently with Promise.all and poll them in parallel — the only practical limit is your account rate cap, not sequential generation time.
Script-to-character video
Describe a character and scene in a prompt to generate talking-character video with no portrait or reference image required. VidVox T2V is the fastest way to prototype a character — define appearance, personality, and movement in a prompt brief, and use the output to validate the concept before sourcing real portrait photos.
Built for speed
VidVox generates output significantly faster than general-purpose video models — built for rapid iteration and high-volume production pipelines.
Multi-modal inputs
Drive video from portrait images, reference clips, or text prompts — one API pattern handles every VidVox generation type.
15-second HD output
Generate lifelike HD character video in one request — fast output with high visual fidelity, ready for production delivery.