VidModel API Platform

Video & Image Generation Models. One API.

Access production-ready video and image models — from industry-known names to high-performance specialists — through a single, credit-based API.

Featured

Popular Models

Most popular models across all categories

WAN

Wan

Alibaba's Wan video models — open-source DiT architecture with 86%+ VBench scores, covering image-to-video, text-to-video, reference-to-video, and instruction-based video editing across Wan 2.6 and 2.7

WAN

Wan 2.7 Image to Video

Alibaba's flagship I2V model — up to 15s at 1080p with 9-grid multi-angle input, first+last frame guidance, and physics-accurate motion from a single source image.

Quickstart

Ship in minutes,
not days.

Two endpoints, one API key. POST to create a task, GET to retrieve the result — the same pattern across every model type.

One key, every model

Your API key works across video generation, characters, face swap, and image editing — no separate accounts per provider.

Async task lifecycle

POST to create a task, GET to poll the result. The same two-step pattern works for every model type.

Predictable response shape

Every task returns taskId, taskStatus, and a videoUrl or imageUrls on success. No undocumented fields.

const BASE = "https://api.vidmodel.ai"
const HEADERS = {
  "Authorization": `Bearer ${process.env.VIDMODEL_API_KEY}`,
  "Content-Type": "application/json",
}

// 1. Create task
const { taskId } = await fetch(
  `${BASE}/openapi/v1/aigc/video-generation/tasks`,
  {
    method: "POST",
    headers: HEADERS,
    body: JSON.stringify({
      model: "vidgen1.0-t2v",
      prompt: "A cinematic shot of a girl walking in the rain",
      promptEnhance: true,
      outputAudio: false,
    }),
  }
).then((r) => r.json())

// 2. Query result
const { videoUrl } = await fetch(
  `${BASE}/openapi/v1/tasks/${taskId}`,
  { headers: HEADERS }
).then((r) => r.json())
Use Cases

What teams ship with VidModel

Three production workflows, each handled by a single API call.

vidgen-i2v
PortraitJPG / PNG
vidgen-i2v
SpokespersonVideo output
E-commerce

Product spokesperson videos

Animate any portrait into a natural talking video. Pass an image URL — receive a production-ready clip. No studio, no shoot.

vidvox-flash
Text PromptCharacter brief
vidvox-flash
Batch OutputHigh volume
Creator Tools

Personalized content at scale

Generate character video in parallel across thousands of users. Flash mode delivers full throughput without sacrificing visual realism.

vidvox-r2v
References2–4 photos
vidvox-r2v
Character SceneHD video
Marketing

Multi-reference character scenes

Blend subjects from separate photos into one coherent scene. R2V reads all references and synthesizes them into high-fidelity video.

Why choose us

One API for every video and image model

VidModel aggregates the best video and image models — from well-known names to high-performance specialists — under a single API key. No separate contracts, no mismatched task lifecycles.

Curated model selection

Access flagship names and high-performing specialists — all production-tested and available under a single API key, with no separate contracts per provider.

Organized by use case

Models are grouped by what you build — video generation, talking characters, AI personas, image editing, and face swap — not just by vendor name.

Granular output controls

Set resolution, duration, aspect ratio, reference faces, and style per request without managing separate vendor integrations.

Production-grade reliability

Typed responses, consistent output shapes, and usage visibility designed for teams shipping media features at scale.

Platform signals

Ship media features with fewer moving parts.

One API key covers every model type. Teams get a unified task lifecycle, consistent output shapes, and a single billing surface — so the focus stays on building, not integrating.

5media pipelines
20smax video output
1 keyall model types
Customer stories

Teams shipping creative AI features faster

From creator tools to production media pipelines, teams use VidModel to consolidate video generation, character creation, and image editing into a single integration.

"

Switching to VidModel unified our entire media pipeline — video generation, face swap, and character creation now run through a single integration.

AR

Ari Renn

Head of Product · ClipForge

VIDEO PLATFORM

We cut our provider count from four to one. VidGen's human-optimized output is the only video generation model our team needs.

MP

Mina Park

VP Engineering

CREATOR STUDIO

The model catalog makes it easy for non-ML teammates to pick the right pipeline for each content type.

TW

Theo Walsh

Creative Lead

EDTECH APP

Talking character generation and face swap now share one workflow. A feature that used to take months shipped in a week.

LC

Leila Chen

Founder

3.8×faster time to ship
60%less integration work
99.9%API availability
Pricing

Pay per generation. No subscriptions.

Buy credits once, spend them across any model. Every generation shows the exact cost before you run it — no surprises.

FASTEST

VidVox Flash

VidVox's realism at the highest generation speed — built for rapid iteration and high-volume pipelines.

$0.25/ 5s
720p$0.25 / 5s
1080p$0.38 / 5s
PORTRAIT SPECIALIST

VidGen 1.0

Near-perfect human body structure and facial accuracy — the go-to for portrait and character video.

$0.50/ 5s
Audio off$0.50 / 5s
Audio on+$0.10 / 5s
SPEED + REALISM

VidVox 1.0

Fast generation paired with high visual realism — a rare combination across both standard and R2V modes.

$0.50/ 5s
720p$0.50 / 5s
1080p$0.75 / 5s
R2V 720p$0.50 + $0.50 ref
See full pricing →Face Swap, Image Edit & AI Character models — all credit-based, full rates on pricing page

All Models

Browse all video generation, character, image editing, and face swap models in one compact index.

VidGen 1.0 Image to Video

Animate images into video with accurate human body structure and natural movement — strong in complex action scenes with background audio support.

VIDEO GEN
VidGen 1.0 Video Extend

Extend VidGen-generated video up to 20 seconds while preserving body structure, character consistency, and scene continuity.

VIDEO GEN
VidGen 1.0 Text to Video

Generate video from text prompts with exceptional human body structure and portrait accuracy — ideal for complex action scenes and audio-rich content.

VIDEO GEN
VidVox 1.0 Reference to Video

Blend multiple images into a video, replace characters, or change actions in video.

VIDEO GEN
VidVox 1.0 Image to Video

Animate a single portrait image into realistic talking-character video — up to 15s HD per generation, with fast output and lifelike results.

VIDEO GEN
VidVox 1.0 Flash Image to Video

The fastest VidVox variant — generate talking-character video with lifelike realism at significantly reduced generation time.

VIDEO GEN
FaceSync Character Generator

Generate AI characters with high facial identity preservation — up to 95%+ similarity to the uploaded reference face.

CHARACTER
StyleSync Character Generator

Generate AI characters that preserve the reference face's original style and hairstyle with strong face similarity (~85%).

CHARACTER
SkinReal Character Generator

Generate full-body AI character images with natural skin texture — less smoothing for a more realistic, photographic result.

CHARACTER
PortraitLive Character Generator

Generate detailed half-body portrait images with vivid expressions and high facial clarity.

CHARACTER
VidGen Image Edit

Prompt-driven image editing powered by VidGen — best for removing objects or subjects from complex, layered scenes.

IMAGE EDIT
VidVox Image Edit

Prompt-driven image editing powered by VidVox — best for preserving fine details and rendering sharp, legible text in the output.

IMAGE EDIT
VidVox Pro Image Edit

The Pro tier of VidVox image editing — enhanced detail rendering, sharper text, and higher-fidelity edits for production-quality output.

IMAGE EDIT
Seedance 2.0 Image to Video

Animate still images into cinematic video with physics-accurate motion, native audio sync, and up to 4K resolution — powered by ByteDance.

SEEDANCE
Seedance 2.0 Text to Video

Generate cinematic video from text prompts with automatic camera planning, multimodal reference support, and native audio sync — up to 4K.

SEEDANCE
Seedance 2.0 Reference to Video

Multi-reference video generation with strong character and style consistency — supply images, video clips, and audio to anchor identity across shots.

SEEDANCE
Loading more…