Model ReferenceVideo Models
Video Models
Browse video models exposed by @dlazy/cli.
happyhorse-1.0— Happy Horse 1.0 video model — one model covers text-to-video (t2v), first-frame-to-video (i2v), reference-to-video (r2v), and video editing (edit). The selected mode is automatically routed to the matching sub-model.heygen-lipsync-speed— HeyGen Lipsync Speed: Fast lip-sync model, ideal for scenarios requiring rapid generation.jimeng-dream-actor— Jimeng character/action-driven video model, supports reference image and video input, suitable for character acting, action transfer, and style-consistent generation.jimeng-i2v-first— Jimeng first-frame-to-video model, uses first frame + text to generate video. Suitable for single-shot scenes that naturally animate static images.jimeng-i2v-first-tail— Jimeng first/last-frame video model; constrains shot start/end frames. Good for transitions and clearly resolved action.jimeng-omnihuman-1.5— Jimeng digital human model: combines any-ratio character/subject image with audio to generate high-quality digital human videos.kling-v3— Kling V3 general video model, supports text + up to 4 reference images, suitable for stable short video clips and daily creative workflows.kling-v3-omni— Kling Omni video model, supports multiple reference images, duration, mode (std/pro), and optional audio. Suitable for highly controlled video synthesis tasks.pixverse-c1— PixVerse C1 video model (strong on action, VFX, and high-motion scenes) — one model covers text-to-video, image-to-video, first/last-frame-to-video, and reference-to-video: t2v when no images, i2v with first frame only, kf2v with first+last frames, r2v with reference images.seedance-2.0— ByteDance's latest video generation model. Supports multi-modal reference (images, video, audio) to generate videos, as well as first/last frame and text-to-video modes.seedance-2.0-fast— Fast version of ByteDance's Seedance 2.0. Generates videos faster with support for multi-modal references, first/last frame, and text-to-video.sync-lipsync-3— fal.ai sync-lipsync v3 — given an input video and audio, generate a new video where the speaker's lip movement matches the audio. Good for dubbing, localization, and re-syncing virtual presenters.veo-3.1— High-quality video generation model, supports text-to-video and single-image-driven video. Suitable for ad shorts and cinematic sequences (slower speed, higher quality).veo-3.1-fast— Fast video generation model, supports text-to-video and single/multi-image/first-last frame driven. Suitable for time-sensitive previews and rapid iterations.video-audio-split— Audio/video split tool: uses ffmpeg to split a video into a silent video track and a standalone audio track, suitable for independent editing.video-replicate— Video replicate tool: extracts the first frame and audio from the source video, runs video understanding for a prompt, and returns a Seedance 2.0 replicate bundle (first frame + audio + video).videoretalk— Tongyi VideoRetalk lip sync / lip-sync (mouth sync, dubbing) video model — takes a talking-person video plus a voice audio track and regenerates the video so the speaker's mouth/lips match the new audio. Use this for lip syncing a person video to new speech. Optionally provide a reference face image to pick the target person when the video contains multiple faces.videoseg— Video human segmentation tool: invokes Aliyun's async SegmentVideoBody and returns a same-length black/white mask video, suitable for downstream compositing or matting.viduq2-i2v— Vidu image-to-video model, supports reference image-driven video, duration/resolution/ratio, and audio settings, suitable for image animation and short clips.wan2.7— Tongyi Wanxiang 2.7 video model — one model covers text-to-video, first/last-frame-to-video, and reference-to-video: uses text-to-video when no images are provided, first/last-frame-to-video when frames are provided, and reference-to-video when reference images are supplied.
viduq2-t2i
Vidu image model with text + reference image, ratio, and resolution control. Good for character art, covers, and high-res output.
happyhorse-1.0
Happy Horse 1.0 video model — one model covers text-to-video (t2v), first-frame-to-video (i2v), reference-to-video (r2v), and video editing (edit). The selected mode is automatically routed to the matching sub-model.