dLazy AIdLazy AI
Model ReferenceVideo Models

Video Models

Browse video models exposed by @dlazy/cli.

  • happyhorse-1.0 — Happy Horse 1.0 video model — one model covers text-to-video (t2v), first-frame-to-video (i2v), reference-to-video (r2v), and video editing (edit). The selected mode is automatically routed to the matching sub-model.
  • heygen-lipsync-speed — HeyGen Lipsync Speed: Fast lip-sync model, ideal for scenarios requiring rapid generation.
  • jimeng-dream-actor — Jimeng character/action-driven video model, supports reference image and video input, suitable for character acting, action transfer, and style-consistent generation.
  • jimeng-i2v-first — Jimeng first-frame-to-video model, uses first frame + text to generate video. Suitable for single-shot scenes that naturally animate static images.
  • jimeng-i2v-first-tail — Jimeng first/last-frame video model; constrains shot start/end frames. Good for transitions and clearly resolved action.
  • jimeng-omnihuman-1.5 — Jimeng digital human model: combines any-ratio character/subject image with audio to generate high-quality digital human videos.
  • kling-v3 — Kling V3 general video model, supports text + up to 4 reference images, suitable for stable short video clips and daily creative workflows.
  • kling-v3-omni — Kling Omni video model, supports multiple reference images, duration, mode (std/pro), and optional audio. Suitable for highly controlled video synthesis tasks.
  • pixverse-c1 — PixVerse C1 video model (strong on action, VFX, and high-motion scenes) — one model covers text-to-video, image-to-video, first/last-frame-to-video, and reference-to-video: t2v when no images, i2v with first frame only, kf2v with first+last frames, r2v with reference images.
  • seedance-2.0 — ByteDance's latest video generation model. Supports multi-modal reference (images, video, audio) to generate videos, as well as first/last frame and text-to-video modes.
  • seedance-2.0-fast — Fast version of ByteDance's Seedance 2.0. Generates videos faster with support for multi-modal references, first/last frame, and text-to-video.
  • sync-lipsync-3 — fal.ai sync-lipsync v3 — given an input video and audio, generate a new video where the speaker's lip movement matches the audio. Good for dubbing, localization, and re-syncing virtual presenters.
  • veo-3.1 — High-quality video generation model, supports text-to-video and single-image-driven video. Suitable for ad shorts and cinematic sequences (slower speed, higher quality).
  • veo-3.1-fast — Fast video generation model, supports text-to-video and single/multi-image/first-last frame driven. Suitable for time-sensitive previews and rapid iterations.
  • video-audio-split — Audio/video split tool: uses ffmpeg to split a video into a silent video track and a standalone audio track, suitable for independent editing.
  • video-replicate — Video replicate tool: extracts the first frame and audio from the source video, runs video understanding for a prompt, and returns a Seedance 2.0 replicate bundle (first frame + audio + video).
  • videoretalk — Tongyi VideoRetalk lip sync / lip-sync (mouth sync, dubbing) video model — takes a talking-person video plus a voice audio track and regenerates the video so the speaker's mouth/lips match the new audio. Use this for lip syncing a person video to new speech. Optionally provide a reference face image to pick the target person when the video contains multiple faces.
  • videoseg — Video human segmentation tool: invokes Aliyun's async SegmentVideoBody and returns a same-length black/white mask video, suitable for downstream compositing or matting.
  • viduq2-i2v — Vidu image-to-video model, supports reference image-driven video, duration/resolution/ratio, and audio settings, suitable for image animation and short clips.
  • wan2.7 — Tongyi Wanxiang 2.7 video model — one model covers text-to-video, first/last-frame-to-video, and reference-to-video: uses text-to-video when no images are provided, first/last-frame-to-video when frames are provided, and reference-to-video when reference images are supplied.