kling-v3
Kling V3 general video model, supports text + up to 4 reference images, suitable for stable short video clips and daily creative workflows.
Kling V3 general video model, supports text + up to 4 reference images, suitable for stable short video clips and daily creative workflows.
Overview
| Field | Value |
|---|---|
| Model ID | kling-v3 |
| CLI | dlazy kling-v3 |
| MCP | kling-v3 (Claude Code surfaces this as mcp__dlazy__kling-v3) |
| Type | video |
| Execution | Async task; the CLI polls until completion (--no-wait returns generateId immediately) |
| Batch | Supports --batch <n> parallel fan-out |
Parameters
| Arg | Type | Required | Notes |
|---|---|---|---|
prompt | string | Yes | Prompt |
generation_mode | "frames" | "components" | No | Generation Mode(frames=Frames; components=Components); default "frames" |
images | array<url> | No | Images; image (accepts URL, local path, or data: URL); max 4 items |
subjects | array<string> | No | Subjects; max 3 items; only when !(generation_mode="frames") |
aspect_ratio | "16:9" | "9:16" | "1:1" | No | Aspect Ratio; default "16:9" |
duration | "3" | "4" | "5" | "6" | "7" | "8" | "9" | "10" | "11" | "12" | "13" | "14" | "15" | No | Duration (s); default "5" |
mode | "std" | "pro" | No | Mode; default "std" |
sound | string | No | Sound Effect; default false |
--input @file.jsonor--input '{...}'can supply all args at once; flags take precedence over--inputkeys.
CLI Examples
dlazy kling-v3 --help
dlazy kling-v3 --prompt "Write your prompt here" --images "https://example.com/reference1.jpg" "https://example.com/reference2.jpg"
dlazy kling-v3 --prompt "Write your prompt here" --images "./local-image.png"
dlazy kling-v3 --prompt "Write your prompt here" --images "https://example.com/reference1.jpg" "https://example.com/reference2.jpg" --dry-run
dlazy kling-v3 --prompt "Write your prompt here" --images "https://example.com/reference1.jpg" "https://example.com/reference2.jpg" --no-wait
dlazy kling-v3 --prompt "Write your prompt here" --images "https://example.com/reference1.jpg" "https://example.com/reference2.jpg" --batch 4Compose with a pipeline
dlazy gpt-image-2 --prompt "reference visual" \
| dlazy kling-v3 --images - --prompt "Write your prompt here"MCP
MCP is consumed by AI clients, not handwritten. Once dLazy is registered as an MCP server the tool appears in the client's tool list — Claude Code surfaces it as mcp__dlazy__kling-v3; generic clients (e.g. OpenClaw) call it as kling-v3.
See MCP setup for how to add the server.
Output
{
"outputs": [
{ "type": "video", "id": "o_...", "url": "https://files.dlazy.com/result.mp4", "mimeType": "video/mp4" }
]
}Media tools emit image / video / audio / file outputs. Use --output url to print only the URLs on stdout.
jimeng-omnihuman-1.5
Jimeng digital human model: combines any-ratio character/subject image with audio to generate high-quality digital human videos.
kling-v3-omni
Kling Omni video model, supports multiple reference images, duration, mode (std/pro), and optional audio. Suitable for highly controlled video synthesis tasks.