Skill detail

aliyun-kling-video

Generates videos with Kling v3 models on DashScope covering text-to-video, image-to-video, reference-to-video, smart storyboard, and video editing.

MatchDirectReviewed for Video Generation
Sourcecinience/​alicloud-skillsExternal source
Reported installs72Popularity signal only

Inspect before use

Automated review checks relevance, not safety or endorsement. Read the source instructions before using this skill.

Saved source preview

SKILL.md

The saved excerpt is a snapshot from review. The external source remains the complete and most current version.

---
name: aliyun-kling-video
description: Use when generating videos with Kling v3 models on DashScope (kling/kling-v3-video-generation, kling/kling-v3-omni-video-generation). Use when implementing text-to-video, image-to-video, reference-to-video, smart storyboard, or video editing via the video-synthesis async API.
---

# Kling V3 Video Generation

## Validation

```bash
mkdir -p output/aliyun-kling-video
python -m py_compile skills/ai/video/aliyun-kling-video/scripts/generate_kling_video.py && echo "py_compile_ok" > output/aliyun-kling-video/validate.txt
```

Pass criteria: command exits 0 and `output/aliyun-kling-video/validate.txt` is generated.

## Output And Evidence

- Save task IDs, polling responses, and final video URLs to `output/aliyun-kling-video/`.
- Keep at least one end-to-end run log for troubleshooting.

## Prerequisites

- Install dependencies (recommended in a venv):

```bash
python3 -m venv .venv
. .venv/bin/activate
python -m pip install requests
```
- Set `DASHSCOPE_API_KEY` in your environment (must be Beijing region API Key).
- Enable Kling in [百炼控制台](https://bailian.console.aliyun.com/cn-beijing/?tab=model#/model-market/all) — search "kling" and activate.

## Critical model names

- `kling/kling-v3-video-generation` — standard model: t2v, i2v (first frame, first+last frame)
- `kling/kling-v3-omni-video-generation` — omni model: adds reference-to-video, video editing, multi-subject references

## Capabilities

| Capability | Model | Required media |
|---|---|---|
| Text-to-video | both | none |
| Smart storyboard (multi-shot) | both | none (use `multi_prompt`) |
| Image-to-video (first frame) | both | `first_frame` |
| Image-to-video (first+last frame) | both | `first_frame` + `last_frame` |
| Reference-to-video | omni only | `refer` and/or `feature` |
| Video editing | omni only | `base` + optional `refer` |

## API endpoint (async only)

```
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/video-generation/video-synthesis
```

Required headers:
- `Authorization: Bearer $DASHSCOPE_API_KEY`
- `Content-Type: application/json`
- `X-DashScope-Async: enable`

**Region**: Beijing only. No Singapore endpoint.

## Normalized interface

### Request (input)
- `prompt` (string, conditional) — up to 2500 characters. Required for `shot_type=intelligence`. For omni reference-to-video, use `<<<element_1>>>`, `<<<image_1>>>`, `<<<video_1>>>` to reference media.
- `negative_prompt` (string, optional) — content to exclude
- `media` (array, optional) — media objects with `type` and `url`:
  - Standard model types: `first_frame`, `last_frame`
  - Omni model types: `first_frame`, `last_frame`, `refer`, `base`, `feature`
- `multi_shot` (boolean, optional) — enable multi-shot generation (default: false)
- `shot_type` (string, conditional) — `intelligence` (AI auto-split) or `customize` (manual). Required when `multi_shot=true`.
- `multi_prompt` (array, optional) — per-shot prompts when `shot_type=customize`
- `element_list` (array, optional) — multi-subject element images (omni model only)
- `keep_original_sound` (string, optional) — `no` (default) or `yes`, for videos (omni model only)

### Request (parameters)
- `mode` (string, optional) — `pro` (default, 1080P) or `std` (720P)
- `aspect_ratio` (string, conditional) — `16:9` (default), `9:16`, `1:1`. Required for t2v and reference-to-video.
- `duration` (integer, optional) — video length [3, 15] seconds (default: 5). When using reference video, [3, 10].
- `audio` (boolean, optional) — generate audio (default: false). Affects pricing.
- `watermark` (boolean, optional) — add "可灵 AI" watermark (default: false)

### Media input limits

**Images** (first_frame, last_frame, refer):
- Formats: JPEG, JPG, PNG (no transparency)
- Resolution: [300, 8000] pixels per side
- Max size: 10MB

**Videos** (base, feature):
- Formats: mp4, mov
- Duration: 3-10s
- Resolution: [720, 2160] pixels per side
- Frame rate: 24-60 fps
- Max size: 200MB

### Media combination rules

**kling/
Read the full source on GitHub (opens external page)
Context

Related work