Detalle del Skill

aliyun-kling-video

Genera videos con los modelos Kling v3 en DashScope, abarcando texto a video, imagen a video, referencia a video, storyboard inteligente y edición de video.

CoincidenciaDirectaRevisado para Generación de vídeo
Fuentecinience/​alicloud-skillsFuente externa
Instalaciones reportadas72Solo señal de popularidad

Revisar antes de usar

La revisión automática comprueba relevancia, no seguridad ni respaldo. Lee las instrucciones de la fuente antes de usar este Skill.

Vista previa guardada

SKILL.md

Este extracto es una copia guardada durante la revisión. La fuente externa contiene la versión completa y actual.

---
name: aliyun-kling-video
description: Use when generating videos with Kling v3 models on DashScope (kling/kling-v3-video-generation, kling/kling-v3-omni-video-generation). Use when implementing text-to-video, image-to-video, reference-to-video, smart storyboard, or video editing via the video-synthesis async API.
---

# Kling V3 Video Generation

## Validation

```bash
mkdir -p output/aliyun-kling-video
python -m py_compile skills/ai/video/aliyun-kling-video/scripts/generate_kling_video.py && echo "py_compile_ok" > output/aliyun-kling-video/validate.txt
```

Pass criteria: command exits 0 and `output/aliyun-kling-video/validate.txt` is generated.

## Output And Evidence

- Save task IDs, polling responses, and final video URLs to `output/aliyun-kling-video/`.
- Keep at least one end-to-end run log for troubleshooting.

## Prerequisites

- Install dependencies (recommended in a venv):

```bash
python3 -m venv .venv
. .venv/bin/activate
python -m pip install requests
```
- Set `DASHSCOPE_API_KEY` in your environment (must be Beijing region API Key).
- Enable Kling in [百炼控制台](https://bailian.console.aliyun.com/cn-beijing/?tab=model#/model-market/all) — search "kling" and activate.

## Critical model names

- `kling/kling-v3-video-generation` — standard model: t2v, i2v (first frame, first+last frame)
- `kling/kling-v3-omni-video-generation` — omni model: adds reference-to-video, video editing, multi-subject references

## Capabilities

| Capability | Model | Required media |
|---|---|---|
| Text-to-video | both | none |
| Smart storyboard (multi-shot) | both | none (use `multi_prompt`) |
| Image-to-video (first frame) | both | `first_frame` |
| Image-to-video (first+last frame) | both | `first_frame` + `last_frame` |
| Reference-to-video | omni only | `refer` and/or `feature` |
| Video editing | omni only | `base` + optional `refer` |

## API endpoint (async only)

```
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/video-generation/video-synthesis
```

Required headers:
- `Authorization: Bearer $DASHSCOPE_API_KEY`
- `Content-Type: application/json`
- `X-DashScope-Async: enable`

**Region**: Beijing only. No Singapore endpoint.

## Normalized interface

### Request (input)
- `prompt` (string, conditional) — up to 2500 characters. Required for `shot_type=intelligence`. For omni reference-to-video, use `<<<element_1>>>`, `<<<image_1>>>`, `<<<video_1>>>` to reference media.
- `negative_prompt` (string, optional) — content to exclude
- `media` (array, optional) — media objects with `type` and `url`:
  - Standard model types: `first_frame`, `last_frame`
  - Omni model types: `first_frame`, `last_frame`, `refer`, `base`, `feature`
- `multi_shot` (boolean, optional) — enable multi-shot generation (default: false)
- `shot_type` (string, conditional) — `intelligence` (AI auto-split) or `customize` (manual). Required when `multi_shot=true`.
- `multi_prompt` (array, optional) — per-shot prompts when `shot_type=customize`
- `element_list` (array, optional) — multi-subject element images (omni model only)
- `keep_original_sound` (string, optional) — `no` (default) or `yes`, for videos (omni model only)

### Request (parameters)
- `mode` (string, optional) — `pro` (default, 1080P) or `std` (720P)
- `aspect_ratio` (string, conditional) — `16:9` (default), `9:16`, `1:1`. Required for t2v and reference-to-video.
- `duration` (integer, optional) — video length [3, 15] seconds (default: 5). When using reference video, [3, 10].
- `audio` (boolean, optional) — generate audio (default: false). Affects pricing.
- `watermark` (boolean, optional) — add "可灵 AI" watermark (default: false)

### Media input limits

**Images** (first_frame, last_frame, refer):
- Formats: JPEG, JPG, PNG (no transparency)
- Resolution: [300, 8000] pixels per side
- Max size: 10MB

**Videos** (base, feature):
- Formats: mp4, mov
- Duration: 3-10s
- Resolution: [720, 2160] pixels per side
- Frame rate: 24-60 fps
- Max size: 200MB

### Media combination rules

**kling/
Leer la fuente completa en GitHub (abre una página externa)
Contexto

Trabajo relacionado