Kling 2.6 is Kuaishou's short-form text-to-video and image-to-video model with native audio generation — toggle the "sound" parameter to generate synchronized speech, effects, and ambience in a single pass. It generates 5s or 10s clips across 16:9, 9:16, and 1:1 aspect ratios, with prompts up to 1,000 characters. It's also the base model for Kling Motion Control, which transfers motion from a reference video onto a static image at 720p or 1080p. Best for social-media clips and dialogue-driven scenes with strong action physics.
Kling 2.6 is Kuaishou's short-form text-to-video and image-to-video model with native audio generation — toggle the "sound" parameter to generate synchronized speech, effects, and ambience in a single pass. It generates 5s or 10s clips across 16:9, 9:16, and 1:1 aspect ratios, with prompts up to 1,000 characters. It's also the base model for Kling Motion Control, which transfers motion from a reference video onto a static image at 720p or 1080p. Best for social-media clips and dialogue-driven scenes with strong action physics.
Credit costs are examples — final credits are calculated during generation based on your selected resolution and duration.
See the full model catalog, pricing plans, or take the model quiz.
We use analytics to improve your experience. See our Privacy Policy.