A high-quality text-to-video and image-to-video model focused on realistic motion, scene consistency, and strong camera control.
Kling 2.1 is a high-quality text-to-video and image-to-video model focused on realistic motion, scene consistency, and strong camera control. It produces visually coherent videos with natural physics and movement, and works well for action sequences, nature and landscape content, and scenes that need smooth motion. It supports cinematic camera language (pan, dolly, tracking, handheld) and negative prompts, and works best with clear, natural-language prompts that describe subject, movement, scene, and lighting. Kling 2.1 does not generate audio; add music or voiceover in post-production.
Action sequences and nature content
Credit costs are examples — final credits are calculated during generation based on your selected resolution and duration.
See the full model catalog, pricing plans, or take the model quiz.
We use analytics to improve your experience. See our Privacy Policy.