Skip to content

EthenEthenEthen

videospeech

Kling Video V2.6

Kling 2.6 Pro image- and text-to-video with integrated speech synthesis and 5 and 10-second control.

Studio is in early access; availability resolves per project after sign-in.

Endpoints
4
Eligible
4
Input schemas imported
4 / 4
Delivery
fal.ai
Weights
Not stated
Developer
Not stated in source

Overview

Kling 2.6 Pro integrates speech synthesis into generation, pairing cinematic motion with native Chinese and English voice output. Dialogue embeds directly in prompts for automatic voicing matched to lips and scene timing. Duration runs 5 or 10 seconds; single images animate with scene continuity. Motion-control endpoints join the pro and standard tiers.

Capabilities

  • Integrated Chinese and English speech synthesis
  • Prompt-embedded dialogue voicing
  • 5 and 10-second duration control
  • Lip-matched AV coherence

Best for

  • Social media creation
  • Marketing production
  • Cinematic prototyping

Use cases

  • Speaking-character clips
  • No-stitch AV content
  • Coherent short narratives

Endpoints

4 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.

EndpointTaskCatalog statusInputs
fal-ai/kling-video/v2.6/pro/image-to-videoimage to videoEligible7 fields · requires prompt, start_image_url
fal-ai/kling-video/v2.6/pro/motion-controlvideo editingEligible5 fields · requires image_url, video_url, character_orientation
fal-ai/kling-video/v2.6/pro/text-to-videotext to videoEligible6 fields · requires prompt
fal-ai/kling-video/v2.6/standard/motion-controlvideo editingEligible5 fields · requires image_url, video_url, character_orientation

Questions

Separate audio workflow?

No: speech synthesizes inside the pipeline.

Which voices?

Chinese and English output.