Skip to content

EthenEthenEthen

videolora training

Cogvideox 5B

Open-weight text-to-video model generating 10-second 720x480 clips with LoRA fine-tuning support.

Studio is in early access; availability resolves per project after sign-in.

Endpoints
3
Eligible
2
Input schemas imported
3 / 3
Delivery
fal.ai
Weights
Not stated
Developer
Not stated in source

Overview

CogVideoX-5B generates 10-second videos from text prompts with temporal coherence. Full model weights and LoRA fine-tuning support give developers customizable open-source flexibility. Default output runs 720x480, configurable through a video-size parameter. Endpoints cover text-driven, image-driven, and video-to-video generation.

Capabilities

  • 10-second text-to-video output
  • Image-to-video and video-to-video endpoints
  • Full weights plus LoRA fine-tuning
  • Configurable 720x480 default resolution

Best for

  • Marketing content creation
  • Product demonstrations
  • Social media video assets

Use cases

  • Customizable open video pipelines
  • Fine-tuned brand motion
  • Short-form social clips

Endpoints

1 under review · 2 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.

EndpointTaskCatalog statusInputs
fal-ai/cogvideox-5bunknownUnder review9 fields · requires prompt
fal-ai/cogvideox-5b/image-to-videoimage to videoEligible10 fields · requires prompt, image_url
fal-ai/cogvideox-5b/video-to-videovideo to videoEligible11 fields · requires prompt, video_url

Questions

How long are clips?

10 seconds continuous with temporal coherence.

Can developers customize it?

Yes: full weights and LoRA fine-tuning support.