Skip to content

EthenEthenEthen

Open Source Model Profile · chetwinlow1

Ovi

Ovi is a chetwinlow1 image-to-video model for synchronized video-plus-audio generation. Captured metadata reports about 11.66B parameters.

Publisher
chetwinlow1
Task
image-to-video
Model type
Unknown
License
apache-2.0
Library
ovi
Publication status
Accepted · not indexed

Model overview

Ovi is published by chetwinlow1 as an image-to-video model for joint audio-video generation. Captured Safetensors metadata reports 11,660,753,108 parameters. According to the model card, it generates synchronized video and audio from text or text-plus-image inputs, with Ovi 1.1 supporting 10-second output at 960x960 resolution.

Recorded capabilities

Synchronized video-plus-audio generation

According to the model card, the model generates synchronized video and audio simultaneously from text-only or text-plus-image conditioning.

10-second 960x960 video output

The model card describes temporally consistent 10-second or 5-second videos at 24 FPS and 960x960 resolution across several aspect ratios.

Speech and audio prompt tags

The card documents an updated prompt format with <S> and <E> speech tags and Audio: descriptions replacing the older AUDCAP format.

Quantized and multi-GPU execution

According to the model card, fp8 and qint8 weights support 24GB VRAM execution, with CPU offload, sequence parallelism, and single- or multi-GPU inference scripts.

Use cases in the source record

  • Text-to-audio-video generation with speech tags and audio descriptions for synchronized clips.
  • Image-to-audio-video generation from text-plus-image conditioning at 960x960 resolution.
  • ComfyUI and Gradio video-creation workflows using the documented example prompts and inference scripts.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current availability.
  • Output quality, timing, and VRAM figures come from the publisher model card and have not been independently measured by Ethen.

Source and provenance

Source: chetwinlow1/Ovi

Captured: Unknown. Processed: 2026-09-07T19:35:54.808431+00:00.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation Chetwin Low * 1 , Weimin Wang * † 1 , Calder Katyal 2 * Equal contribution, † Project Lead 1 Character AI, 2 Yale University 🎥 Video Demo 🆕 Ovi 1.1 10-Second Demo Ovi 1.1 – 10-second temporally consistent video generation (960 × 960 resolution) 🎬 Original 5-Second Demo 🆕 Ovi 1.1 Update (10 November 2025) Key Feature: Enables temporal-consistent 10-second video generation at 960 × 960 resolution Training Improvements: - Trained natively on 960×960 resolution videos - Dataset includes 100% more videos for greater diversity Prompt Format Update: Audio descriptions sho…

F001F002F003F004F005F006F007F008F009F010F011F014F015F018F024F025F028