Skip to content

EthenEthenEthen

Open Source Model Profile · CohereLabs

cohere-transcribe-03-2026

cohere-transcribe-03-2026 is a 2.07B-parameter Conformer-based speech-recognition model from CohereLabs. According to the model card, it transcribes audio to text in 14 languages.

Publisher
CohereLabs
Task
automatic-speech-recognition
Model type
cohere_asr
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

cohere-transcribe-03-2026 is published by CohereLabs as a cohere_asr automatic-speech-recognition model. The captured configuration identifies CohereAsrForConditionalGeneration, and Safetensors metadata reports 2065804048 parameters. According to the model card, it is a 2B-parameter audio-in, text-out model trained from scratch on 14 languages.

Recorded capabilities

Conformer encoder-decoder ASR

Captured config identifies CohereAsrForConditionalGeneration, and the card describes a conformer-based encoder-decoder with a large Conformer encoder and lightweight Transformer decoder.

14-language transcription

According to the model card, the model supports 14 languages across European, APAC, and MENA groupings, with language-code selection for non-English transcription.

Transformers and batched inference

The model card describes native Transformers support for offline inference and batched processing of multiple audio files with chunking and reassembly.

Use cases in the source record

  • Audio transcription into text in any of the 14 documented languages, with explicit language-code selection for non-English audio.
  • Offline batched transcription experiments using the card's documented Transformers workflow, including mixed short-form and long-form audio.

Limitations and unknowns

  • No independent benchmark measurements were extracted; publisher accuracy and speed comparisons remain publisher claims.
  • The card describes strong human-preference results and best-in-class accuracy claims, but these are publisher statements without independently extracted evaluation data.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: CohereLabs/cohere-transcribe-03-2026

Captured: Unknown. Processed: 2026-09-07T19:34:29.796709+00:00.

Cohere Transcribe Cohere Transcribe is an open source release of a 2B parameter dedicated audio-in, text-out automatic speech recognition (ASR) model. The model supports 14 languages. Developed by: Cohere and Cohere Labs . Point of Contact: Cohere Labs . Name cohere-transcribe-03-2026 Architecture conformer-based encoder-decoder Input audio waveform → log-Mel spectrogram. Audio is automatically resampled to 16kHz if necessary during preprocessing. Similarly, multi-channel (stereo) inputs are averaged to produce a single channel signal. Output transcribed text Model size 2B Model a large Conformer encoder extracts acoustic representa…

F001F002F003F004F005F006F007F010F011F013F014F015F016F017F018F019F020