Skip to content

EthenEthenEthen

Open Source Model Profile · hexgrad

Kokoro-82M

Kokoro-82M is an open-weight text-to-speech model from hexgrad. According to the model card, it has 82 million parameters with Apache-licensed weights for production or personal deployment.

Publisher
hexgrad
Task
text-to-speech
Model type
Unknown
License
apache-2.0
Library
Unknown
Publication status
Accepted · not indexed

Model overview

Kokoro-82M is published by hexgrad as a text-to-speech checkpoint. Hub tags record English association and references to arXiv 2306.07691 and 2203.02395. According to the model card, it is an 82M-parameter open-weight TTS model built on StyleTTS 2 with an ISTFTNet vocoder path, using a decoder-only release without diffusion or encoder release.

Recorded capabilities

82M-parameter TTS weights

According to the model card, Kokoro has 82 million parameters and the publisher describes it as lightweight, faster, and more cost-efficient than larger models.

StyleTTS 2 plus ISTFTNet path

According to the model card, the model facts list StyleTTS 2 and ISTFTNet architectures with a misaki grapheme-to-phoneme library and 24kHz sample output.

Apache-2.0 deployment

The captured record lists Apache-2.0, and according to the model card, the Apache-licensed model has been deployed in projects and commercial APIs.

Permissive-data training claim

According to the model card, Kokoro was trained exclusively on permissive or non-copyrighted audio and IPA phoneme labels, with a few hundred hours of audio in total.

Use cases in the source record

  • Text-to-speech synthesis from English text input, matching the text-to-speech pipeline and the publisher's production-to-personal deployment description.
  • Multilingual speech experiments and sample listening referenced by the model card, which points to samples and an advanced-usage section for more languages and details.

Limitations and unknowns

  • According to the model card, API market-rate figures under $1 per million characters are dated April 2025 with named sources and should not be treated as current pricing.
  • According to the model card, similarly named third-party websites may be scams; users should rely on the publisher repository and demo links rather than those sites.
  • No independent evaluation results were extracted; quality comparisons to larger models are publisher descriptions only.
  • No Safetensors parameter count, quantization, or context-window value was extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: hexgrad/Kokoro-82M

Captured: Unknown. Processed: 2026-09-07T19:34:46.579121+00:00.

Kokoro is an open-weight TTS model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects. 🐈 GitHub : https://github.com/hexgrad/kokoro 🚀 Demo : https://hf.co/spaces/hexgrad/Kokoro-TTS As of April 2025, the market rate of Kokoro served over API is under $1 per million characters of text input , or under $0.06 per hour of audio output. (On average, 1000 characters of input is about 1 minute of output.) Sour…

F001F002F003F004F005F006F007F008F009F010F012F013F014F015F016