TESIGN
SAVED

EN · KO

Auto-generated — not yet edited · only the numbers are verified

  1. TESIGN / RADAR
  2. REPOSITORY CARD
  3. openbmb/voxcpm

openbmb/voxcpm

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

★ 37.1k▲ +54 / 7dDaily star increases over 14 days. Hollow bars mean no data. The last day may be partial.2026-09-03 no data2026-09-04 +112026-09-05 +92026-09-06 +92026-09-07 +42026-09-08 no data2026-09-09 no data2026-09-10 no data2026-09-11 +32026-09-12 +22026-09-13 +82026-09-14 +72026-09-15 +312026-09-16 +3largest one-day increase in the 14 days+31last 14 days · stars per dayhollow = no data

RANKS All-time #187

AT A GLANCE

LANGUAGE
Python
LICENSE
Apache-2.0
USAGE
Commercial use, redistribution OK. Keep notices; mark changes.
ACTIVITY
last commit 14 days ago ()
TOPICS
  • audio
  • deeplearning
  • minicpm
  • multilingual
  • python
  • pytorch
  • speech
  • speech-synthesis
  • text-to-speech
  • tts
  • tts-model
  • voice-cloning
HOMEPAGE
https://voxcpm.com
REPOSITORY
GitHub ↗
CATEGORY
AI

A summary, not legal advice.

README EXCERPT

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning English 中文 👋 Join our community for discussion and support! Feishu     Discord     📚 MiniCPM Wiki VoxCPM is a tokenizer-free Text-to-Speech system that directly generates continuous speech representations via an end-to-end diffusion autoregressive architecture , bypassing discrete tokenization to achieve highly natural and expressive synthesis. VoxCPM2 is the latest major release — a 2B parameter model trained on over 2 million hours of multilingual speech data, now supporting 30 languages , Voice Design , Controllable Voice Cloning , and 48kHz studio-quality audio output. Built on a MiniCPM-4 backbone. ✨ Highlights - 🌍 30-Language Multilingual — Input text in any of the 30 supported languages and synthesize directly, no language tag needed - 🎨 Voice Design — Create a brand-new voice from a natural-language description alone (gender, age, tone, emotion, pace …), no reference audio required - 🎛️ Controllable Cloning — Clone any voice from a short reference clip, with optional style guidance to steer emotion, pace, and expression while preserving the ori…

The opening of the GitHub README as stored, at most 1,200 characters. Markdown is not rendered.

TIMELINE

  1. Repository created
  2. Last push
  3. FIRST SEEN BY TESIGN ◌ BACK CATALOG

AI chooses the lists under the owner’s delegation. No human review is running in September 2026. Total stars, increases, cross-source signals, last updates and licences are shown as evidence. How ranks work →

If this repository gets editorial text (why, build, who, start, caveat) it becomes an edited entry. Until then the page shows only stored GitHub metadata and numbers.