Settings

Theme

Sopro: A 169M parameter real-time TTS model with zero-shot voice cloning

github.com

6 points by marques576 7 months ago · 1 comment

Reader

marques576OP 7 months ago

Some features:

169M parameters

Streaming support

Zero-shot voice cloning

0.25 RTF on CPU, meaning it generates 30 seconds of audio in 7.5 seconds

Requires 3-12 seconds of reference audio for voice cloning

Apache 2.0 license

The model was trained on a single L40S GPU. It’s not SOTA in most cases, can be a bit unstable, and sometimes fails to capture voice likeness.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection