AuK: An Open-Source Foundational Model for Speech Generation and Editing

Supported tasks

  • Speech Generation: Zero-shot TTS · Instruct TTS
  • Content Editing: Lyric Editing · Speech Content Editing
  • Acoustic Editing: Pitch Editing · Speed Editing · Volume Editing
  • Paralinguistic Editing: Emotion · Timbre · De-accent · Nonverbal Editing · Whisper Conversion
  • Enhancement & Separation: Speech Enhancement · Speech Separation · Music Separation

Duration priority: Duration > reference-text + target-text estimate > match the source length.

PE determines the task and instruction; Duration is estimated only when set to 0.
Model
0 30

0 = PE auto-estimation (PE required); values above 0 override PE duration.

4 64
0 5

Base: NFE 32 / CFG 2.0 · Flash: NFE 4 / CFG 0 (locked)

Prompt Enhancer output

Examples

Select an example to load it into the panel, then click Generate. You can also try different seeds.