F5-TTS
Verified BenchmarkZero-shot voice cloning TTS via flow-matching diffusion transformer
Quality & reliability of primary audio output.
Flow-matching DiT
Onboarding & editor ergonomics.
Free open source (MIT code · CC-BY-NC model checkpoints)
Scoring Vectors
Deterministic evaluation across the AiRecMark five-dimension rubric (v2-5dim)
# F5-TTS — official documentation
"https://github.com"
# Pricing source (T0): https://github.com/SWivid/F5-TTS
# Entry tier: Free open source (MIT code · CC-BY-NC model checkpoints) (checked 2026-09-17)
# Rubric: v2-5dim · recorded features: 4
Quantitative Technical Specification
Core architectural subsystems in production release v2.5
Flow-matching DiT
Diffusion transformer with ConvNeXt text handling for fluent speech.
Zero-shot cloning
Clone voices from a short reference clip.
Multilingual
English and Chinese generation with romanization prep.
Apps included
Gradio inference app, API server and speech editing tools.
Direct Peer Matrix: Voice Synthesis & Conversational Engines
Standardized comparative evaluation using normalized 1,000-character test payloads
| Model & Provider | AiRecMark Score | Pricing Base | Performance (rubric) | Languages | Prosody Fidelity |
|---|---|---|---|---|---|
| F5-TTS Leader | 81.5 | Free open source (MIT code | 82/100 | — | 84% |
| ElevenLabs | 95.4 | Usage API / Tiered | — | — | — |
| PlayHT | 89.2 | Freemium | — | — | — |
| Cartesia | 85.8 | $5/mo | — | — | — |
| Adobe Podcast | 85.6 | $0/mo | — | — | — |
Commercial Tiers & Compute Allocation
Transparent character pools, concurrent socket capacities, and API rate limits
Free open source (MIT code · CC-BY-NC model checkpoints)
- check Flow-matching DiT
- check Zero-shot cloning
- check Multilingual
Dedicated infrastructure, security review, and compliance support.
- check SSO / SAML & audit controls
- check Dedicated support channel
- check Custom quota & SLA
F5-TTS stands out for diffusion transformer with flow matching for natural speech. Zero-shot voice cloning TTS via flow-matching diffusion transformer anchors its proposition, and AiRecMark's five-dimension audit lands it at 81.5/100 — a pragmatic default for cloning Voices from Short References.