INDEX / AUDIO_SPEECH / F5_TTS
github.com verified
F5

F5-TTS

Verified Benchmark

Zero-shot voice cloning TTS via flow-matching diffusion transformer

FREE PRICING SCORE 81.5/100 4 RECORDED FEATURES UPDATED 2026-09-17
AiRecMark Score
81.5 /100
workspace_premium #20 in Voice & Music AI
Launch Console north_east
Output Quality record_voice_over
84%

Quality & reliability of primary audio output.

Feature Depth bolt

Flow-matching DiT

Usability fingerprint
70%

Onboarding & editor ergonomics.

Entry Price payments
Free open source (MIT code

Free open source (MIT code · CC-BY-NC model checkpoints)

Scoring Vectors

Deterministic evaluation across the AiRecMark five-dimension rubric (v2-5dim)

SNAPSHOT 2026-09-17
sentiment_satisfied Output Quality & Reliability (Weighted 25%) 84%
format_quote Feature Depth & Integrations (Weighted 20%) 74%
noise_aware Onboarding & Usability (Weighted 15%) 70%
auto_stories Runtime Performance & Latency (Weighted 20%) 82%
graphic_eq Price-to-Value Efficiency (Weighted 20%) 94%
PACKET TRANSIT DISTRIBUTION (PING JITTER) 5-dim spread ±12.0 pts (5-dim spread)
40ms 65ms (Median) 75ms (Flash TTFT) 120ms
graphic_eq ARCHIVE RECORD
v2-5dim
QUALITY DIM 84%
FEATURES DIM 74%
VALUE DIM 94%
terminal quickstart.py
# F5-TTS — official documentation
"https://github.com"

# Pricing source (T0): https://github.com/SWivid/F5-TTS
# Entry tier: Free open source (MIT code · CC-BY-NC model checkpoints) (checked 2026-09-17)
# Rubric: v2-5dim · recorded features: 4

Quantitative Technical Specification

Core architectural subsystems in production release v2.5

speed

Flow-matching DiT

Diffusion transformer with ConvNeXt text handling for fluent speech.

Capability T0 Verified
mic_double

Zero-shot cloning

Clone voices from a short reference clip.

Capability T0 Verified
support_agent

Multilingual

English and Chinese generation with romanization prep.

Capability T0 Verified
surround_sound

Apps included

Gradio inference app, API server and speech editing tools.

Capability T0 Verified

Direct Peer Matrix: Voice Synthesis & Conversational Engines

Standardized comparative evaluation using normalized 1,000-character test payloads

Benchmarks Updated 2026-09-14
Model & Provider AiRecMark Score Pricing Base Performance (rubric) Languages Prosody Fidelity
F5-TTS Leader 81.5 Free open source (MIT code 84%
ElevenLabs 95.4 Usage API / Tiered
PlayHT 89.2 Freemium
Cartesia 85.8 $5/mo
Adobe Podcast 85.6 $0/mo

Commercial Tiers & Compute Allocation

Transparent character pools, concurrent socket capacities, and API rate limits

Pro Popular
Custom /mo

Free open source (MIT code · CC-BY-NC model checkpoints)

  • check Flow-matching DiT
  • check Zero-shot cloning
  • check Multilingual
Enterprise
Custom SLA terms

Dedicated infrastructure, security review, and compliance support.

  • check SSO / SAML & audit controls
  • check Dedicated support channel
  • check Custom quota & SLA
verified_user
AIrecmark Analyst Consensus Verdict SCORING MODEL: v2-5dim

F5-TTS stands out for diffusion transformer with flow matching for natural speech. Zero-shot voice cloning TTS via flow-matching diffusion transformer anchors its proposition, and AiRecMark's five-dimension audit lands it at 81.5/100 — a pragmatic default for cloning Voices from Short References.

Lead Evaluator: AiRecMark Audio Analyst Group
ARCHIVE SNAPSHOT: 2026-09-17 · github.com