AiRecMark/Intelligence/Output Quality leaders
C-04 · Output Quality leadersarchive-recorded · snapshot 2026-09-17

Output Quality Leaders: Top 25

The 25 highest output quality scores (score.dims.quality) across all 448 archives, each shown against its overall score so the deviation is visible. Aggregation is a plain sort of recorded values — no re-weighting; method and N are published below.

Key numbers

Archives ranked
448
data/tools · N=448
Output Quality median
80
score.dims.quality · N=448
Top score
96
score.dims.quality · N=448
Leader
Midjourney
score.dims.quality · N=448
Correlation with overall
r=0.82
pearson(dim, overall) · N=448

Output Quality — top 25 of 448

sort: dims desc, then overall desc, then slug; deviation = dim − overall

#ToolCategoryOutput QualityOverallDeviation
1MidjourneyVideo9685.7+10.3
2ElevenLabsAudio & Voice9587.8+7.2
3ClaudeResearch9488.9+5.1
4Claude CodeCoding9486.1+7.9
5ChatGPTResearch9390.9+2.1
6CursorCoding9389.2+3.8
7Google VeoVideo9283.2+8.8
8PlayHTAudio & Voice9189.2+1.8
9PerplexityResearch9188.4+2.6
10OpenAI CodexCoding9084.6+5.4
11Nano BananaDesign9083.4+6.6
12AlphaSenseResearch9080.3+9.7
13RespeecherAudio & Voice9079.7+10.3
14Magnific AIDesign9077.4+12.6
15SoraVideo9076.5+13.5
16GeminiResearch8988.5+0.5
17FLUXDesign8988.3+0.7
18GitHub CopilotCoding8888.1-0.1
19ZedCoding8886.1+1.9
20AiderCoding8886+2
21CartesiaAudio & Voice8885.8+2.2
22DeepL WriteWriting8885.8+2.2
23Adobe PodcastAudio & Voice8885.6+2.4
24NotebookLMResearch8885.2+2.8
25DeepgramAudio & Voice8885.1+2.9

Output Quality — top 10

Midjourney96ElevenLabs95Claude94Claude Code94ChatGPT93Cursor93Google Veo92PlayHT91Perplexity91OpenAI Codex90

Recorded score.dims.quality of the ten leading archives (scale 0–100).

Method

  • Population: all 448 archives in data/tools with a recorded score.dims.quality.
  • Ordering: dimension score descending; ties broken by overall descending, then slug ascending — a deterministic total order.
  • Deviation column = score.dims.quality − score.overall per archive; positive values lead their own overall.
  • Pearson r between the dimension and overall is computed over all 448 pairs (standard formula, rounded to 2 decimals).
  • Scores are quoted as recorded in each archive; nothing is recomputed or re-weighted.
Data appendix. Source: data/tools/*.json · snapshot 2026-09-17 · method: sort(score.dims.quality desc, overall desc, slug) top 25 + deviation + pearson · aggregates published with method and N. T1 benchmarks recorded: 3 / 448.