Performance varies meaningfully across the three datasets with different audio lengths, accents, vocabulary, and …

Performance varies meaningfully across the three datasets with different audio lengths, accents, vocabulary, and background noise. On AA-AgentTalk, our private test set, ElevenLabs Scribe v2 Realtime leads both final (2.8%) and partial (2.9
Ranked #60 on backlist 2026-05-28 (28 May 2026 UTC) · by (Artificial Analysis) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.