Past battle · 2024-11-03 UTC
Speech Recognition Showdown — November 3, 2024
From the Speech Recognition category. 27 marks placed across 6 fighters. WithAudio took the crown.
Final standings
The line-up
The fighters
Profiles of every tool that competed in this battle, ranked by their final score.

WithAudio
One-time purchase text-to-speech reader for Mac and Windows with natural AI voices.

WithAudio is a desktop text-to-speech application for Mac and Windows that converts written content into spoken audio. Users can paste text, load documents, or import articles and have them read aloud using a selection of AI-generated voices, making it useful for proofreading, accessibility, and hands-free reading. Unlike most TTS tools that rely on monthly subscriptions, WithAudio is offered as a one-time purchase, which appeals to readers and writers who want predictable costs. The app focuses on a straightforward listening experience rather than complex audio production, with playback controls and the ability to export generated audio for later use.
Criteria breakdown
- Text-to-speech conversion with AI voices
- Cross-platform support for macOS and Windows
- Audio export for offline listening
- Document and text input options
- Adjustable playback controls
- One-time license activation

CallFluent
AI call analytics platform that transcribes, analyzes, and automates business phone conversations.

CallFluent is an AI-driven call analytics platform designed to help businesses get more value from their phone interactions. It combines speech-to-text transcription, conversation intelligence, and workflow automation to surface insights from sales calls, support tickets, and customer inquiries. The platform analyzes tone, intent, and key topics across calls, giving teams visibility into customer sentiment, agent performance, and common objections. Managers can review trends across high volumes of conversations without manually listening to each recording. With automation features like call summaries, CRM logging, and follow-up triggers, CallFluent aims to reduce repetitive post-call work and help teams act faster on what customers actually say.
Criteria breakdown
- AI speech-to-text transcription
- Sentiment and intent analysis
- Automated call summaries
- Conversation insights dashboard
- CRM and workflow integrations
- Post-call automation triggers

Inworld AI
Build interactive AI-driven characters for games and immersive virtual experiences.

Inworld AI is a character engine designed for developers building games, virtual worlds, and interactive media. It lets creators design NPCs and digital personas with distinct personalities, memories, voices, and motivations, then bring them to life through real-time conversation and contextual behavior. The platform combines language models with goals, knowledge bases, and emotional states so characters can respond dynamically to players rather than following scripted lines. Integrations with engines like Unreal and Unity, plus SDKs and APIs, make it possible to deploy characters across game projects, chat experiences, training simulations, and entertainment apps.
Criteria breakdown
- Customizable AI character personalities
- Long-term memory and knowledge bases
- Voice synthesis and emotional expression
- Unity and Unreal Engine SDKs
- Goals, triggers, and behavior controls
- Multiplayer and multi-character interactions

Molly Personal Assistant
AI assistant for automating workflows and streamlining team collaboration.

Molly is an AI personal assistant designed to automate workflows and streamline team collaboration. It aims to provide a deeply personalized experience through its 'infinite memory' feature, which allows it to remember user preferences and tailor interactions accordingly. Users can delegate tasks to Molly, which it executes in a manner that aligns with the user's style. The assistant also offers proactive suggestions to keep users ahead by anticipating their future needs. Molly is currently in a limited private beta, and interested users can join the waitlist to experience its capabilities.
Criteria breakdown
- Workflow automation
- Task and project tracking
- Team collaboration tools
- Productivity-focused AI assistant
- Smart scheduling and reminders
- Cross-tool integrations

Deepgram
Speech-to-text and text-to-speech APIs for building real-time voice applications.

Deepgram is a voice AI platform that provides developers with APIs for transcribing audio and generating natural-sounding speech. Its models are designed for low-latency, high-accuracy performance across a wide range of languages, accents, and audio conditions, making it suitable for live captioning, call analytics, voice assistants, and conversational agents. Beyond core transcription, Deepgram offers features like speaker diarization, sentiment and topic detection, custom model training, and streaming support. The platform targets engineering teams that need to embed voice capabilities into products without building speech infrastructure from scratch.
Criteria breakdown
- Real-time streaming speech-to-text
- Neural text-to-speech voices
- Speaker diarization and word-level timestamps
- Custom model fine-tuning
- Audio intelligence (sentiment, topics, summarization)
- REST and WebSocket APIs with multi-language SDKs
Kokoro TTS
Open-source multilingual text-to-speech that turns written text into natural-sounding voices.

Kokoro TTS is a text-to-speech system designed to convert written input into clear, natural-sounding speech across a range of languages and voice styles. It aims to make high-quality voice synthesis accessible to developers, content creators, and hobbyists who need realistic audio output for projects like videos, audiobooks, accessibility tools, and voice assistants. The model focuses on producing fluent prosody and recognizable speaker characteristics while remaining lightweight enough to run in a variety of environments. Users can generate spoken audio from text snippets, choose between different voices, and integrate the output into their own workflows or applications.
Criteria breakdown
- Multilingual text-to-speech generation
- Multiple selectable voice profiles
- Natural intonation and pacing
- Exportable audio output
- Suitable for apps, videos, and narration
- Developer-friendly integration




