Past battle · 2025-05-20 UTC
Audio Generation Showdown — May 20, 2025
From the Audio Generation category. 25 marks placed across 8 fighters. AI Music Generator - Create Songs from Text with AI took the crown.
Final standings
The line-up
The fighters
Profiles of every tool that competed in this battle, ranked by their final score.

AI Music Generator - Create Songs from Text with AI
Generate original songs with AI vocals from text prompts.

AI Music Generator is a web‑based service that creates complete songs—including vocals and optional custom lyrics—directly from user‑provided text descriptions. By turning simple prompts about genre, mood, instruments, tempo, and style into audio files, it removes the need for musical training or production software. The platform targets a broad audience: content creators who need royalty‑free background tracks, game developers looking for adaptive soundtracks, marketers seeking jingles or brand themes, and musicians who want quick demos or genre exploration. Users type a description, optionally supply lyrics or let the AI write them, select a male or female vocal style, and let the system generate a track. Generated music covers a wide range of genres such as pop, rock, hip‑hop, electronic, jazz, classical, country, and more, with the ability to blend styles. The AI‑driven vocals are described as realistic and expressive, and every track is released under a commercial‑use license, allowing unrestricted use in videos, podcasts, games, and advertisements. The creation process follows four steps: describe the desired music, add or generate lyrics, generate and customize the track, then download the final MP3. According to the site, a complete song is produced in one to three minutes, and users can regenerate or tweak parameters until satisfied. While the service delivers fast, professional‑quality output without any musical expertise, it is limited to the preset vocal timbres (male and female) and the genre/style options offered by the platform. Users seeking highly nuanced arrangements or unique vocal characteristics may still need traditional composition tools.
Criteria breakdown
- Text‑to‑music generation with optional lyric creation
- AI‑generated male or female vocal tracks
- Genre and mood selection across many styles
- MP3 download of completed songs
- Commercial‑use licensing for all outputs


Stories is an AI-powered audio walking tour application that allows users to explore the history and stories of various locations around the world. It is a perfect companion for history buffs and travelers, requiring no reading or searching. The app offers a diverse range of stories, from the Indigenous Palawa People to the Magna Carta's genesis and global influence, and even the history of the Mill Valley lumber industry. Users can find and listen to these stories while traveling, with the app working anywhere, including cities, villages, towns, forest hikes, and more. It is available for iOS and Android devices.
Criteria breakdown
- AI-generated audio walking tours
- Global location coverage
- On-demand narration as you walk
- Mobile app with location awareness
- Content about history, architecture and landmarks

AI Song Generator
Turn text prompts and ideas into original AI-generated songs in minutes.

AI Song Generator is a music creation tool that converts written prompts, themes, or lyrics into full audio tracks. Users describe the mood, genre, or story they want, and the system produces a song complete with instrumentation and vocals. It is aimed at hobbyists, content creators, and songwriters who want a fast way to prototype musical ideas without needing instruments or production skills. Generated tracks can be used as inspiration, background music for videos, or starting points for further editing.
Criteria breakdown
- Text-to-song generation
- Genre and mood selection
- Lyric input or AI-written lyrics
- Vocal and instrumental tracks
- Downloadable audio files
- Multiple song length options

Cartesia Sonic-3
Real-time, multilingual text-to-speech with sub‑90 ms latency and voice cloning

Cartesia Sonic-3 is a text‑to‑speech model that focuses on delivering natural, expressive speech in real time. It targets enterprises and developers who need high‑quality audio for customer‑facing applications such as marketing calls, sales outreach, and automated support. The model is built on state‑space architecture, which the vendor claims provides sub‑90 ms latency while maintaining a ranking of #1 for naturalness. The service supports more than 40 languages and a variety of regional accents, allowing a single voice model to be used across global markets. Voice cloning is offered with as little as ten seconds of source audio, producing a synthetic voice that retains the speaker’s identity. Users can also upload custom pronunciation dictionaries to ensure proper rendering of domain‑specific terms, proper nouns, or brand names. Sonic‑3’s expressive capabilities include the ability to convey emotion, tone, and even laughter, aiming to preserve the nuance of the original speaker during localization. The platform is positioned as an enterprise‑grade solution, with compliance certifications such as HIPAA, SOC 2 Type 2, GDPR, and PCI, and options for both cloud and on‑premise deployment. Typical workflows involve integrating the API into marketing automation, CRM, or contact‑center platforms to generate personalized audio messages at scale. The vendor highlights use cases like warm‑lead outreach, real‑time sales objection handling, automated customer authentication, and lifecycle‑stage follow‑ups. Security and compliance are emphasized for regulated industries. Limitations noted in the public material include the need to contact sales for pricing and onboarding, and the reliance on a cloud or managed deployment model for most customers. While the language coverage is broad, it is limited to the 40+ languages explicitly supported, and the voice cloning feature may not capture the full expressive range of a speaker with only ten seconds of audio.
Criteria breakdown
- Sub‑90 ms low‑latency inference
- Multilingual TTS across 40+ languages
- Instant voice cloning with 10 s audio sample
- Custom pronunciation dictionary
- Enterprise‑grade security and compliance

Lyrics To Song AI
Turn written lyrics into finished, studio-style songs in seconds using AI.

Lyrics To Song AI is a generative music tool that converts plain text lyrics into fully produced tracks, complete with vocals, instrumentation, and mixing. Users paste or write their lyrics, pick a style or mood, and receive a ready-to-share song without needing musicians, recording gear, or production skills. The platform is designed for songwriters, content creators, hobbyists, and marketers who want quick musical output for demos, social posts, videos, or personal projects. Because the entire pipeline—from melody to vocal delivery—is automated, turnaround is fast and the barrier to producing original music is low.
Criteria breakdown
- Lyrics-to-song generation pipeline
- AI-generated vocal performances
- Automatic instrumentation and arrangement
- Multiple genre and mood presets
- Fast export of ready-to-release tracks
- Browser-based workflow with no setup

PodMind AI Podcast Generator
Turn PDFs and text into natural-sounding AI podcasts in minutes, with multi-language support.

PodMind AI Podcast Generator converts written content such as PDFs, articles, and raw text into spoken-word audio that mimics the pacing and tone of a real podcast. Users upload or paste source material and the tool produces a finished episode without the need for scripting, recording, or editing. The generator supports multiple languages, making it useful for educators, marketers, researchers, and content creators who want to repurpose documents into audio for global audiences. Output is intended to sound conversational rather than robotic, so listeners can absorb long-form material on the go.
Criteria breakdown
- PDF and text-to-podcast conversion
- Natural-sounding AI voice generation
- Multi-language output
- Fast turnaround in minutes
- Works with long-form documents
- Content repurposing for creators and educators

TuneX AI Music, Song Generator
AI-powered app for generating royalty-free songs, beats, and vocals across any genre.

TuneX AI Music is a song generation tool that lets users create original tracks without musical training or instruments. By describing a mood, genre, or theme, users can produce full songs complete with beats and vocals in a matter of moments. The platform targets creators who need quick background music, demo tracks, or inspiration for their own projects. Outputs are royalty-free, making them suitable for videos, podcasts, social media, and other personal or commercial uses. With genre flexibility and a low barrier to entry, TuneX AI aims to make music creation accessible to anyone with an idea, regardless of their production experience.
Criteria breakdown
- Text-to-song AI generation
- Custom beat creation
- AI vocal synthesis
- Multi-genre support
- Royalty-free licensing
- Beginner-friendly interface
Natural TTS Labs
Free AI text-to-speech tool for generating natural-sounding voiceovers

Natural TTS Labs is a text-to-speech platform that converts written content into lifelike spoken audio using AI-generated voices. It targets creators, educators, and developers who need quick voiceovers without recording equipment or professional narrators. The tool emphasizes accessibility, offering a straightforward interface where users paste text, pick a voice, and generate audio output. It aims to lower the barrier to producing audio content for videos, presentations, podcasts, and accessibility use cases. With free access and a focus on natural intonation, it positions itself as an entry-level option for anyone exploring AI voice synthesis before committing to paid enterprise solutions.
Criteria breakdown
- Text-to-speech conversion with AI voices
- Multiple voice options
- Browser-based, no installation needed
- Audio file download
- Natural-sounding intonation and pacing
- Suitable for videos, e-learning, and accessibility






