Past battle · 2026-07-24 UTC
Observability Showdown — July 24, 2026
From the Observability category. 21 marks placed across 9 fighters. Coval (YC S24) took the crown.
Final standings
The line-up
The fighters
Profiles of every tool that competed in this battle, ranked by their final score.

Coval (YC S24)
Simulation and evaluation platform for testing AI voice and chat agents at scale.

Coval is a developer platform built to simulate, test, and evaluate AI agents before they reach production. It lets teams run thousands of synthetic conversations against their voice or chat agents, measuring how they handle edge cases, interruptions, tool calls, and multi-turn dialogue. Backed by Y Combinator (S24), Coval positions itself as a 'self-driving cars approach' to agent reliability, applying rigorous simulation-based testing to conversational AI. Engineers can define scenarios, replay production traffic, score outputs against custom metrics, and track regressions across agent versions. The platform targets teams shipping customer-facing agents in support, sales, and operations, where reliability and consistency are critical for deployment.
Criteria breakdown
- Large-scale conversation simulation
- Voice agent testing with realistic dialogue
- Custom evaluation metrics and scoring
- Regression tracking across agent versions
- Scenario and edge-case generation
- Production traffic replay

Fiddler AI
AI observability and security platform for monitoring, explaining, and governing ML and LLM applications.

Fiddler AI is an enterprise platform that helps teams monitor, analyze, and secure machine learning models and generative AI applications in production. It provides visibility into model performance, data drift, bias, and quality issues, while also offering safeguards against risks specific to LLMs such as hallucinations, prompt injection, and unsafe outputs. Designed for ML engineers, data scientists, and risk and compliance teams, Fiddler combines explainability, real-time monitoring, and guardrails in a single workflow. It integrates with common ML pipelines and cloud environments, helping organizations operationalize responsible AI practices at scale.
Criteria breakdown
- Model performance and drift monitoring
- LLM hallucination and safety detection
- Prompt injection and jailbreak protection
- Explainable AI and root cause analysis
- Bias and fairness assessments
- Dashboards and alerts for production AI

Future AGI
A platform enhancing AI accuracy through comprehensive evaluation and optimization tools.

Future AGI is a platform that enhances AI accuracy through comprehensive evaluation and optimization tools. It allows users to build self-improving agents, catch what breaks, and know why. The platform offers a range of features, including agent evaluation, optimization, and simulated conversations. Users can create and manage scenarios, and the platform provides analytics and optimization tools to improve agent performance. Future AGI seems to cater to developers and businesses looking to deploy AI-powered chatbots and agents. It allows users to simulate conversations, evaluate and optimize agent performance, and integrate with knowledge bases. The platform appears to be a tool for developers to create and refine AI agents, rather than a customer-facing interface. The platform offers features such as agent evaluation, simulated conversations, and scenario management. It allows users to edit and manage prompts, add retrieval steps for knowledge base articles, and integrate with global prompts for handling complex situations. While Future AGI provides valuable features for AI development and optimization, it also comes with limitations. For example, it relies on general knowledge and may require additional steps for more complex scenarios. Additionally, the platform's reliance on simulated conversations may limit its ability to handle real-world uncertainties. Future AGI seems to offer a unique set of features for AI development and optimization. Its comprehensive evaluation and optimization tools make it a valuable resource for developers looking to refine their AI agents. However, it may require additional development or integration to handle more complex scenarios and real-world uncertainties.
Criteria breakdown
- Agent evaluation and ranking
- Simulated conversations and scenario management
- Knowledge base integration and retrieval
- Global prompt handling for complex situations
- Analytics and optimization tools


On the AI2AI project website, users can observe real-time conversations between two AI agents. To initiate an interaction, users must create two entities, assigning them a prompt, context, voice, and gender, and select a language model. The interaction can be started, saved, and shared with other users on the site. Example interactions are provided, showcasing different scenarios, such as seduction, anger, curiosity, and sales pitches. New users need to register and receive $1 in credit to explore the platform. Pricing is based on a per-minute rate, starting from 1 cent per minute.
Criteria breakdown
- Live AI-to-AI dialogue viewing
- Two distinct AI entities in conversation
- Observational, hands-off user experience
- Showcase of emergent AI behavior
- Accessible browser-based interface

Arize AI
An AI observability and LLM evaluation platform that assists AI developers and data scientists in monitoring, troubleshooting, and enhancing the performance...

Arize AI is an AI observability and LLM evaluation platform. It assists AI developers and data scientists in monitoring, troubleshooting, and enhancing the performance of their AI models. The platform provides tools for identifying issues and proposing fixes, helping users improve their models' accuracy and reliability. Arize AI is designed for AI developers and data scientists who need to optimize their models' performance. The platform's capabilities include monitoring and troubleshooting, allowing users to quickly identify and address issues. Arize AI also provides features for evaluating and enhancing model performance, enabling users to refine their models over time. By using Arize AI, users can ensure their AI models are operating at peak performance, leading to better decision-making and outcomes. The platform's focus on observability and evaluation sets it apart from other AI development tools, making it a valuable resource for those working with large language models and other complex AI systems.
Criteria breakdown
- AI model monitoring
- Troubleshooting and issue identification
- Proposed fix generation
- Model performance evaluation
- LLM evaluation and optimization

Edwin AI
AI agent for IT operations that speeds up incident detection, triage, and resolution.

Edwin AI is an AI agent for IT operations designed to speed up incident detection, triage, and resolution. It provides a centralized platform for IT teams to investigate incidents, understand their impact, find or generate fixes, and apply them across existing tools without needing to switch between systems. Edwin AI correlates alerts, identifies root causes, and initiates remediation automatically, starting from the first alert through to verified resolution. It uses historical patterns and observability data to predict and prevent outages. The tool integrates with over 3,000 tools across observability, APM, security, and CMDB, enabling real-time, actionable insights and eliminating silos. A Forrester study found that Edwin AI delivered a 313% ROI for a composite organization, with a payback period of 6 months or less.
Criteria breakdown
- Alert correlation and noise reduction
- AI-driven root cause suggestions
- Natural-language incident summaries
- Integrations with ITSM and observability platforms
- Automated triage workflows
- Knowledge enrichment from past incidents


FoundryAI is a development platform focused on creating AI agents that handle real business workflows. It combines agent design, testing, and continuous improvement tools so teams can move from prototype to production without stitching together separate systems. The platform emphasizes evaluation, giving builders ways to measure agent performance against defined tasks and refine behavior over time. This makes it suited for organizations automating customer support, internal operations, or repetitive knowledge work where reliability matters. FoundryAI targets technical teams who need more control than no-code builders offer but want faster iteration than building agents entirely from scratch.
Criteria breakdown
- Agent building environment
- Evaluation and testing tools
- Performance monitoring
- Workflow automation support
- Iterative improvement loops
- Integration with business systems
Helicone AI
All-in-one observability platform to monitor, debug, and improve production LLM apps.
Helicone AI is a developer-focused observability platform built specifically for applications powered by large language models. It captures requests, responses, costs, and latency across providers, giving engineering teams a unified view of how their LLM features behave in production. Beyond logging, Helicone offers tools for debugging prompts, tracing multi-step agent workflows, running evaluations, and tracking user-level usage. Teams can identify regressions, control spend, and iterate on prompts with data rather than guesswork. It integrates with popular model providers and frameworks through a lightweight proxy or async logging, making it straightforward to add to existing stacks without major code changes.
Criteria breakdown
- Request and response logging
- Cost and token usage tracking
- Prompt management and versioning
- Agent and session tracing
- Custom evaluations and dashboards
- User and rate-limit analytics

llm scout
Monitor how your brand appears across ChatGPT, Claude, Perplexity, and Google AI Overviews.

LLM Scout is a brand monitoring tool built for the era of generative search. It tracks how your company, products, and competitors are mentioned across major AI assistants and answer engines, giving marketing and SEO teams visibility into a channel that traditional analytics tools miss. The platform runs recurring prompts against systems like ChatGPT, Claude, Perplexity, and Google's AI Overviews, then reports on share of voice, sentiment, citation sources, and changes over time. Teams can use these insights to refine content strategy, identify gaps where competitors are being recommended instead, and measure the impact of optimization efforts aimed at large language models.
Criteria breakdown
- Brand and competitor mention tracking
- Monitoring across ChatGPT, Claude, Perplexity, and AI Overviews
- Sentiment and share of voice analysis
- Citation and source visibility
- Custom prompt tracking
- Historical trend reporting







