Relari (YC W24)Testing, evaluation, and synthetic data generation platform for AI agents.
Overview
Key features
- Synthetic dataset generation
- Automated agent evaluation pipelines
- Scenario and conversation simulation
- Customizable evaluation metrics
- Regression testing for LLM apps
- Performance benchmarking and reporting
Pricing
- Model
- Free
- Category
- Observability
- Rating
- 4.3 / 5 (6)
Use cases
Testing AI agents
Reliable and testable agents evaluation for AI agents.
Pros & Cons
Pros
- Purpose-built for evaluating multi-step AI agents
- Generates synthetic test data at scale
- Supports custom metrics and evaluators
- Backed by Y Combinator with active development
Cons
- Primarily aimed at technical teams, not non-developers
- Newer platform with an evolving feature set
- May require integration work to fit existing stacks
Battle record
Across 3 battles in the Pantheon.
Last 3 battles
Reviews
Average from 6 ratings.
Sign in to leave a review.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.
Solid for our team
We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.
Compared a few options
Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.
Use it every day
Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.
Q&A
What are the pros of using Relari?
Relari is purpose-built for evaluating multi-step AI agents, generates synthetic test data at scale, and supports custom metrics and evaluators, with active development backed by Y Combinator.
Asked by Anya Sokolova · Jun 27, 2026
Is Relari suitable for non-technical teams?
Relari is primarily aimed at technical teams, not non-developers, and may require integration work to fit existing stacks.
Asked by Hasan Demir · May 27, 2026
What features does Relari support?
Relari supports features like synthetic dataset generation, automated agent evaluation pipelines, scenario simulation, and customizable evaluation metrics.
Asked by Marisol Pena · Apr 17, 2026
What is Relari used for?
Relari is a platform for testing, evaluation, and synthetic data generation for AI agents, helping teams improve their reliability through systematic testing and evaluation.
Asked by Hana Kobayashi · Mar 25, 2026
Ask a question
Observability alternatives

Unified developer platform for building, monitoring, and scaling LLM applications.

Security and governance platform for autonomous AI agents and intelligent systems.

End-to-end platform for evaluating, monitoring, and improving AI agents

Monitor how your brand appears across ChatGPT, Claude, Perplexity, and Google AI Overviews.

A no-code AI workflow builder that enables businesses to automate operations by integrating multiple large language models (LLMs) and connecting prompts seam...

Build, evaluate, and improve AI agents for business automation
All-in-one observability platform to monitor, debug, and improve production LLM apps.

AI agent for IT operations that speeds up incident detection, triage, and resolution.
Trending now

Accurate Homework Help with Full Explanations

Document intelligence API that parses, splits, OCRs, and extracts structured data from complex PDFs, slides, and spreadsheets.

Open multimodal 12B model handling interleaved images and text with a 128K context window.

Sponsored answers, paid per click.
