Relari (YC W24) logo

Relari (YC W24)Testing, evaluation, and synthetic data generation platform for AI agents.

4.3 (6)
Daniel NikulshynReviewed by Daniel Nikulshyn·Updated July 2026

Overview

Relari is a developer platform focused on improving the reliability of AI agents through systematic testing and evaluation. It helps teams generate synthetic datasets, run automated evaluations, and benchmark agent performance across realistic scenarios before shipping to production. Backed by Y Combinator (W24), Relari targets engineering teams building complex LLM applications and multi-step agents where traditional QA falls short. Its tooling aims to bring software-engineering rigor—unit tests, regression checks, and measurable metrics—to non-deterministic AI systems. The platform supports custom evaluators, scenario simulation, and continuous monitoring, making it useful for both pre-launch validation and ongoing quality assurance of production agents.

Key features

  • Synthetic dataset generation
  • Automated agent evaluation pipelines
  • Scenario and conversation simulation
  • Customizable evaluation metrics
  • Regression testing for LLM apps
  • Performance benchmarking and reporting

Pricing

Model
Free
Rating
4.3 / 5 (6)

Use cases

Testing AI agents

Reliable and testable agents evaluation for AI agents.

Pros & Cons

Pros

  • Purpose-built for evaluating multi-step AI agents
  • Generates synthetic test data at scale
  • Supports custom metrics and evaluators
  • Backed by Y Combinator with active development

Cons

  • Primarily aimed at technical teams, not non-developers
  • Newer platform with an evolving feature set
  • May require integration work to fit existing stacks

Battle record

Across 3 battles in the Pantheon.

0
1st
0
2nd
0
3rd

Last 3 battles

Reviews

4.3

Average from 6 ratings.

5
2
4
4
3
0
2
0
1
0

Sign in to leave a review.

Fatima Zahra

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

Robert Ainsworth

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

DW

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Carlos Mendoza

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Yuki Mori

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

Leila Hassan

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Q&A

What are the pros of using Relari?

Relari is purpose-built for evaluating multi-step AI agents, generates synthetic test data at scale, and supports custom metrics and evaluators, with active development backed by Y Combinator.

Asked by Anya Sokolova · Jun 27, 2026

Is Relari suitable for non-technical teams?

Relari is primarily aimed at technical teams, not non-developers, and may require integration work to fit existing stacks.

Asked by Hasan Demir · May 27, 2026

What features does Relari support?

Relari supports features like synthetic dataset generation, automated agent evaluation pipelines, scenario simulation, and customizable evaluation metrics.

Asked by Marisol Pena · Apr 17, 2026

What is Relari used for?

Relari is a platform for testing, evaluation, and synthetic data generation for AI agents, helping teams improve their reliability through systematic testing and evaluation.

Asked by Hana Kobayashi · Mar 25, 2026

Ask a question

Observability alternatives