Relari (YC W24) logo

Relari (YC W24)Platforma za testiranje, ocenjevanje in generiranje sintetičnih podatkov za AI agente.

4.3 (6)
Daniel NikulshynPregledal Daniel Nikulshyn·Posodobljeno julij 2026

Pregled

Relari je razvijalna platforma, osredotočena na izboljšanje zanesljivosti AI agentov z sistematičnim testiranjem in ocenjevanjem. Omogoča ekipam ustvarjanje sintetičnih podatkovnih naborov, izvajanje samodejnih evaluacij ter primerjavo zmogljivosti agentov v realističnih scenarijih, preden se uvedejo v proizvodnjo. Podprto z Y Combinatorjem (W24), Relari cilja na inženirske ekipe, ki razvijajo kompleksne LLM aplikacije in večkorakove agente, kjer tradicionalni QA ne zadovoljuje zahtev. Njegova orodja se osredotočajo na vnos stroge programske inženirstva – enotne teste, regresijske preglede in merljive metrike – v nedeeterministične AI sisteme. Platforma podpira prilagojene ocenjevalce, simulacijo scenarijev in kontinuirano spremljanje, kar jo naredi uporabno tako za pre-lansiranje preverjanje kot za stalno zagotavljanje kakovosti produkcijskih agentov.

Ključne funkcije

  • Generiranje sintetičnih podatkovnih nizov
  • Samodejni postopki za ocenjevanje agentov
  • Simulacija scenarijev in pogovorov
  • Prilagodljivi kazalniki ocenjevanja
  • Regresijsko testiranje za LLM aplikacije
  • Primerjalno merjenje zmogljivosti in poročanje

Cene

Model
Free
Kategorija
Observabilnost
Ocena
4.3 / 5 (6)

Primeri uporabe

Testiranje AI agentov

Zanesljivo in testno ocenjevanje agentov za AI agente.

Prednosti in slabosti

Prednosti

  • Zasnovan posebej za ocenjevanje večstopniških AI agentov
  • Generira sintetične testne podatke v obsegu
  • Podpira prilagojene kazalnike in ocenjevalce
  • Podprt z Y Combinatorjem z aktivnim razvojem

Slabosti

  • Primarno namenjen tehničnim ekipam, ne pa uporabnikom brez programiranja
  • Novejša platforma z razvijajočim se nabirčkom funkcij
  • Morda zahteva delo za integracijo v obstoječe okolje

Rekord bitk

V 3 bitkah v Panteonu.

0
1.
0
2.
0
3.

Last 3 battles

Ocene

4.3

Povprečje iz 6 ocen.

5
2
4
4
3
0
2
0
1
0

Prijavi se za oddajo ocene.

Fatima Zahra

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

Robert Ainsworth

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

DW

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Carlos Mendoza

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Yuki Mori

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

Leila Hassan

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Vprašanja

Kakšni so prednosti uporabe Relari?

Relari je zasnovan posebej za ocenjevanje večkoraknih AI agentov, generira sintetične testne podatke v obsegu, podpira prilagojene metrike in ocenjevalce ter ga razvija aktivno pod Y Combinator.

Asked by Anya Sokolova · Jun 27, 2026

Ali je Relari primeren za netehnike ekipe?

Relari je predvsem namenjen tehničnim ekipam, ne pa netehnikom, in je morda potrebna integracija, da se prilega obstoječim sistemom.

Asked by Hasan Demir · May 27, 2026

Katere funkcionalnosti podpira Relari?

Relari podpira funkcionalnosti, kot so generiranje sintetičnih nabora podatkov, avtomatizirani cevovodi za ocenjevanje agentov, simulacija scenarijev in prilagodljivi metriki ocenjevanja.

Asked by Marisol Pena · Apr 17, 2026

Za kaj se uporablja Relari?

Relari je platforma za testiranje, ocenjevanje in generiranje sintetičnih podatkov za AI agente, ki ekipam pomaga izboljšati zanesljivost z sistematičnim testiranjem in ocenjevanjem.

Asked by Hana Kobayashi · Mar 25, 2026

Postavi vprašanje

Alternative za Observabilnost