Relari (YC W24)Platforma do testowania, oceny i generowania danych syntetycznych dla agentów AI.
Przegląd
Kluczowe funkcje
- Generowanie zestawów danych syntetycznych
- Automatyczne potoki oceny agentów
- Symulacja scenariuszy i konwersacji
- Dostosowywalne metryki oceny
- Testy regresji dla aplikacji LLM
- Benchmarking wydajności i raportowanie
Cennik
- Model
- Free
- Kategoria
- Widoczność Systemów
- Ocena
- 4.3 / 5 (6)
Zastosowania
Testowanie agentów AI
Niezawodna i testowalna ocena agentów AI
Plusy i minusy
Plusy
- Specjalnie zaprojektowane do oceny wieloetapowych agentów AI
- Generuje dane testowe syntetyczne w dużej skali
- Obsługuje niestandardowe metryki i ewaluatory
- Wspierane przez Y Combinator z aktywnym rozwojem
Minusy
- Przede wszystkim skierowane do zespołów technicznych, nie do osób nietechnicznych
- Nowsza platforma z rozwijającym się zestawem funkcji
- Może wymagać pracy integracyjnej, aby dopasować do istniejących stosów
Wynik bitew
W 3 bitwach w Panteonie.
Last 3 battles
Recenzje
Średnia z 6 ocen.
Zaloguj się, aby zostawić recenzję.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.
Solid for our team
We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.
Compared a few options
Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.
Use it every day
Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.
Pytania i odpowiedzi
What are the pros of using Relari?
Relari is purpose-built for evaluating multi-step AI agents, generates synthetic test data at scale, and supports custom metrics and evaluators, with active development backed by Y Combinator.
Asked by Anya Sokolova · Jun 27, 2026
Is Relari suitable for non-technical teams?
Relari is primarily aimed at technical teams, not non-developers, and may require integration work to fit existing stacks.
Asked by Hasan Demir · May 27, 2026
What features does Relari support?
Relari supports features like synthetic dataset generation, automated agent evaluation pipelines, scenario simulation, and customizable evaluation metrics.
Asked by Marisol Pena · Apr 17, 2026
What is Relari used for?
Relari is a platform for testing, evaluation, and synthetic data generation for AI agents, helping teams improve their reliability through systematic testing and evaluation.
Asked by Hana Kobayashi · Mar 25, 2026
Zadaj pytanie
Alternatywy dla Widoczność Systemów

Zunifikowana platforma deweloperska do budowania, monitorowania i skalowania aplikacji LLM.

Platforma bezpieczeństwa i zarządzania dla autonomicznych agentów AI i systemów inteligentnych.

Platforma end-to-end do oceny, monitorowania i ulepszania agentów AI

Monitoruj, jak Twoja marka jest prezentowana w ChatGPT, Claude, Perplexity i Google AI Overviews.

Kreator przepływów AI bez kodu, który umożliwia firmom automatyzację operacji poprzez integrację wielu dużych modeli językowych (LLM) i płynne łączenie promptów.

Twórz, oceniaj i udoskonalaj agenty AI w automatyzacji procesów biznesowych
All-in-one platforma obserwacyjna do monitorowania, debugowania i ulepszania produkcyjnych aplikacji LLM.

Agent AI dla operacji IT przyspieszający wykrywanie incydentów, triage oraz ich rozwiązywanie.
Trending now

API inteligencji dokumentowej, które analizuje, dzieli, wykonuje OCR i wyodrębnia strukturalne dane z złożonych PDF-ów, slajdów i arkuszy kalkulacyjnych.

Sponsorowane odpowiedzi, płatne za kliknięcie.

Precyzyjna pomoc z zadaniami domowymi z pełnymi wyjaśnieniami

Otwarte, multimodalne model 12B obsługujący przeplatanie obrazów i tekstu przy 128K oknie kontekstowym
