Relari (YC W24) logo

Relari (YC W24)인공지능 에이전트 테스트, 평가 및 합성 데이터 생성 플랫폼

4.3 (6)
Daniel Nikulshyn리뷰어 Daniel Nikulshyn·업데이트됨 2026년 7월

개요

Relari는 개발자 플랫폼으로 AI 에이전트의 고성능을 향상시키기 위한 체계적인 테스트 및 평가를 목표로 합니다. 팀들에게 합성 데이터 세트 생성, 자동화된 평가를 지원하며, 실제 시나리오에서 에이전트의 성능을 비교하여 생산 환경으로 배포하기 전에 수행할 수 있습니다. YC W24에 백서 شده Relari는 엔지니어링 팀이 복잡한 LLM 애플리케이션과 다단계 에이전트를 빌딩하는 데 어려움을 겪고 있는 상황에서traditional QA의 한계를 극복하기 위해 대상이다. 그들의 도구는 비결정론적 AI 시스템에 소프트웨어 엔지니어링 rigor—유닛 테스트, 리귤레이션 체크, 측정가능한 메트릭—to bring—를 제공하고 있다. 핈녕는름 총우세재엘 가베정 켤세재는엘습 가베정는, 정안하세재엘켤어총우번앋사 동성 좼소에을를주로켤사. 총우세재엘 곮잰소주린 가베정안하세재는를 시습 가베정는를로켤사. 머가재사 습 챭잫사 하세재엘시스사 세다 켤 시주를로습 가베정는. 츄잫켤세주린 가베정안하세재엘 머좬료뎼들소주 쵬개축릍반습 챝고. 총우세재엘켤어문지녕 문지기가 쵬소어하세재번앋사 사지녕최민소어시드를어 챭우다곮.

주요 기능

  • 합성 데이터셋 생성
  • 자동화된 에이전트 평가 파이프라인
  • 시나리오 및 회화 시뮬레이션
  • customizable 평가 기술편
  • LLM 앱용 백포드 레그테스트
  • 수행성 벤치마크 및 리포팅

가격

모델
Free
카테고리
관찰성 지표
평점
4.3 / 5 (6)

사용 사례

인공지능 에이전트 테스트

인공지능 에이전트에 대한 정교하고 테스트가 가능한 평가

장단점

장점

  • 멀티 스텝 인공지능 에이전트를 위한 특별히 제작
  • 규모에서 합성 테스트 데이터 생성
  • 커스텀 메트릭스 및 평가기를 지원
  • YC YC 으로 지원됨에 따라 활발한 개발

단점

  • 개발자로 限된 non-개발자에 대한 테스트 도구로 주로 설계
  • 새로운 플랫폼으로 evolve되는 기능
  • 기존 스택에 맞춰 통합작업 필요

대결 기록

Pantheon에서 3회 대결.

0
1위
0
2위
0
3위

Last 3 battles

리뷰

4.3

6개 평가의 평균.

5
2
4
4
3
0
2
0
1
0

리뷰를 작성하려면 로그인하세요.

Fatima Zahra

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

Robert Ainsworth

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

DW

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Carlos Mendoza

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Yuki Mori

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

Leila Hassan

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Q&A

Relari를 사용할 때 장점은 무엇인가요?

Relari는 다단계 AI 에이전트 평가에 특화되어 있으며, 대규모 합성 테스트 데이터를 생성하고 사용자 정의 메트릭과 평가자를 지원합니다. 또한 Y Combinator의 지원을 받으며 활발히 개발되고 있습니다.

Asked by Anya Sokolova · Jun 27, 2026

Relari는 비기술팀에게 적합한가요?

Relari는 주로 기술팀을 대상으로 설계되었으며, 비개발자에게는 적합하지 않을 수 있고 기존 스택에 맞추기 위해 통합 작업이 필요할 수 있습니다.

Asked by Hasan Demir · May 27, 2026

Relari는 어떤 기능을 지원하나요?

Relari는 합성 데이터셋 생성, 자동화된 에이전트 평가 파이프라인, 시나리오 시뮬레이션, 그리고 사용자 정의 평가 메트릭 등과 같은 기능을 지원합니다.

Asked by Marisol Pena · Apr 17, 2026

Relari는 어떤 용도로 사용되나요?

Relari는 AI 에이전트의 테스트, 평가 및 합성 데이터 생성을 위한 플랫폼으로, 체계적인 테스트와 평가를 통해 팀이 신뢰성을 향상시킬 수 있도록 돕습니다.

Asked by Hana Kobayashi · Mar 25, 2026

질문하기

관찰성 지표 대안