개요
주요 기능
- 합성 데이터셋 생성
- 자동화된 에이전트 평가 파이프라인
- 시나리오 및 회화 시뮬레이션
- customizable 평가 기술편
- LLM 앱용 백포드 레그테스트
- 수행성 벤치마크 및 리포팅
가격
- 모델
- Free
- 카테고리
- 관찰성 지표
- 평점
- 4.3 / 5 (6)
사용 사례
인공지능 에이전트 테스트
인공지능 에이전트에 대한 정교하고 테스트가 가능한 평가
장단점
장점
- 멀티 스텝 인공지능 에이전트를 위한 특별히 제작
- 규모에서 합성 테스트 데이터 생성
- 커스텀 메트릭스 및 평가기를 지원
- YC YC 으로 지원됨에 따라 활발한 개발
단점
- 개발자로 限된 non-개발자에 대한 테스트 도구로 주로 설계
- 새로운 플랫폼으로 evolve되는 기능
- 기존 스택에 맞춰 통합작업 필요
대결 기록
Pantheon에서 3회 대결.
Last 3 battles
리뷰
6개 평가의 평균.
리뷰를 작성하려면 로그인하세요.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.
Solid for our team
We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.
Compared a few options
Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.
Use it every day
Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.
Skeptical, then convinced
I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.
Q&A
Relari를 사용할 때 장점은 무엇인가요?
Relari는 다단계 AI 에이전트 평가에 특화되어 있으며, 대규모 합성 테스트 데이터를 생성하고 사용자 정의 메트릭과 평가자를 지원합니다. 또한 Y Combinator의 지원을 받으며 활발히 개발되고 있습니다.
Asked by Anya Sokolova · Jun 27, 2026
Relari는 비기술팀에게 적합한가요?
Relari는 주로 기술팀을 대상으로 설계되었으며, 비개발자에게는 적합하지 않을 수 있고 기존 스택에 맞추기 위해 통합 작업이 필요할 수 있습니다.
Asked by Hasan Demir · May 27, 2026
Relari는 어떤 기능을 지원하나요?
Relari는 합성 데이터셋 생성, 자동화된 에이전트 평가 파이프라인, 시나리오 시뮬레이션, 그리고 사용자 정의 평가 메트릭 등과 같은 기능을 지원합니다.
Asked by Marisol Pena · Apr 17, 2026
Relari는 어떤 용도로 사용되나요?
Relari는 AI 에이전트의 테스트, 평가 및 합성 데이터 생성을 위한 플랫폼으로, 체계적인 테스트와 평가를 통해 팀이 신뢰성을 향상시킬 수 있도록 돕습니다.
Asked by Hana Kobayashi · Mar 25, 2026
질문하기
관찰성 지표 대안

개발자 플랫폼을 위한統합된 AI 모델 플랫폼

자율 AI 에이전트와 지능형 시스템에 대한 보안 및 관리 플랫폼

AI 에이전트 평가, 모니터링, 성능 최적화를 위한 종합 플랫폼입니다.

브랜드가 ChatGPT, Claude, Perplexity, 및 Google AI Overviews에서 어떻게 표현되는지 모니터링하세요.

일련의 대용량 언어 모델(LLM)을 통합하고 PROMPT 연결을 용이하게하며 기업은 비즈니스 운영의 자동화를 지원하는 노 코딩 AI 워크 플로우 빌더입니다.

사업 자동화에 위한 인공지능 에이전트를 구축, 성과 평가, 향상하세요.
LLM 프로덕션 앱 모니터링, 디버깅, 개선에 필요한 모든 것을 한 곳에서 관리한다.

IT 운영을 위한 AI 에이전트가 사고 검출, 분류 및 해결을 가속화합니다.
Trending now

복잡한 PDF, 슬라이드 및 스프레드시트에서 구조화된 데이터를 추출할 수 있는 문서 지능 API입니다. 이들은 텍스트, 이미지, 탭 및 레이아웃 정보까지 모든 요소를 파싱 및 추출합니다.

광고 주도 답변, 클릭당 지불

정확한 과제 도움말과 전체 설명

여러 모드의 오픈 12B 모델은 128K 컨텍스트 윈도우가 있는 이미지를 처리하고 텍스트를 함께 사용합니다.
