Past battle · 2026-05-10 UTC

Research AI Agents Showdown — May 10, 2026

From the Research AI Agents category. 35 marks placed across 10 fighters. Google AI Co-Scientist took the crown.

Final standings

The line-up

The fighters

Profiles of every tool that competed in this battle, ranked by their final score.

1Google AI Co-Scientist logo

Google AI Co-Scientist

A multi-agent AI system that collaborates with scientists to generate hypotheses and accelerate biomedical research.

4.5 (4)
Free
Google AI Co-Scientist screenshot

The Google AI Co-Scientist is a multi-agent AI system designed to collaborate with scientists and accelerate biomedical research. It is built with Gemini 2.0 and functions as a virtual scientific collaborator to help generate novel hypotheses and research proposals. The system is intended to mirror the reasoning process underpinning the scientific method and uncover new, original knowledge. Given a scientist's research goal specified in natural language, the AI Co-Scientist generates novel research hypotheses, a detailed research overview, and experimental protocols. It uses a coalition of specialized agents, including Generation, Reflection, Ranking, Evolution, Proximity, and Meta-review, which use automated feedback to iteratively generate, evaluate, and refine hypotheses. The system allows scientists to interact with it in various ways, including providing seed ideas for exploration or feedback on generated outputs in natural language. The AI Co-Scientist also uses tools like web-search and specialized AI models to enhance the grounding and quality of generated hypotheses. The system is designed to flexibly scale compute and iteratively improve its scientific reasoning towards the specified research goal. The AI Co-Scientist is particularly useful for scientists who face challenges in navigating the rapid growth in scientific publications and integrating insights from unfamiliar domains. By leveraging recent AI advances, including the ability to synthesize across complex subjects and perform long-term planning and reasoning, the AI Co-Scientist aims to accelerate the clock speed of scientific and biomedical discoveries. The system has been tested in various applications, including gene transfer discovery and transfer re-discovery. While the AI Co-Scientist presents several advantages, its effectiveness depends on the quality of the input research goal and the scientist's ability to provide meaningful feedback. The AI Co-Scientist is a new tool that has the potential to significantly impact the scientific community, but its limitations and potential biases need to be carefully evaluated. The development of the AI Co-Scientist is an ongoing effort, and its future applications and improvements will depend on continued research and testing.

Criteria breakdown

Ease of use1
Value for money1
Features & power1
Integrations1
Support & docs1
Reliability1
  • Hypothesis generation
  • Research overview creation
  • Experimental protocol development
  • Specialized agents for scientific reasoning
  • Automated feedback loop
  • Web-search integration
2GLM-4.6V logo

GLM-4.6V

Open-source multimodal GLM from Z.ai unifying vision, text, and tool calling for long-context reasoning, search, coding, and UI-to-code.

4.3 (6)
Free
GLM-4.6V screenshot

GLM-4.6V is a multimodal large language model that scales its context window to 128k tokens and achieves state-of-the-art performance in visual understanding. It integrates native Function Calling capabilities, enabling a unified technical foundation for multimodal agents in real-world business scenarios. GLM-4.6V can accept multimodal inputs and automatically generate high-quality, structured image-text interleaved content, perform complex document understanding, and perform visual search and retrieval. This model provides several capabilities and scenarios, including intelligent image-text content creation and layout, visual web search and rich media report generation, and long-context understanding. It can accept mixed text-image inputs, generate high-quality content, and perform a visual audit on candidate images. GLM-4.6V's multimodal search-and-analysis workflow enables seamless movement from visual perception to online retrieval and structured report generation. It maintains multimodal context awareness and performs reasoning grounded in both textual and visual information. This model provides several capabilities and scenarios, including intent recognition and search planning, multimodal results alignment, and reasoning with rich media report generation.

Criteria breakdown

Ease of use1
Value for money1
Features & power1
Integrations1
Support & docs1
Reliability0
  • Multimodal input support (images, text, files)
  • Native multimodal tool calling
  • 128k context length
  • Function Calling capabilities
  • Long-context understanding
  • Visual comprehension and tool retrieval
3AlphaSense logo

AlphaSense

An AI-powered market intelligence platform that autonomously sources expert insights and delivers analyst-level research at scale.

4.4 (5)
Contact
AlphaSense screenshot

AlphaSense is an AI-powered market intelligence platform that provides trusted insights to guide business decisions. It offers a comprehensive collection of curated sources, including Tegus expert transcripts, broker research, and company filings. The platform uses advanced AI models to generate high-quality insights with sentence-level citations. AlphaSense integrates workflows, eliminating fragmented sources and providing streamlined, connected insights. Its GenAI capabilities provide real-time insights that ground decisions in reliable expertise. The platform is designed for investment banking, hedge funds, private equity, asset management, consulting, life sciences, and other industries. AlphaSense differentiates itself with its highly synthesized insights, synthesized from an unrivaled content set. The platform also features customizable agents to automate repeatable research tasks, transforming findings into executive-ready reports and slides. Its Generative Search provides a 360-degree view of markets and companies, connecting qualitative insights with structured financial data. One of AlphaSense's key strengths is its ability to empower users to make bold moves with confidence and speed. The platform gives teams the precise answers, unique perspectives, and critical updates they need to stay ahead of the competition.,

Criteria breakdown

Ease of use0
Value for money0
Features & power1
Integrations1
Support & docs1
Reliability1
  • Curated sources
  • Tegus expert transcripts
  • Broker research
  • Company filings
  • Private and public financial data
  • GenAI for real-time insights
4Autoresearch logo

Autoresearch

An open-source project that lets AI agents autonomously run LLM training experiments and keep the best model changes.

4.8 (5)
Free
Autoresearch screenshot

Autoresearch is an open-source project that enables AI agents to autonomously run LLM training experiments and retain the best model changes. The project allows users to set up a small but real LLM training environment and let an AI agent experiment with it overnight, modifying the code, training for a short period, and checking if the results improve. The goal is to automate the research process, letting the AI agent explore different model architectures, hyperparameters, and optimization strategies without human intervention. The project includes a simplified single-GPU implementation of nanochat and provides a basic structure for programming the AI agent's research process using Markdown files. The project is designed to be extensible, allowing users to add more agents and improve the research process over time.

Criteria breakdown

Ease of use1
Value for money1
Features & power1
Integrations1
Support & docs0
Reliability0
  • Autonomous LLM training experiments
  • AI agent-driven research process
  • Single-GPU implementation of nanochat
  • Markdown-based programming for the research process
  • 5-minute training time budget with evaluation metric (val_bpb)
5Cell2Sentence logo

Cell2Sentence

Open-source framework that turns single-cell gene expression into 'cell sentences' so LLMs can analyze and generate biology insights.

4.3 (4)
Free
Cell2Sentence screenshot

Cell2Sentence is an open-source framework that transforms single-cell gene expression data into 'cell sentences' for analysis and insight generation by Large Language Models (LLMs). It proposes a rank-ordering transformation of expression vectors into cell sentences, which are space-separated gene names ordered by descending expression. This allows LLMs to natively model single-cell RNA sequencing (scRNA-seq) data using natural language. The framework includes the C2S-Scale models, which unify transcriptomic and textual data and enable advanced single-cell tasks such as perturbation prediction, dataset summarization, cluster captioning, and biological question answering. The C2S-Scale models are available on Hugging Face and are based on architectures like Pythia and Gemma-2. Cell2Sentence is aimed at researchers and scientists working with single-cell transcriptomics data. The framework has been updated with new models and features, including support for fine-tuning on custom prompt templates and multi-cell prompt formatting. It also includes a suite of Pythia models for cell type prediction, cell type conditioned generation, and a diverse multi-cell multi-task model trained on over 57 million human and mouse cells. The Cell2Sentence framework is documented and has tutorials for use, including examples of fine-tuning and multi-cell prompt formatting. The development of Cell2Sentence involves the van Dijk Lab and has been published in a preprint on bioRxiv. Cell2Sentence enables next-generation single-cell discovery with LLMs.

Criteria breakdown

Ease of use1
Value for money1
Features & power1
Integrations1
Support & docs0
Reliability0
  • Transformation of expression vectors into cell sentences
  • C2S-Scale models for advanced single-cell tasks
  • Support for fine-tuning on custom prompt templates
  • Multi-cell prompt formatting
  • Pre-trained models based on Pythia and Gemma-2 architectures
6THEUS (Aigora) logo

THEUS (Aigora)

Enterprise AI research system that turns proprietary studies into traceable insights with fact IDs, citations, and expert avatars for synthesis.

4.6 (5)
Freemium
THEUS (Aigora) screenshot

THEUS is an enterprise AI research system that transforms proprietary studies into traceable insights with fact IDs, citations, and expert avatars. It's built on Google's Agent Development Kit (ADK) and uses bidirectional modeling for data-grounded knowledge agents. Users can extract, synthesize, and explore their research data, and every claim is backed by a specific Fact ID with page-level citations. THEUS offers two ways to explore: generating new insights with Dr. Reed or analyzing existing knowledge with Dr. Sinclair. This platform is designed for VPs of Insights, Global Sensory Leads, and Innovation leaders evaluating strategic capability build-out. It uses a purpose-built research methodology, not generic enterprise AI, and creates data-grounded knowledge agents trained on actual institutional data. With 85% of new products failing within two years (Nielsen), evidence-grounded exploration dramatically improves decision quality. THEUS offers a solo seat, which includes features such as up to 30 documents ingested per month, up to 1,500 pages parsed per month, and up to 500 questions to Dr. Sinclair per month. One of the key benefits of THEUS is its ability to turn decades of proprietary data into traceable, decision-ready intelligence with AI-powered knowledge agents.

Criteria breakdown

Ease of use1
Value for money0
Features & power1
Integrations1
Support & docs1
Reliability0
  • Extract, synthesize, and explore research data
  • Generate new insights with Dr. Reed or analyze existing knowledge with Dr. Sinclair
  • Bidirectional modeling for data-grounded knowledge agents
  • Connect product and consumer knowledge for bidirectional insight
  • Supports cross-study synthesis and research gap analysis
7StartupValidator logo

StartupValidator

AI startup idea validator using live Twitter signals and web research to generate a scored verdict with risks, strengths, and suggested pivots.

4.5 (6)
Free
StartupValidator screenshot

StartupValidator is an AI tool designed to validate startup ideas using live Twitter signals and web research. It aims to provide a scored verdict on the viability of a startup idea, highlighting risks, strengths, and suggesting potential pivots. This tool appears to be targeted at entrepreneurs and startup founders looking to assess the feasibility of their ideas before investing significant resources. By leveraging real-time data from Twitter and the web, StartupValidator offers a data-driven approach to startup validation. However, without more specific information on its methodology and data sources, the tool's reliability and accuracy are uncertain. The tool likely integrates with Twitter's API and possibly other web data sources to gather relevant information.

Criteria breakdown

Ease of use0
Value for money1
Features & power0
Integrations1
Support & docs0
Reliability1
  • Live Twitter signal analysis
  • Web research integration
  • Scored verdict on startup viability
  • Risk and strength identification
  • Suggested pivots for improvement
8Kosmos logo

Kosmos

Autonomous AI scientist for long research campaigns that analyzes data and literature to produce fully cited scientific reports.

4.8 (6)
Freemium
Kosmos screenshot

Kosmos is an autonomous AI scientist designed for biopharma R&D teams and researchers. It analyzes data and literature to produce fully cited scientific reports, accelerating the drug development process from hypothesis to registration in days, not months. Kosmos works autonomously, reading literature, generating hypotheses, steering investigations, and executing scientific workflows. It supports the full lifecycle of drug development, from target identification through clinical development and filing. The AI tool collaborates with human scientists through Slack, Teams, and email, and learns from the organization's knowledge, building on experimental history and proprietary scientific data.

Criteria breakdown

Ease of use1
Value for money1
Features & power0
Integrations0
Support & docs0
Reliability0
  • Autonomous hypothesis generation
  • Literature analysis
  • Data analysis
  • Scientific workflow execution
  • Collaboration through Slack, Teams, and email
  • Model agnostic for future state-of-the-art models
9Table Agent logo

Table Agent

AI-powered data assistant for tabular research Automated data gathering on any topic Customizable data schemas AI agent-driven data search

4.5 (6)
Freemium

Table Agent is an AI-powered data assistant designed to aid in tabular research. It is intended to automate the process of gathering data on any given topic, making research more efficient. The tool allows for customizable data schemas, which enables users to tailor their data collection to specific needs. This flexibility is beneficial for a wide range of applications and industries where data structure can vary significantly. At the core of Table Agent's functionality is its AI agent-driven data search capability. This means the tool utilizes artificial intelligence to actively seek out and compile relevant data, potentially reducing the time and effort required for manual searches. While the specifics of its workflow and integrations are not detailed, the concept of an AI-driven data assistant suggests it could seamlessly integrate with various data analysis tools, enhancing the overall research process. The strengths of Table Agent would likely include its ability to automate tedious data collection tasks and its adaptability to different data structures. However, without more specific information, it's challenging to determine its limitations or how it compares to alternative data assistant tools.

Criteria breakdown

Ease of use1
Value for money0
Features & power0
Integrations0
Support & docs0
Reliability1
  • Automated data gathering
  • Customizable data schemas
  • AI agent-driven data search
  • Quick and accurate data retrieval
10AMIE logo

AMIE

A multimodal AI diagnostic agent that conducts clinical conversations and interprets medical images for accurate diagnoses.

4.7 (6)
Free
AMIE screenshot

AMIE (Articulate Medical Intelligence Explorer) is a research AI diagnostic agent developed by Google DeepMind and Google Research. It is designed to conduct clinical conversations and interpret medical images for accurate diagnoses. Initially, AMIE was a text-based medical diagnostic conversational AI agent published in Nature. The latest advancement, multimodal AMIE, integrates the ability to intelligently request, interpret, and reason about visual medical information during clinical conversations. This development aims to improve diagnostic accuracy by incorporating multimodal data, such as images and documents, into the diagnostic process. AMIE uses a state-aware reasoning framework and is built on multimodal Gemini models. It has been evaluated through expert assessments, including Objective Structured Clinical Examinations (OSCEs), comparing its performance to primary care physicians (PCPs) in various patient scenarios.

Criteria breakdown

Ease of use1
Value for money0
Features & power0
Integrations0
Support & docs0
Reliability0
  • Multimodal diagnostic dialogue
  • Visual medical information interpretation
  • State-aware reasoning framework
  • Integration with Gemini models
  • Simulation environment for dialogue evaluation