Skip to main content.
AllNewsProductResearch
Text Arena and Search Arena scores with the factuality adjustment on

Factuality in the Arena

Factuality remains one of the most persistent questions users face when using AI models. Today, we are launching a leaderboard that ranks models not only by human preference, but also by the factual accuracy of their responses.

Read Article
Research
Product
Arena Team—14 Jul 2026
Post Card
Read Article

Build, Deploy, and Evaluate with Fullstack Code Arena

Code Arena has evolved beyond frontend prototyping into a complete fullstack AI development platform. Explore capabilities like databases, authentication, third-party integrations, and deployments, and see how Fullstack Code Arena helps developers, entrepreneurs, and AI labs build, deploy, and evaluate production-ready applications.

Product
Arena Team—2 Jul 2026
Arena reaches $100M in eight months
Read Article

Arena Reaches $100M in 8 Months

Arena achieves $100M annualized run rate in eight months, driven by 10M+ monthly users who have contributed 700M+ conversations and 82M+ votes to build the world's largest human-preference dataset for AI evaluation.

News
Anastasios Angelopoulos—29 Jun 2026
Agent Arena: Causal Evaluation of Agents in the Real World
Read Article

Agent Arena: Causal Evaluation of Agents in the Real World

Agents are increasingly doing real work. The resulting task distribution has greatly expanded. We desire an agent evaluation that scales along with usage and capability.

Product
Research
Arena Team—4 Jun 2026
Empowering Users to Get More Done With Agent Mode
Read Article

Empowering Users to Get More Done With Agent Mode

The future of AI is not single-modality chat; it is in powerful agentic capabilities. Today, we are excited to introduce Agent Mode, designed to help everyone from everyday users looking to get more done to entrepreneurs looking to maximize agentic efficacy across complex use cases.

Product
Arena Team—4 Jun 2026
New Categories for Web Development in Code Arena
Read Article

New Categories for Web Development in Code Arena

AI coding models are increasingly used to build web apps, but aggregated leaderboards obscure key performance differences. After analyzing 250k+ Code Arena prompts, we identified major front-end task categories and built new leaderboard views to compare model strengths and weaknesses.

Product
Arena Team—8 May 2026
Multimodal Max
Read Article

Multimodal Max

Max, Arena's model router powered by 5M+ community votes, is now multimodal. Starting today, Max will be available as the default option in direct chat for all modalities, with expanded capabilities including search, vision, image generation, image editing, and front-end coding. Similar to our original Max for text, the multimodal variants are latency-controlled to provide a fast and performant experience. Try it now at arena.ai/max!

Product
Arena Team—5 May 2026

Subscribe
to Arena news and research

Insights at the frontier of AI.

Invalid email address
Arena Leaderboard Dataset
Read Article

Arena Leaderboard Dataset

For almost three years, Arena has been publishing leaderboards covering frontier AI capabilities across 10 arenas, dozens of categories, and hundreds of models, and today we're releasing the entire history of those leaderboards as a public-access dataset.

Research
Arena Team—2 Apr 2026
March 2026: Arena Updates across Product, Leaderboard Rankings & Research
Read Article

March 2026: Arena Updates across Product, Leaderboard Rankings & Research

March 2026 brought major updates to the Arena leaderboard, including new rankings across document, video, text, and code models. In this monthly roundup, we break down the best AI models, latest LLM benchmarks, and key trends shaping AI evaluation.

News
Product
Arena Team—31 Mar 2026
Inside BullshitBench: AI Models and Nonsense Detection
Read Article

Inside BullshitBench: AI Models and Nonsense Detection

AI failures like hallucinations are well documented. A less examined problem is that models will accept nonsensical premises without question and produce confident, detailed answers to questions that have no valid answer. BullshitBench measures whether models challenge broken premises or play along. We tested over 80 models from all major providers. Clear pushback rates range from 2% to 91%.

News
Peter Gostev—18 Mar 2026
Supporting Independent Research in AI Evaluation
Read Article

Supporting Independent Research in AI Evaluation

Arena’s Academic Partnerships Program provides funding and support for independent research advancing the scientific foundations of AI evaluation.

Research
News
Arena Team—10 Feb 2026
Image Arena Improvements: New Categories & Quality Filtering
Read Article

Image Arena Improvements: New Categories & Quality Filtering

After analyzing over 4 million user prompts, it is clear that a single global leaderboard no longer captures the full picture. Today we introduce new categories and quality filtering.

Product
Arena Team—9 Feb 2026
Introducing Max
Read Article

Introducing Max

Today we are releasing Max, Arena's model router powered by our community’s 5+ million real-world votes. Max acts as an intelligent orchestrator—it routes each user prompt to the most capable model for that specific prompt.

Research
Arena Team—4 Feb 2026
Try Arena
Try Arena
Leaderboard Rankings
Overall
Agent
Text
WebDev
Image-to-WebDev
Text to Image
Image Edit
Text to Video
Image to Video
Video Edit
Vision
Document
Search
Use Cases
Chat with AI
Build Apps & Websites
Write & Edit Text
Search the Web
Generate Images
Generate Videos
Chose any Model
Compare Models Side by Side
Complete Multi-step Tasks
Company
About Us
How It Works
Careers
Changelog
Help Center
FAQ
Blog
Follow
X
LinkedIn
YouTube
Discord
TermsPrivacyCookies
Try Arena
Try Arena

Ⓒ 2026 Arena Intelligence Inc.