Your AI identity and your AI team, portable to every AI platform.
Sign in Get Started
Menu
Build an Agent of Me Explore Styles Professional Agents Community Agents Leaderboard AI News
AI Platforms Directory Model Matrix Compare Which AI should I use? Integration Guides Set up OpenClaw Prompt Fit
Learn & Tools Learn Ask the Data Agent Builder API
About About us Contact Disclaimers
Sign in Get Started
Account
Your AI identity, portable

Create a free account to build your profile. Private by default. Nothing is shared unless you publish it.

Get Started Sign in
Dark mode

🧭 Guided View
New to prompts, system instructions, context windows, tokens? We explain every term as you browse, in plain English. Same pages, with the help built in.

⚡ Expert View
You already know how prompting works. Just the substance, clean and compact, with no extra explanations. This is the default view.

Interface language

AI leaderboard

Which model leads, overall and by category. These are the public benchmarks the charts you see on YouTube and X are built from. We publish no scores of our own: every number belongs to the benchmark that produced it, captured on a date and credited back to the source.

Hard Prompts, what this measures: Human preference on open-ended chat.

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

Limitations: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

Source: LMArena (formerly LMSYS Chatbot Arena) ↗ · published 2026-08-11 · underlying data ↗ · Dataset released under Creative Commons Attribution 4.0 (CC BY 4.0). Reuse is permitted with attribution - credit LMArena and link to the leaderboard.

# Model Organization Arena score (human preference)
1 claude-opus-4-6-thinking Anthropic 1533 95% CI 1529 to 1538 45,891 votes
2 claude-fable-5 Anthropic 1533 95% CI 1526 to 1539 13,966 votes
3 claude-opus-4-6 Anthropic 1527 95% CI 1522 to 1531 49,442 votes
4 claude-opus-4-7-thinking Anthropic 1526 95% CI 1521 to 1531 40,299 votes
5 claude-opus-5-high Anthropic 1520 95% CI 1514 to 1527 13,016 votes
6 claude-opus-4-7 Anthropic 1518 95% CI 1514 to 1523 40,987 votes
7 claude-opus-5-max Anthropic 1517 95% CI 1508 to 1525 6,207 votes
8 qwen3.8-max Alibaba 1516 95% CI 1506 to 1525 4,653 votes
9 kimi-k3-max Moonshot AI 1515 95% CI 1508 to 1522 7,689 votes
10 muse-spark-1.2 (xHigh) Meta 1512 95% CI 1499 to 1525 2,149 votes

Bars are scaled across the visible range so small differences stay readable. They do not start at zero. Where confidence intervals overlap, the models are statistically tied: read the top group as a group, not a strict order.

How to read these, and why they disagree

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

Watch out for: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

Published 2026-08-11 · live leaderboard ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Watch out for: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Published · live leaderboard ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

Two different units, never one chart

A human-preference rating and a percent-correct score cannot share an axis, so this page never mixes them in one table. A model can top one and not the other, and that is a real signal about what it is good at rather than a contradiction: one asks “which answer did people prefer?”, the other asks “which answer was right?”.

What this page will not do

It will not tell you which AI to use. The benchmark leader is often not the right tool for your work, price, availability, integrations, context length and how well it follows YOUR instructions usually matter more than a point or two of score. Which AI should I use? →

Whoever leads this month: your profile and your AI team are portable, moving to a new leader is a paste, not a rebuild. Integration guides

Business

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

Finance

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

Legal

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

Marketing

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

Personal

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

Real Estate

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

Research

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

Technology

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer