你的 AI 身份与 AI 团队,可携带至任意 AI 平台。
登录 立即开始
菜单
创建 Agent of Me 探索风格 专业 Agent 社区 Agents 排行榜 AI 资讯
AI 平台 目录 模型矩阵 对比 我该用哪个 AI? 集成指南 设置 OpenClaw Prompt 匹配度
学习与工具 学习 数据问答 Agent 构建器 API
关于 关于我们 联系我们 免责声明
登录 立即开始
账号
你的 AI 身份,随时可携带

注册免费账号以构建你的档案。默认私密,除非你主动发布,否则不会分享任何内容。

立即开始 登录
深色模式

🧭 引导视图
对 prompt、系统指令、上下文窗口、token 感到陌生?我们在浏览过程中以通俗语言解释每一个术语,内置于相同页面中,无需额外跳转。

⚡ 专业视图
你已经懂得如何写 prompt。只给核心内容,简洁紧凑,无多余说明。这是默认视图。

界面语言

AI 排行榜

哪个模型综合领先,以及各类别的表现。这些是公开基准测评,YouTube 和 X 上的各种榜单图表均源于此。我们不发布任何自有评分:每一个数字都归属于产出它的基准测评,标注采集日期并回溯至原始来源。

Overall, 衡量内容: Human preference on open-ended chat.

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

局限性: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

来源: LMArena (formerly LMSYS Chatbot Arena) ↗ · 已发布 2026-08-11 · 底层数据 ↗ · Dataset released under Creative Commons Attribution 4.0 (CC BY 4.0). Reuse is permitted with attribution - credit LMArena and link to the leaderboard.

# 模型 组织 Arena score (human preference)
1 claude-fable-5 Anthropic 1506 95% CI 1501 to 1512 21,304 votes
2 claude-opus-4-6-thinking Anthropic 1505 95% CI 1501 to 1508 72,425 votes
3 claude-opus-4-7-thinking Anthropic 1502 95% CI 1498 to 1506 60,222 votes
4 muse-spark-1.2 (xHigh) Meta 1498 95% CI 1488 to 1509 3,278 votes
5 claude-opus-4-6 Anthropic 1498 95% CI 1494 to 1501 76,386 votes
6 claude-opus-5-high Anthropic 1494 95% CI 1489 to 1499 19,498 votes
7 claude-opus-4-7 Anthropic 1494 95% CI 1490 to 1498 61,308 votes
8 claude-opus-5-max Anthropic 1490 95% CI 1483 to 1497 9,419 votes
9 qwen3.8-max Alibaba 1490 95% CI 1482 to 1498 6,789 votes
10 muse-spark-1.1 Meta 1489 95% CI 1483 to 1494 16,648 votes
11 muse-spark Meta 1488 95% CI 1482 to 1494 13,600 votes
12 kimi-k3-max Moonshot AI 1487 95% CI 1481 to 1493 11,762 votes
13 gemini-3.1-pro-preview Google 1486 95% CI 1483 to 1490 94,814 votes
14 gemini-3-pro Google 1486 95% CI 1482 to 1489 41,509 votes
15 gemini-3.6-flash Google 1484 95% CI 1478 to 1490 13,559 votes
16 gpt-5.5-high OpenAI 1482 95% CI 1477 to 1486 55,210 votes
17 claude-opus-4-8-thinking Anthropic 1481 95% CI 1477 to 1486 40,409 votes
18 gpt-5.6-sol-xhigh OpenAI 1481 95% CI 1475 to 1487 15,304 votes
19 gemini-3.5-flash-high Google 1477 95% CI 1473 to 1482 25,613 votes
20 gpt-5.5 OpenAI 1477 95% CI 1473 to 1481 56,513 votes

柱状图按可见范围缩放,以保持细微差异的可读性,起点不为零。置信区间存在重叠时,各模型在统计上处于并列状态,请将顶部组视为一个整体,而非严格排名。

如何解读这些结果,以及为何各方结论不一致

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

注意: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

已发布 2026-08-11 · 实时排行榜 ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

注意: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

已发布 · 实时排行榜 ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

两种不同的单位,绝不合并为一张图表

人工偏好评分与正确率不能共用同一坐标轴,因此本页不会将两者混在同一张表中。一个模型可能在其中一项领先,却在另一项落后,这是真实的信号,说明它各自擅长什么,而非自相矛盾:前者问的是"人们更偏好哪个答案?",后者问的是"哪个答案是正确的?"

本页面不会做的事

它不会告诉你该用哪个 AI。基准测试榜首往往并非最适合你工作的工具,价格、可用性、集成能力、上下文长度,以及它对你指令的遵从程度,通常比一两分的差距更重要。 我该用哪个 AI? →

本月的领跑者: 你的个人档案和 AI 团队可随身携带,换新平台只需粘贴,无需从头重建。 集成指南

商业

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

金融

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

法律

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

营销

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

个人

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

房地产

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

研究

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

科技

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer