Bảng xếp hạng AI
Model nào dẫn đầu, tổng thể và theo từng danh mục. Đây là những benchmark công khai mà các biểu đồ bạn thấy trên YouTube và X được xây dựng từ đó. Chúng tôi không công bố điểm số của riêng mình: mọi con số đều thuộc về benchmark đã tạo ra nó, được ghi lại theo ngày và ghi nguồn về tác giả.
Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.
Giới hạn: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.
Nguồn: LMArena (formerly LMSYS Chatbot Arena) ↗ · đã công bố 2026-08-11 · dữ liệu gốc ↗ · Dataset released under Creative Commons Attribution 4.0 (CC BY 4.0). Reuse is permitted with attribution - credit LMArena and link to the leaderboard.
| # | Model | Tổ chức | Arena score (human preference) |
|---|---|---|---|
| 1 | claude-fable-5 | Anthropic | 1506 95% CI 1501 to 1512 21,304 votes |
| 2 | claude-opus-4-6-thinking | Anthropic | 1505 95% CI 1501 to 1508 72,425 votes |
| 3 | claude-opus-4-7-thinking | Anthropic | 1502 95% CI 1498 to 1506 60,222 votes |
| 4 | muse-spark-1.2 (xHigh) | Meta | 1498 95% CI 1488 to 1509 3,278 votes |
| 5 | claude-opus-4-6 | Anthropic | 1498 95% CI 1494 to 1501 76,386 votes |
| 6 | claude-opus-5-high | Anthropic | 1494 95% CI 1489 to 1499 19,498 votes |
| 7 | claude-opus-4-7 | Anthropic | 1494 95% CI 1490 to 1498 61,308 votes |
| 8 | claude-opus-5-max | Anthropic | 1490 95% CI 1483 to 1497 9,419 votes |
| 9 | qwen3.8-max | Alibaba | 1490 95% CI 1482 to 1498 6,789 votes |
| 10 | muse-spark-1.1 | Meta | 1489 95% CI 1483 to 1494 16,648 votes |
| 11 | muse-spark | Meta | 1488 95% CI 1482 to 1494 13,600 votes |
| 12 | kimi-k3-max | Moonshot AI | 1487 95% CI 1481 to 1493 11,762 votes |
| 13 | gemini-3.1-pro-preview | 1486 95% CI 1483 to 1490 94,814 votes | |
| 14 | gemini-3-pro | 1486 95% CI 1482 to 1489 41,509 votes | |
| 15 | gemini-3.6-flash | 1484 95% CI 1478 to 1490 13,559 votes | |
| 16 | gpt-5.5-high | OpenAI | 1482 95% CI 1477 to 1486 55,210 votes |
| 17 | claude-opus-4-8-thinking | Anthropic | 1481 95% CI 1477 to 1486 40,409 votes |
| 18 | gpt-5.6-sol-xhigh | OpenAI | 1481 95% CI 1475 to 1487 15,304 votes |
| 19 | gemini-3.5-flash-high | 1477 95% CI 1473 to 1482 25,613 votes | |
| 20 | gpt-5.5 | OpenAI | 1477 95% CI 1473 to 1481 56,513 votes |
Các thanh được chia tỷ lệ theo phạm vi hiển thị để những khác biệt nhỏ vẫn dễ đọc. Chúng không bắt đầu từ không. Khi các khoảng tin cậy chồng lên nhau, các mô hình được coi là ngang nhau về mặt thống kê: hãy đọc nhóm đầu như một nhóm, không phải theo thứ tự nghiêm ngặt.
Cách đọc những kết quả này, và lý do chúng không đồng nhất
LMArena (formerly LMSYS Chatbot Arena)
Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.
Cẩn thận với: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.
Đã công khai 2026-08-11 · bảng xếp hạng trực tiếp ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).
LiveBench
A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.
Cẩn thận với: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.
Đã công khai · bảng xếp hạng trực tiếp ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.
Hai đơn vị khác nhau, không bao giờ dùng chung một biểu đồ
Điểm đánh giá theo sở thích của người dùng và điểm trả lời đúng không thể dùng chung một trục, vì vậy trang này không bao giờ gộp chúng vào một bảng. Một mô hình có thể dẫn đầu chỉ số này nhưng không phải chỉ số kia, đó là tín hiệu thực sự về điểm mạnh của nó, không phải mâu thuẫn: một câu hỏi là "câu trả lời nào người dùng thích hơn?", câu kia là "câu trả lời nào đúng hơn?".
Trang này sẽ không làm gì
Nó sẽ không cho bạn biết nên dùng AI nào. Mô hình dẫn đầu benchmark thường không phải là công cụ phù hợp nhất cho công việc của bạn, giá cả, tính sẵn có, khả năng tích hợp, độ dài ngữ cảnh và mức độ tuân theo hướng dẫn CỦA BẠN thường quan trọng hơn một hai điểm số. Tôi nên dùng AI nào? →