AI leaderboard
कौन सा model आगे है, overall और category के अनुसार। ये वही public benchmarks हैं जिनसे YouTube और X पर दिखने वाले charts बनते हैं। हम अपने कोई scores publish नहीं करते: हर number उस benchmark का है जिसने इसे produce किया, एक date पर capture किया और source को credit दिया।
Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.
सीमाएँ: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.
स्रोत: LMArena (formerly LMSYS Chatbot Arena) ↗ · प्रकाशित 2026-08-11 · underlying data ↗ · Dataset released under Creative Commons Attribution 4.0 (CC BY 4.0). Reuse is permitted with attribution - credit LMArena and link to the leaderboard.
| # | Model | Organization | Arena score (human preference) |
|---|---|---|---|
| 1 | muse-spark-1.2 (xHigh) | Meta | 1519 95% CI 1493 to 1545 516 votes |
| 2 | claude-opus-4-6-thinking | Anthropic | 1517 95% CI 1511 to 1524 12,338 votes |
| 3 | claude-opus-4-7-thinking | Anthropic | 1516 95% CI 1509 to 1524 10,557 votes |
| 4 | claude-opus-4-7 | Anthropic | 1515 95% CI 1508 to 1522 10,869 votes |
| 5 | claude-fable-5 | Anthropic | 1515 95% CI 1504 to 1526 3,595 votes |
| 6 | claude-opus-4-6 | Anthropic | 1512 95% CI 1506 to 1518 13,542 votes |
| 7 | qwen3.8-max | Alibaba | 1505 95% CI 1486 to 1524 1,049 votes |
| 8 | kimi-k3-max | Moonshot AI | 1503 95% CI 1489 to 1517 1,860 votes |
| 9 | claude-opus-4-8-thinking | Anthropic | 1501 95% CI 1493 to 1509 7,249 votes |
| 10 | claude-opus-4-8 | Anthropic | 1499 95% CI 1491 to 1508 7,463 votes |
Bars को visible range में scale किया गया है ताकि छोटे अंतर भी पढ़े जा सकें। ये zero से शुरू नहीं होते। जहाँ confidence intervals overlap करते हैं, वहाँ models statistically tied हैं: top group को एक group के रूप में पढ़ें, strict order के रूप में नहीं।
इन्हें कैसे पढ़ें, और ये क्यों disagree करते हैं
LMArena (formerly LMSYS Chatbot Arena)
Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.
इनसे सावधान रहें: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.
Published 2026-08-11 · live leaderboard ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).
LiveBench
A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.
इनसे सावधान रहें: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.
Published · live leaderboard ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.
दो अलग units, कभी एक chart नहीं
एक human-preference रेटिंग और एक percent-correct स्कोर एक ही axis साझा नहीं कर सकते, इसलिए यह पेज उन्हें कभी एक टेबल में नहीं मिलाता। कोई मॉडल एक में शीर्ष पर हो सकता है और दूसरे में नहीं, और यह एक वास्तविक संकेत है कि वह किसमें अच्छा है, कोई विरोधाभास नहीं: एक पूछता है "लोगों ने कौन-सा जवाब पसंद किया?", दूसरा पूछता है "कौन-सा जवाब सही था?"।
यह page क्या नहीं करेगा
यह आपको नहीं बताएगा कि कौन सा AI use करें। Benchmark leader अक्सर आपके काम के लिए सही tool नहीं होता, price, availability, integrations, context length और यह कि वह YOUR instructions कितनी अच्छी तरह follow करता है, आमतौर पर एक-दो score points से ज़्यादा matter करते हैं। मुझे कौन-सा AI उपयोग करना चाहिए? →