आपकी AI identity और आपकी AI टीम, हर AI प्लेटफ़ॉर्म पर portable।
साइन इन करें शुरू करें
मेनू
एक Agent of Me बनाएं Styles explore करें Professional Agents Community Agents Leaderboard AI समाचार
AI Platforms Directory Model Matrix तुलना करें मुझे कौन-सा AI उपयोग करना चाहिए? Integration Guides OpenClaw set up करें Prompt Fit
जानें और उपकरण जानें डेटा से पूछें Agent Builder API
परिचय हमारे बारे में संपर्क Disclaimers
साइन इन करें शुरू करें
अकाउंट
आपकी AI identity, portable

अपना profile बनाने के लिए free account बनाएं। Default रूप से private। जब तक आप publish न करें, कुछ भी share नहीं होता।

शुरू करें साइन इन करें
डार्क मोड

🧭 निर्देशित दृश्य
Prompts, system instructions, context windows, tokens से नए हैं? हम हर term को plain English में explain करते हैं जैसे आप browse करते हैं। वही pages, help built-in के साथ।

⚡ विशेषज्ञ दृश्य
आप prompting पहले से जानते हैं। बस सार, साफ़ और कॉम्पैक्ट, बिना अतिरिक्त स्पष्टीकरण के। यह डिफ़ॉल्ट व्यू है।

इंटरफ़ेस भाषा

AI leaderboard

कौन सा model आगे है, overall और category के अनुसार। ये वही public benchmarks हैं जिनसे YouTube और X पर दिखने वाले charts बनते हैं। हम अपने कोई scores publish नहीं करते: हर number उस benchmark का है जिसने इसे produce किया, एक date पर capture किया और source को credit दिया।

Reasoning, यह क्या measure करता है: Objective correctness on verifiable tasks.

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

सीमाएँ: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

स्रोत: LiveBench ↗ · प्रकाशित · underlying data ↗ · Benchmark code and data are public on GitHub and Hugging Face under the project's own terms; cite LiveBench and link to livebench.ai when reusing scores.

# Model Organization Score out of 100 (objective tasks)
1 GPT-5.6 Sol Max Effort OpenAI 91.7
2 claude-opus-5-max-effort Anthropic 91.2
3 Kimi K3 Moonshot AI 90.7
4 GPT-5.6 Terra Max Effort OpenAI 90.6
5 Smaug-Agentic Abacus.AI 90.3
6 Muse Spark 1.2 xHigh Effort Meta 90.0
7 Claude Fable 5 Max Effort Anthropic 89.7
8 gpt-5.5-xhigh OpenAI 89.7
9 Claude 4.8 Opus Thinking Max Effort Anthropic 89.2
10 Claude Sonnet 5 xHigh Effort Anthropic 88.7
11 Claude 4.6 Opus Thinking High Effort Anthropic 88.7
12 Qwen 3.8 Max Alibaba 88.2
13 gpt-5.4-xhigh OpenAI 88.1
14 Muse Spark 1.1 xHigh Effort Meta 87.7
15 claude-opus-4-7-xhigh-effort Anthropic 87.2
16 Grok 4.5 xAI 87.2
17 DeepSeek V4 Flash 0731 DeepSeek 86.6
18 Gemini 3.1 Pro Preview High Google 84.0
19 gpt-5.2-2025-12-11-high OpenAI 83.2
20 gemini-3.5-flash-high Google 82.0

Bars को visible range में scale किया गया है ताकि छोटे अंतर भी पढ़े जा सकें। ये zero से शुरू नहीं होते। जहाँ confidence intervals overlap करते हैं, वहाँ models statistically tied हैं: top group को एक group के रूप में पढ़ें, strict order के रूप में नहीं। कुछ entries maximum-effort configurations हैं, जो default settings से ज़्यादा score करती हैं जो अधिकतर लोग actually use करते हैं। model name बताता है कि किसे test किया गया।

इन्हें कैसे पढ़ें, और ये क्यों disagree करते हैं

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

इनसे सावधान रहें: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

Published 2026-08-11 · live leaderboard ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

इनसे सावधान रहें: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Published · live leaderboard ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

दो अलग units, कभी एक chart नहीं

एक human-preference रेटिंग और एक percent-correct स्कोर एक ही axis साझा नहीं कर सकते, इसलिए यह पेज उन्हें कभी एक टेबल में नहीं मिलाता। कोई मॉडल एक में शीर्ष पर हो सकता है और दूसरे में नहीं, और यह एक वास्तविक संकेत है कि वह किसमें अच्छा है, कोई विरोधाभास नहीं: एक पूछता है "लोगों ने कौन-सा जवाब पसंद किया?", दूसरा पूछता है "कौन-सा जवाब सही था?"।

यह page क्या नहीं करेगा

यह आपको नहीं बताएगा कि कौन सा AI use करें। Benchmark leader अक्सर आपके काम के लिए सही tool नहीं होता, price, availability, integrations, context length और यह कि वह YOUR instructions कितनी अच्छी तरह follow करता है, आमतौर पर एक-दो score points से ज़्यादा matter करते हैं। मुझे कौन-सा AI उपयोग करना चाहिए? →

इस महीने जो आगे है: आपकी profile और आपकी AI टीम portable हैं, नए leader पर जाना एक paste है, rebuild नहीं। Integration guides

व्यवसाय

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

फ़ाइनेंस

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

Legal

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

Marketing

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

Personal

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

रियल एस्टेट

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

Research

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

प्रौद्योगिकी

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer