Ваша AI-идентичность и ваша AI-команда, которые работают на любой AI-платформе.
Войти Начать
Меню
Создать Agent of Me Обзор стилей Профессиональные агенты Агенты сообщества Рейтинг Новости об AI
AI-платформы Каталог Model Matrix Сравнить Какой AI мне использовать? Руководства по интеграции Настроить OpenClaw Prompt Fit
Обучение и инструменты Обучение Запрос к данным Конструктор агентов API
О сайте О нас Контакты Дисклеймеры
Войти Начать
Аккаунт
Ваша AI-идентичность, всегда с вами

Создайте бесплатный аккаунт и постройте свой профиль. По умолчанию всё приватно. Ничего не публикуется без вашего согласия.

Начать Войти
Тёмная тема

🧭 Режим с подсказками
Не знакомы с prompt, system instructions, context window, token? Мы объясняем каждый термин прямо при просмотре, простым языком. Те же страницы, со встроенными подсказками.

⚡ Экспертный режим
Вы уже знаете, как работает prompting. Только суть, чётко и компактно, без лишних пояснений. Это вид по умолчанию.

Язык интерфейса

AI-рейтинг

Какая модель лидирует, в целом и по категориям. Это публичные бенчмарки, на которых основаны графики, которые вы видите на YouTube и X. Мы не публикуем собственных оценок: каждая цифра принадлежит бенчмарку, который её создал, зафиксирована с датой и ссылкой на источник.

Coding, что это измеряет: Objective correctness on verifiable tasks.

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Ограничения: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Источник: LiveBench ↗ · опубликовано · исходные данные ↗ · Benchmark code and data are public on GitHub and Hugging Face under the project's own terms; cite LiveBench and link to livebench.ai when reusing scores.

# Модель Организация Score out of 100 (objective tasks)
1 Claude Fable 5 Max Effort Anthropic 86.0
2 GPT-5.6 Sol Max Effort OpenAI 83.9
3 Smaug-Agentic Abacus.AI 82.5
4 gpt-5.5-xhigh OpenAI 82.2
5 claude-opus-4-7-xhigh-effort Anthropic 82.1
6 Claude 4.8 Opus Thinking Max Effort Anthropic 81.8
7 claude-opus-5-max-effort Anthropic 81.5
8 Kimi K3 Moonshot AI 81.5
9 Claude Sonnet 5 xHigh Effort Anthropic 80.7
10 GPT-5.6 Terra Max Effort OpenAI 78.2
11 gemini-3.5-flash-high Google 78.2
12 Claude 4.6 Opus Thinking High Effort Anthropic 78.2
13 gpt-5.4-xhigh OpenAI 77.5
14 Muse Spark 1.2 xHigh Effort Meta 77.5
15 Muse Spark 1.1 xHigh Effort Meta 77.2
16 Gemini 3.1 Pro Preview High Google 76.5
17 gpt-5.2-2025-12-11-high OpenAI 76.1
18 DeepSeek V4 Flash 0731 DeepSeek 75.0
19 Qwen 3.8 Max Alibaba 72.9
20 Grok 4.5 xAI 68.6

Столбцы масштабированы по видимому диапазону, чтобы небольшие различия оставались читаемыми. Отсчёт не идёт от нуля. Там, где доверительные интервалы пересекаются, модели статистически равны, воспринимайте верхнюю группу как группу, а не строгий рейтинг. Некоторые записи, это конфигурации с максимальными усилиями, которые дают более высокий балл, чем настройки по умолчанию, используемые большинством. Рядом с результатом указано, какая модель тестировалась.

Как читать эти данные и почему они расходятся

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

Обратите внимание на: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

Опубликовано 2026-08-11 · live leaderboard ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Обратите внимание на: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Опубликовано · live leaderboard ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

Две разные единицы измерения, никогда не на одном графике

Рейтинг предпочтений людей и показатель правильных ответов не могут находиться на одной оси, поэтому на этой странице они никогда не смешиваются в одной таблице. Модель может лидировать по одному показателю и уступать по другому, это реальный сигнал о её сильных сторонах, а не противоречие: первый спрашивает «какой ответ люди предпочли?», второй, «какой ответ был верным?».

Чего эта страница не делает

Он не скажет вам, какой AI использовать. Лидер бенчмарков часто не лучший инструмент для вашей задачи, цена, доступность, интеграции, длина контекста и то, насколько точно он следует ВАШИМ инструкциям, обычно важнее одного-двух пунктов в рейтинге. Какой AI мне использовать? →

Кто лидирует в этом месяце: ваш профиль и ваша AI-команда портативны, переход к новому руководителю, это вставка, а не пересборка. Руководства по интеграции

Деловая активность

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

Финансы

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

Юридическое

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

Маркетинг

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

Личное

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

Недвижимость

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

Исследование

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

Технологии

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer