あなたのAI identityとAIチームを、あらゆるAIプラットフォームにそのまま持ち込める。
サインイン はじめる
メニュー
Agent of Me を作成 スタイルを探索 プロフェッショナルエージェント コミュニティエージェント リーダーボード AIニュース
AIプラットフォーム ディレクトリ Model Matrix 比較 どのAIを使えばいいですか? インテグレーションガイド OpenClawをセットアップ Prompt Fit
学ぶ・ツール 学ぶ データに質問する エージェントビルダー API
概要 サイトについて お問い合わせ 免責事項
サインイン はじめる
アカウント
持ち運べるAI identity

無料アカウントを作成してプロフィールを構築しましょう。デフォルトは非公開。あなたが公開しない限り、何も共有されません。

はじめる サインイン
ダークモード

🧭 ガイドビュー
プロンプト、システム指示、コンテキストウィンドウ、トークンが初めてですか?このサイトを閲覧しながら、すべての用語をわかりやすい言葉で説明します。ヘルプが組み込まれた同じページで確認できます。

⚡ エキスパートビュー
プロンプトの仕組みはわかっている前提で。余計な説明なし、要点だけをコンパクトに。これがデフォルト表示。

表示言語

AI リーダーボード

総合および各カテゴリのトップモデル。YouTubeやXで使われているグラフの元データとなる公開ベンチマークです。独自スコアの公開は行いません:すべての数値は産出元のベンチマークに帰属し、取得日とともに出典を明示しています。

Coding, この指標の測定内容: Human preference on open-ended chat.

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

制限事項: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

出所: LMArena (formerly LMSYS Chatbot Arena) ↗ · 公開済み 2026-08-11 · 基礎データ ↗ · Dataset released under Creative Commons Attribution 4.0 (CC BY 4.0). Reuse is permitted with attribution - credit LMArena and link to the leaderboard.

# モデル 組織 Arena score (human preference)
1 claude-fable-5 Anthropic 1554 95% CI 1545 to 1562 5,751 votes
2 claude-opus-4-7-thinking Anthropic 1552 95% CI 1546 to 1558 17,231 votes
3 claude-opus-4-6-thinking Anthropic 1552 95% CI 1546 to 1557 18,855 votes
4 claude-opus-4-7 Anthropic 1548 95% CI 1542 to 1554 17,377 votes
5 claude-opus-4-6 Anthropic 1547 95% CI 1541 to 1552 21,309 votes
6 kimi-k3-max Moonshot AI 1542 95% CI 1531 to 1553 3,111 votes
7 muse-spark-1.2 (xHigh) Meta 1533 95% CI 1514 to 1553 947 votes
8 claude-opus-4-8-thinking Anthropic 1533 95% CI 1526 to 1540 11,325 votes
9 claude-opus-5-high Anthropic 1531 95% CI 1522 to 1540 5,281 votes
10 muse-spark-1.1 Meta 1531 95% CI 1522 to 1540 4,911 votes

グラフは表示範囲内でスケーリングされているため、小さな差異も読み取りやすくなっています。ゼロ起点ではありません。信頼区間が重複している場合、モデルは統計的に同率です。上位グループは厳密な順位ではなく、グループとして解釈してください。

この結果の読み方と、なぜ数値が異なるのか

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

注意すべきこと: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

公開済み 2026-08-11 · ライブリーダーボード ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

注意すべきこと: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

公開済み · ライブリーダーボード ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

2つの異なる単位。同一チャートには表示しない

人間の選好評価と正答率スコアは同じ軸で表せないため、このページでは両者を1つの表に混在させていません。片方でトップに立つモデルがもう片方でそうでない場合も、それは矛盾ではなく「どちらの回答が好まれたか」と「どちらの回答が正しかったか」という、本質的に異なる問いへの答えです。

このページでできないこと

どのAIを使うべきかは教えてくれません。ベンチマークのトップモデルが、あなたの用途に最適とは限らないからです。価格・可用性・インテグレーション・コンテキスト長、そしてあなたの指示にどれだけ従うかは、スコアの数点差より重要です。 どのAIを使えばいいですか? →

今月のリーダー: プロフィールとAIチームはポータブルです。新しいリーダーへの移行は、貼り付けるだけ、作り直し不要。 インテグレーションガイド

企業

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

ファイナンス

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

法律

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

マーケティング

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

個人

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

不動産

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

リサーチ

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

情報技術

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer