Deine AI-Identität und dein AI-Team, portabel für jede AI-Plattform.
Anmelden Loslegen
Menü
Einen Agent of Me erstellen Stile erkunden Professionelle Agenten Community-Agents Leaderboard KI-News
KI-Plattformen Verzeichnis Model Matrix Vergleichen Welche KI sollte ich verwenden? Integrationsleitfäden OpenClaw einrichten Prompt Fit
Lernen & Tools Lernen Daten befragen Agent Builder API
Über Über uns Kontakt Haftungsausschlüsse
Anmelden Loslegen
Konto
Deine KI-Identität, portabel

Erstelle ein kostenloses Konto, um dein Profil aufzubauen. Standardmäßig privat. Nichts wird geteilt, es sei denn, du veröffentlichst es.

Loslegen Anmelden
Dunkelmodus

🧭 Geführte Ansicht
Neu bei Prompts, System Instructions, Context Windows, Tokens? Wir erklären jeden Begriff beim Stöbern verständlich. Dieselben Seiten, mit integrierter Hilfe.

⚡ Expertenansicht
Du weißt, wie Prompting funktioniert. Nur das Wesentliche, klar und kompakt, ohne zusätzliche Erklärungen. Das ist die Standardansicht.

Oberflächensprache

AI-Leaderboard

Welches Modell führt, insgesamt und nach Kategorie. Das sind die öffentlichen Benchmarks, aus denen die Diagramme auf YouTube und X gebaut werden. Wir veröffentlichen keine eigenen Scores: Jede Zahl gehört dem Benchmark, der sie erstellt hat, erfasst an einem Datum und mit Quellenangabe.

Data Analysis, was das misst: Objective correctness on verifiable tasks.

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Einschränkungen: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Quelle: LiveBench ↗ · veröffentlicht · zugrunde liegende Daten ↗ · Benchmark code and data are public on GitHub and Hugging Face under the project's own terms; cite LiveBench and link to livebench.ai when reusing scores.

# Modell Organisation Score out of 100 (objective tasks)
1 gpt-5.5-xhigh OpenAI 81.6
2 Claude Fable 5 Max Effort Anthropic 80.5
3 Smaug-Agentic Abacus.AI 79.9
4 GPT-5.6 Sol Max Effort OpenAI 79.8
5 DeepSeek V4 Flash 0731 DeepSeek 79.3
6 gpt-5.4-xhigh OpenAI 79.3
7 GPT-5.6 Terra Max Effort OpenAI 79.3
8 Kimi K3 Moonshot AI 78.7
9 Gemini 3.1 Pro Preview High Google 78.5
10 Qwen 3.8 Max Alibaba 78.4
11 claude-opus-4-7-xhigh-effort Anthropic 78.3
12 gpt-5.2-2025-12-11-high OpenAI 78.2
13 Muse Spark 1.2 xHigh Effort Meta 76.5
14 claude-opus-5-max-effort Anthropic 74.5
15 Grok 4.5 xAI 73.0
16 Muse Spark 1.1 xHigh Effort Meta 72.5
17 Claude Sonnet 5 xHigh Effort Anthropic 71.7
18 Claude 4.6 Opus Thinking High Effort Anthropic 69.9
19 Claude 4.8 Opus Thinking Max Effort Anthropic 66.0
20 gemini-3.5-flash-high Google 64.9

Balken sind auf den sichtbaren Bereich skaliert, damit kleine Unterschiede lesbar bleiben. Sie beginnen nicht bei null. Wenn sich Konfidenzintervalle überschneiden, sind die Modelle statistisch gleichauf, lies die Spitzengruppe als Gruppe, nicht als strenge Reihenfolge. Manche Einträge sind Konfigurationen mit maximalem Aufwand, die höher abschneiden als die Standardeinstellungen, die die meisten tatsächlich verwenden. Der Modellname zeigt, welches getestet wurde.

Wie man diese liest, und warum sie voneinander abweichen

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

Achtung vor: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

Veröffentlicht 2026-08-11 · Live-Leaderboard ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Achtung vor: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Veröffentlicht · Live-Leaderboard ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

Zwei verschiedene Einheiten, niemals ein gemeinsames Diagramm

Eine Bewertung nach menschlicher Präferenz und ein prozentualer Korrektheitswert lassen sich nicht auf einer Achse darstellen, daher werden sie auf dieser Seite nie in einer Tabelle vermischt. Ein Modell kann bei einem der beiden führen, beim anderen nicht, das ist ein aussagekräftiges Signal dafür, worin es stark ist, kein Widerspruch: Das eine fragt „Welche Antwort haben Menschen bevorzugt?", das andere „Welche Antwort war richtig?".

Was diese Seite nicht tut

Es wird dir nicht sagen, welche KI du nutzen sollst. Der Benchmark-Spitzenreiter ist oft nicht das richtige Werkzeug für deine Arbeit. Preis, Verfügbarkeit, Integrationen, Kontextlänge und wie gut es DEINE Anweisungen befolgt, sind meist wichtiger als ein paar Scorepunkte. Welche KI sollte ich verwenden? →

Wer diesen Monat führt: Dein Profil und dein AI-Team sind portabel, zu einem neuen Tool zu wechseln ist ein Einfügen, kein Neuaufbau. Integrationsleitfäden

Konjunktur

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

Finanzen

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

Rechtliches

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

Marketing

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

Persönlich

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

Immobilien

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

Recherche

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

Technologie

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer