AI identity dan AI team-mu, portabel ke setiap platform AI.
Masuk Mulai
Menu
Buat Agent of Me Jelajahi Gaya Professional Agents Community Agents Leaderboard Berita AI
Platform AI Direktori Model Matrix Bandingkan AI mana yang sebaiknya aku gunakan? Panduan Integrasi Siapkan OpenClaw Prompt Fit
Pelajari & Alat Pelajari Tanya Data Agent Builder API
Tentang Tentang kami Kontak Disclaimers
Masuk Mulai
Akun
Identitas AI Anda, portabel

Buat akun gratis untuk membangun profil Anda. Privat secara default. Tidak ada yang dibagikan kecuali Anda mempublikasikannya.

Mulai Masuk
Mode Gelap

🧭 Tampilan Terpandu
Belum familiar dengan prompt, system instruction, context window, token? Kami menjelaskan setiap istilah saat kamu menjelajah, dalam bahasa yang mudah dipahami. Halaman yang sama, dengan panduan terintegrasi.

⚡ Tampilan Ahli
Kamu sudah paham cara prompting bekerja. Cukup isinya, ringkas dan to the point, tanpa penjelasan tambahan. Ini adalah tampilan default.

Bahasa antarmuka

AI leaderboard

Model mana yang unggul, secara keseluruhan maupun per kategori. Ini adalah benchmark publik yang menjadi dasar grafik yang kamu lihat di YouTube dan X. Kami tidak menerbitkan skor sendiri: setiap angka milik benchmark yang menghasilkannya, diambil pada tanggal tertentu dan dikreditkan ke sumbernya.

Agentic Coding, apa yang ini ukur: Objective correctness on verifiable tasks.

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Keterbatasan: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Sumber: LiveBench ↗ · diterbitkan · data yang mendasari ↗ · Benchmark code and data are public on GitHub and Hugging Face under the project's own terms; cite LiveBench and link to livebench.ai when reusing scores.

# Model Organisasi Score out of 100 (objective tasks)
1 claude-opus-5-max-effort Anthropic 65.2
2 Smaug-Agentic Abacus.AI 64.7
3 Qwen 3.8 Max Alibaba 64.7
4 Claude Fable 5 Max Effort Anthropic 62.2
5 Kimi K3 Moonshot AI 62.2
6 Claude Sonnet 5 xHigh Effort Anthropic 59.4
7 Muse Spark 1.1 xHigh Effort Meta 58.5
8 Muse Spark 1.2 xHigh Effort Meta 57.6
9 Grok 4.5 xAI 56.5
10 GPT-5.6 Sol Max Effort OpenAI 56.2
11 GPT-5.6 Terra Max Effort OpenAI 55.0
12 gpt-5.5-xhigh OpenAI 54.0
13 gpt-5.4-xhigh OpenAI 53.8
14 claude-opus-4-7-xhigh-effort Anthropic 50.7
15 Claude 4.8 Opus Thinking Max Effort Anthropic 50.5
16 gpt-5.2-2025-12-11-high OpenAI 50.2
17 gemini-3.5-flash-high Google 49.0
18 Claude 4.6 Opus Thinking High Effort Anthropic 49.0
19 DeepSeek V4 Flash 0731 DeepSeek 46.8
20 Gemini 3.1 Pro Preview High Google 44.1

Batang diskalakan pada rentang yang terlihat agar perbedaan kecil tetap mudah dibaca. Batang tidak dimulai dari nol. Di mana confidence interval saling tumpang tindih, model-model tersebut seri secara statistik: baca kelompok teratas sebagai satu kelompok, bukan urutan ketat. Beberapa entri adalah konfigurasi dengan upaya maksimal, yang mendapat skor lebih tinggi dari pengaturan default yang umumnya digunakan orang. Nama model menunjukkan mana yang diuji.

Cara membaca ini, dan mengapa keduanya berbeda

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

Waspadai: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

Diterbitkan 2026-08-11 · live leaderboard ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Waspadai: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Diterbitkan · live leaderboard ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

Dua unit yang berbeda, jangan dalam satu grafik

Penilaian preferensi manusia dan skor persentase-benar tidak dapat berbagi satu sumbu, sehingga halaman ini tidak pernah mencampurnya dalam satu tabel. Sebuah model bisa unggul di satu sisi namun tidak di sisi lain, dan itu adalah sinyal nyata tentang keunggulannya, bukan kontradiksi: satu bertanya "jawaban mana yang lebih disukai orang?", yang lain bertanya "jawaban mana yang benar?".

Yang tidak akan dilakukan halaman ini

Ini tidak akan memberi tahu AI mana yang harus digunakan. Pemimpin benchmark sering kali bukan alat yang tepat untuk pekerjaanmu, harga, ketersediaan, integrasi, panjang konteks, dan seberapa baik ia mengikuti instruksi KAMU biasanya lebih penting daripada selisih satu dua poin skor. AI mana yang sebaiknya aku gunakan? →

Siapa yang memimpin bulan ini: profilmu dan AI team-mu portabel, berpindah ke pemimpin baru hanya perlu paste, bukan membangun ulang dari awal. Panduan integrasi

Bisnis

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

Keuangan

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

Hukum

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

Pemasaran

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

Personal

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

Properti

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

Riset

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

Teknologi

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer