AI identity và AI team của bạn, portable trên mọi nền tảng AI.
Đăng nhập Bắt đầu
Menu
Xây dựng Agent of Me Khám phá Phong cách Professional Agents Agent cộng đồng Bảng xếp hạng Tin tức AI
Nền tảng AI Danh mục Model Matrix So sánh Tôi nên dùng AI nào? Hướng dẫn tích hợp Thiết lập OpenClaw Prompt Fit
Học & Công cụ Học Hỏi dữ liệu Agent Builder API
Giới thiệu Về chúng tôi Liên hệ Tuyên bố miễn trách
Đăng nhập Bắt đầu
Tài khoản
AI identity của bạn, mang đi được

Tạo tài khoản miễn phí để xây dựng profile của bạn. Mặc định riêng tư. Không có gì được chia sẻ trừ khi bạn tự công bố.

Bắt đầu Đăng nhập
Chế độ tối

🧭 Chế độ hướng dẫn
Chưa quen với prompt, system instruction, context window, token? Chúng tôi giải thích từng thuật ngữ ngay khi bạn duyệt, bằng ngôn ngữ dễ hiểu. Cùng một trang, tích hợp sẵn phần hướng dẫn.

⚡ Chế độ chuyên gia
Bạn đã biết cách prompting hoạt động. Chỉ phần nội dung thực chất, gọn gàng và súc tích, không có giải thích thêm. Đây là chế độ xem mặc định.

Ngôn ngữ giao diện

Bảng xếp hạng AI

Model nào dẫn đầu, tổng thể và theo từng danh mục. Đây là những benchmark công khai mà các biểu đồ bạn thấy trên YouTube và X được xây dựng từ đó. Chúng tôi không công bố điểm số của riêng mình: mọi con số đều thuộc về benchmark đã tạo ra nó, được ghi lại theo ngày và ghi nguồn về tác giả.

Language, chỉ số này đo gì: Objective correctness on verifiable tasks.

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Giới hạn: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Nguồn: LiveBench ↗ · đã công bố · dữ liệu gốc ↗ · Benchmark code and data are public on GitHub and Hugging Face under the project's own terms; cite LiveBench and link to livebench.ai when reusing scores.

# Model Tổ chức Score out of 100 (objective tasks)
1 Claude Fable 5 Max Effort Anthropic 90.7
2 claude-opus-5-max-effort Anthropic 88.7
3 GPT-5.6 Sol Max Effort OpenAI 87.7
4 gpt-5.5-xhigh OpenAI 87.4
5 Kimi K3 Moonshot AI 85.5
6 Gemini 3.1 Pro Preview High Google 85.4
7 gemini-3.5-flash-high Google 84.6
8 Smaug-Agentic Abacus.AI 84.4
9 Claude 4.6 Opus Thinking High Effort Anthropic 83.3
10 GPT-5.6 Terra Max Effort OpenAI 82.9
11 Grok 4.5 xAI 82.8
12 gpt-5.4-xhigh OpenAI 82.6
13 gpt-5.2-2025-12-11-high OpenAI 79.8
14 Qwen 3.8 Max Alibaba 79.7
15 Claude 4.8 Opus Thinking Max Effort Anthropic 79.7
16 DeepSeek V4 Flash 0731 DeepSeek 79.2
17 Muse Spark 1.2 xHigh Effort Meta 78.6
18 claude-opus-4-7-xhigh-effort Anthropic 77.9
19 Claude Sonnet 5 xHigh Effort Anthropic 75.0
20 Muse Spark 1.1 xHigh Effort Meta 74.3

Các thanh được chia tỷ lệ theo phạm vi hiển thị để những khác biệt nhỏ vẫn dễ đọc. Chúng không bắt đầu từ không. Khi các khoảng tin cậy chồng lên nhau, các mô hình được coi là ngang nhau về mặt thống kê: hãy đọc nhóm đầu như một nhóm, không phải theo thứ tự nghiêm ngặt. Một số mục là cấu hình nỗ lực tối đa, cho điểm số cao hơn cài đặt mặc định mà hầu hết mọi người thực sự dùng. Tên mô hình hiển thị mô hình nào đã được kiểm thử.

Cách đọc những kết quả này, và lý do chúng không đồng nhất

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

Cẩn thận với: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

Đã công khai 2026-08-11 · bảng xếp hạng trực tiếp ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

Cẩn thận với: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

Đã công khai · bảng xếp hạng trực tiếp ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

Hai đơn vị khác nhau, không bao giờ dùng chung một biểu đồ

Điểm đánh giá theo sở thích của người dùng và điểm trả lời đúng không thể dùng chung một trục, vì vậy trang này không bao giờ gộp chúng vào một bảng. Một mô hình có thể dẫn đầu chỉ số này nhưng không phải chỉ số kia, đó là tín hiệu thực sự về điểm mạnh của nó, không phải mâu thuẫn: một câu hỏi là "câu trả lời nào người dùng thích hơn?", câu kia là "câu trả lời nào đúng hơn?".

Trang này sẽ không làm gì

Nó sẽ không cho bạn biết nên dùng AI nào. Mô hình dẫn đầu benchmark thường không phải là công cụ phù hợp nhất cho công việc của bạn, giá cả, tính sẵn có, khả năng tích hợp, độ dài ngữ cảnh và mức độ tuân theo hướng dẫn CỦA BẠN thường quan trọng hơn một hai điểm số. Tôi nên dùng AI nào? →

Ai dẫn đầu tháng này: profile và AI team của bạn hoàn toàn di động, chuyển sang một leader mới chỉ cần dán vào, không cần xây lại từ đầu. Hướng dẫn tích hợp

Doanh nghiệp

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

Tài chính

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

Pháp lý

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

Marketing

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

Cá nhân

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

Bất động sản

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

Nghiên cứu

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

Công nghệ

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer