나의 AI 정체성과 AI 팀을 모든 AI 플랫폼에서 그대로 사용하세요.
로그인 시작하기
메뉴
Agent of Me 만들기 스타일 탐색 전문 에이전트 커뮤니티 에이전트 리더보드 AI 뉴스
AI 플랫폼 디렉터리 Model Matrix 비교 어떤 AI를 사용해야 하나요? 연동 가이드 OpenClaw 설정 Prompt Fit
학습 및 도구 학습 데이터에 질문하기 에이전트 빌더 API
소개 회사 소개 문의 면책 조항
로그인 시작하기
계정
이식 가능한 AI 정체성

무료 계정을 만들어 프로필을 구축하세요. 기본적으로 비공개입니다. 직접 공개하지 않는 한 아무것도 공유되지 않습니다.

시작하기 로그인
다크 모드

🧭 가이드 보기
프롬프트, 시스템 지침, 컨텍스트 윈도우, 토큰이 처음이신가요? 모든 용어를 쉬운 설명으로 탐색하면서 익힐 수 있습니다. 동일한 페이지에 도움말이 내장되어 있습니다.

⚡ 전문가 보기
프롬프트 작동 방식은 이미 알고 계시죠. 군더더기 없이, 핵심만 간결하게. 기본 보기입니다.

인터페이스 언어

AI 리더보드

전체 및 카테고리별 선두 모델을 확인하세요. YouTube와 X에서 볼 수 있는 차트의 기반이 되는 공개 벤치마크입니다. 자체 점수는 공개하지 않으며, 모든 수치는 해당 벤치마크에서 산출된 것으로 날짜와 출처를 명시합니다.

Overall, 측정 항목: Human preference on open-ended chat.

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

한계: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

출처: LMArena (formerly LMSYS Chatbot Arena) ↗ · 공개됨 2026-08-11 · 기반 데이터 ↗ · Dataset released under Creative Commons Attribution 4.0 (CC BY 4.0). Reuse is permitted with attribution - credit LMArena and link to the leaderboard.

# 모델 조직 Arena score (human preference)
1 claude-fable-5 Anthropic 1506 95% CI 1501 to 1512 21,304 votes
2 claude-opus-4-6-thinking Anthropic 1505 95% CI 1501 to 1508 72,425 votes
3 claude-opus-4-7-thinking Anthropic 1502 95% CI 1498 to 1506 60,222 votes
4 muse-spark-1.2 (xHigh) Meta 1498 95% CI 1488 to 1509 3,278 votes
5 claude-opus-4-6 Anthropic 1498 95% CI 1494 to 1501 76,386 votes
6 claude-opus-5-high Anthropic 1494 95% CI 1489 to 1499 19,498 votes
7 claude-opus-4-7 Anthropic 1494 95% CI 1490 to 1498 61,308 votes
8 claude-opus-5-max Anthropic 1490 95% CI 1483 to 1497 9,419 votes
9 qwen3.8-max Alibaba 1490 95% CI 1482 to 1498 6,789 votes
10 muse-spark-1.1 Meta 1489 95% CI 1483 to 1494 16,648 votes
11 muse-spark Meta 1488 95% CI 1482 to 1494 13,600 votes
12 kimi-k3-max Moonshot AI 1487 95% CI 1481 to 1493 11,762 votes
13 gemini-3.1-pro-preview Google 1486 95% CI 1483 to 1490 94,814 votes
14 gemini-3-pro Google 1486 95% CI 1482 to 1489 41,509 votes
15 gemini-3.6-flash Google 1484 95% CI 1478 to 1490 13,559 votes
16 gpt-5.5-high OpenAI 1482 95% CI 1477 to 1486 55,210 votes
17 claude-opus-4-8-thinking Anthropic 1481 95% CI 1477 to 1486 40,409 votes
18 gpt-5.6-sol-xhigh OpenAI 1481 95% CI 1475 to 1487 15,304 votes
19 gemini-3.5-flash-high Google 1477 95% CI 1473 to 1482 25,613 votes
20 gpt-5.5 OpenAI 1477 95% CI 1473 to 1481 56,513 votes

막대 그래프는 가시 범위 내에서 작은 차이도 읽기 쉽도록 조정되어 있으며, 0부터 시작하지 않습니다. 신뢰 구간이 겹치는 경우 모델 간 통계적 차이가 없습니다. 상위 그룹은 엄격한 순위가 아닌 하나의 그룹으로 읽어주세요.

이 결과를 읽는 방법, 그리고 왜 서로 다른지

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

주의하세요: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

게시됨 2026-08-11 · 실시간 리더보드 ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

주의하세요: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

게시됨 · 실시간 리더보드 ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

두 가지 단위는 서로 다릅니다. 하나의 차트로 표시하지 않습니다

인간 선호도 평가와 정답률은 같은 축을 공유할 수 없으므로, 이 페이지에서는 두 지표를 하나의 표에 혼합하지 않습니다. 한 모델이 한쪽에서 1위를 차지하고 다른 쪽에서는 그렇지 않을 수 있는데, 이는 모순이 아니라 해당 모델이 무엇에 강한지를 보여주는 실질적인 신호입니다. 하나는 "사람들이 어떤 답변을 선호했는가?"를 묻고, 다른 하나는 "어떤 답변이 옳았는가?"를 묻습니다.

이 페이지에서 하지 않는 것

어떤 AI를 사용해야 하는지 알려드리지 않습니다. 벤치마크 1위가 항상 최선의 도구는 아닙니다. 가격, 가용성, 연동 기능, 컨텍스트 길이, 그리고 내 지침을 얼마나 잘 따르는지가 보통 점수 차이 몇 점보다 훨씬 중요합니다. 어떤 AI를 사용해야 하나요? →

이번 달 선두: 프로필과 AI 팀은 그대로 이식됩니다. 새로운 도구로 전환해도 붙여넣기 한 번이면 충분합니다. 처음부터 다시 만들 필요가 없습니다. 연동 가이드

기업

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

금융

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

법률

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

마케팅

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

개인

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

부동산

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

리서치

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

기술

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer