هويتك في الـ AI وفريقك من الـ AI، قابلان للنقل إلى كل منصة.
تسجيل الدخول ابدأ الآن
القائمة
بناء Agent of Me استكشاف الأساليب الوكلاء المهنيون وكلاء المجتمع لوحة التصنيفات أخبار AI
منصات AI الدليل Model Matrix مقارنة أي AI يجب أن أستخدم؟ أدلة التكامل إعداد OpenClaw مدى ملاءمة الـ Prompt
تعلّم وأدوات تعلّم استفسر عن البيانات منشئ الـ Agent API
حول من نحن تواصل معنا إخلاءات المسؤولية
تسجيل الدخول ابدأ الآن
الحساب
هويتك بالذكاء الاصطناعي، في كل مكان

أنشئ حساباً مجانياً لبناء ملفك الشخصي. خاص بالافتراضي. لا يُشارك أي شيء ما لم تنشره أنت.

ابدأ الآن تسجيل الدخول
الوضع الداكن

🧭 العرض الإرشادي
جديد على الـ prompts وتعليمات النظام ونوافذ السياق والـ tokens؟ نشرح كل مصطلح أثناء تصفحك، بلغة واضحة. الصفحات ذاتها، مع المساعدة مدمجةً فيها.

⚡ عرض الخبراء
أنت تعرف كيف يعمل الـ prompting. فقط الجوهر، واضحاً ومكثفاً، بلا شروحات زائدة. هذا هو العرض الافتراضي.

لغة الواجهة

لوحة تصدر الذكاء الاصطناعي

أي نموذج يتصدر، إجمالاً وبحسب الفئة. هذه هي المعايير العامة التي تُبنى منها الرسوم البيانية التي تراها على YouTube وX. لا ننشر درجاتنا الخاصة: كل رقم ينتمي إلى المعيار الذي أنتجه، مسجَّلاً بتاريخ ومنسوباً إلى مصدره.

Overall, ما يقيسه هذا: Human preference on open-ended chat.

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

القيود: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

المصدر: LMArena (formerly LMSYS Chatbot Arena) ↗ · منشور 2026-08-11 · البيانات الأساسية ↗ · Dataset released under Creative Commons Attribution 4.0 (CC BY 4.0). Reuse is permitted with attribution - credit LMArena and link to the leaderboard.

# النموذج المؤسسة Arena score (human preference)
1 claude-fable-5 Anthropic 1506 95% CI 1501 to 1512 21,304 votes
2 claude-opus-4-6-thinking Anthropic 1505 95% CI 1501 to 1508 72,425 votes
3 claude-opus-4-7-thinking Anthropic 1502 95% CI 1498 to 1506 60,222 votes
4 muse-spark-1.2 (xHigh) Meta 1498 95% CI 1488 to 1509 3,278 votes
5 claude-opus-4-6 Anthropic 1498 95% CI 1494 to 1501 76,386 votes
6 claude-opus-5-high Anthropic 1494 95% CI 1489 to 1499 19,498 votes
7 claude-opus-4-7 Anthropic 1494 95% CI 1490 to 1498 61,308 votes
8 claude-opus-5-max Anthropic 1490 95% CI 1483 to 1497 9,419 votes
9 qwen3.8-max Alibaba 1490 95% CI 1482 to 1498 6,789 votes
10 muse-spark-1.1 Meta 1489 95% CI 1483 to 1494 16,648 votes
11 muse-spark Meta 1488 95% CI 1482 to 1494 13,600 votes
12 kimi-k3-max Moonshot AI 1487 95% CI 1481 to 1493 11,762 votes
13 gemini-3.1-pro-preview Google 1486 95% CI 1483 to 1490 94,814 votes
14 gemini-3-pro Google 1486 95% CI 1482 to 1489 41,509 votes
15 gemini-3.6-flash Google 1484 95% CI 1478 to 1490 13,559 votes
16 gpt-5.5-high OpenAI 1482 95% CI 1477 to 1486 55,210 votes
17 claude-opus-4-8-thinking Anthropic 1481 95% CI 1477 to 1486 40,409 votes
18 gpt-5.6-sol-xhigh OpenAI 1481 95% CI 1475 to 1487 15,304 votes
19 gemini-3.5-flash-high Google 1477 95% CI 1473 to 1482 25,613 votes
20 gpt-5.5 OpenAI 1477 95% CI 1473 to 1481 56,513 votes

تُحجَّم الأعمدة وفق النطاق المرئي لتظل الفوارق الصغيرة مقروءة. لا تبدأ من الصفر. حيث تتداخل فترات الثقة، تكون النماذج متعادلة إحصائياً: اقرأ المجموعة العليا كمجموعة واحدة، لا كترتيب صارم.

كيفية قراءة هذه النتائج، ولماذا تتباين

LMArena (formerly LMSYS Chatbot Arena)

Real people are shown the same prompt answered by two anonymous models side by side and vote for the better answer. Millions of these blind head-to-head votes are fed into a Bradley-Terry statistical model (the successor to the Elo system it started with) which converts win/loss pairs into a single rating per model. A higher rating means people picked that model more often against strong opposition. This snapshot uses the 'style control' variant, which is the site's default: it statistically adjusts for answer length and formatting so a model cannot climb simply by writing longer, prettier replies.

انتبه لـ: It measures which answer people LIKE, not which answer is CORRECT - a confident, well-written wrong answer can still win a vote. Voters are self-selected volunteers rather than a representative sample, prompts skew toward what that crowd chooses to type, and models with few votes have wide confidence intervals (ci_low/ci_high) that often overlap the models ranked above and below them. Treat small rank gaps as ties.

منشور 2026-08-11 · لوحة نتائج مباشرة ↗ · Data: LMArena leaderboard dataset (CC BY 4.0).

LiveBench

A fixed set of test questions with objectively verifiable answers is run against each model and scored automatically against ground truth - no human voting and no AI judge, so the score is repeatable. This release spans 23 tasks grouped into 7 categories. Each category score is the average of its tasks, and the headline 'global average' is the average of the 7 category scores, so every category counts equally regardless of how many tasks it contains. Scores are percentages: 100 is perfect.

انتبه لـ: Contamination-LIMITED, not contamination-proof: questions are refreshed from recent sources to reduce the chance a model simply memorised them during training, but that cannot be guaranteed. Scores reflect only these 23 tasks - they say nothing about tone, safety, speed or cost. Many entries are effort/thinking variants of the same underlying model (model_id shows the exact configuration tested), and a variant given more reasoning budget will usually outscore the cheaper default that most people actually use.

منشور · لوحة نتائج مباشرة ↗ · Data: LiveBench 2026-06-25 release, livebench.ai.

وحدتا قياس مختلفتان، لا مخطط واحد أبداً

لا يمكن لتقييم التفضيل البشري ونسبة الإجابات الصحيحة أن يتشاركا محوراً واحداً، لذا لا تجمع هذه الصفحة بينهما في جدول واحد. قد يتصدر نموذج معين أحد المعيارين دون الآخر، وهذا مؤشر حقيقي على طبيعة أدائه لا تناقض فيه: أحدهما يسأل "أي إجابة فضّلها الناس؟"، والآخر يسأل "أي إجابة كانت صحيحة؟".

ما لن تفعله هذه الصفحة

لن يخبرك بأي ذكاء اصطناعي تستخدم. المتصدر في المعيار كثيراً ما لا يكون الأداة المثلى لعملك؛ السعر والتوفر والتكاملات وطول نافذة السياق ومدى اتباعه لتعليماتك أنت عادةً أهم من نقطة أو نقطتين في الدرجة. أي AI يجب أن أستخدم؟ →

أياً كان المتصدر هذا الشهر: ملفك الشخصي وفريقك من الـ AI قابلان للنقل، الانتقال إلى قائد جديد هو عملية لصق لا إعادة بناء. أدلة التكامل

قطاع الأعمال

Business AnalystChief of StaffExecutive AssistantM&A AnalystManagement ConsultantOperations AnalystProject ManagerRecruiter

المالية

AccountantDue Diligence AnalystEquity Research AnalystFamily Office AnalystFinancial AnalystFixed Income AnalystInvestment Banking AnalystPortfolio Analyst

قانوني

Contract Review AssistantLegal Due Diligence AssistantLegal Research AssistantParalegal

التسويق

Brand StrategistContent StrategistGEO AnalystMarketing StrategistSEO AnalystSales Strategist

شخصي

Career CoachLearning TutorReflection AssistantResearch AssistantTravel PlannerWriting Assistant

العقارات

Acquisition AnalystAsset Management AnalystCommercial Real Estate AnalystDevelopment AnalystLease AnalystProperty Financial Analyst

بحث

Competitive Intelligence AnalystDeep Research AnalystIndustry Research AnalystJournalist ResearcherMarket Research AnalystMedical Research Assistant

التكنولوجيا

AI Strategy AdvisorCybersecurity Research AssistantData AnalystProduct ManagerSoftware Engineer