INTRODUCTION

Narrative

This dashboard shows that it is practical to collect live benchmark data and provides visual intuition into the value of this data for model selection. By the end of the story, we conclude that AI alignment through benchmark comparisons is operationalizable and requires a nuanced understanding of the details to be enacted safely. The dashboard then points to a second dashboard that explains some nuances and provides a mockup of a potential community practice.

The live data sources were gathered from the sources acknowledged below. Through their efforts, this dashboard was made viable.

Acknowledgements

The live datasets are gathered from the following sources:

  1. Artificial.AI: Artificial Analysis performs intelligence, quality, performance and price benchmarking on AI models, inference API endpoints and systems. This section of our website describes our benchmarking methodology, including both our quality benchmarking and performance benchmarking.
  2. /oolong-tea-2026: Auto-updated daily snapshots of every Arena AI (formerly LMSYS Chatbot Arena) leaderboard in structured JSON.
  3. Epoch: Investigating the trajectory of AI for the benefit of society.

1 a: Coverage


Artificial Analysis (https://artificialanalysis.ai) provides its gathered benchmarks for free through a REST API. Their coverage spans a range of benchmarks (around 10) based on their own AGI benchmark. Artificial Analysis also provides a detailed dashboard that you may consider interacting with in your own research.

The graph is dynamic. Hovering over a point will display a bar with further information.

1 b: Best models

Model Domain Benchmark Details
Qwen3.8 Max Agentic tau_banking Link t2-Bench: benchmark for Tool-Agent-User interaction in real-world domains.
GLM-5.2 (max) Agentic tau2 Link t2-Bench: benchmark for Tool-Agent-User interaction in real-world domains.
JT-35B-Flash Agentic tau2 Link t2-Bench: benchmark for Tool-Agent-User interaction in real-world domains.
GPT-5.6 Sol (max) Agentic terminalbench_hard Link Terminal-Bench: agentic benchmark evaluating agents on software engineering, sysadmin, and game-playing via a terminal.
GPT-5.6 Sol (xhigh) Agentic terminalbench_v2_1 Link Terminal-Bench: agentic benchmark evaluating agents on software engineering, sysadmin, and game-playing via a terminal.
Claude Opus 5 (Adaptive Reasoning, Max Effort) AGI artificial_analysis_intelligence_index Link Artificial Analysis Intelligence Index combines performance across ten evaluations: GDPval-AA, t2-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity’s Last Exam, GPQA Diamond, CritPt.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) AGI hle Link Humanity’s Last Exam: 2,500 challenging questions across over a hundred subjects at the frontier of human knowledge.
GPT-5.6 Sol (xhigh) Coding artificial_analysis_coding_index Link Artificial Analysis Intelligence Index combines performance across ten evaluations: GDPval-AA, t2-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity’s Last Exam, GPQA Diamond, CritPt.
Gemini 3 Pro Preview (high) Coding livecodebench Link LiveCodeBench: holistic, contamination-free evaluation of LLMs for code, updated continuously with new problems.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Coding scicode Link SciCode: models a realistic scientist workflow of identifying science concepts and transforming them into simulation code.
Grok 4.3 (medium) Instructions ifbench Link IFBench
Muse Spark 1.2 (xhigh) Long context Reasoning lcr Link AA-LCR: measures ability to extract and synthesise information from long-form documents ranging from 10k to 100k tokens.
GPT-5 (high) Maths aime Link AIME 2025: all 30 problems from the 2025 American Invitational Mathematics Examination.
GPT-5.2 (xhigh) Maths aime_25 Link AIM25
GPT-5.2 (xhigh) Maths artificial_analysis_math_index Link Artificial Analysis Intelligence Index combines performance across ten evaluations: GDPval-AA, t2-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity’s Last Exam, GPQA Diamond, CritPt.
Grok 4.6 (high) Reasoning gpqa Link GPQA: 448 multiple-choice questions written by domain experts in biology, physics, and chemistry.
Gemini 3 Pro Preview (high) Reasoning mmlu_pro Link MMLU-Pro: a more challenging and robust benchmark for language models across 12K complex questions in various disciplines.
GPT-5 (high) Saturated math_500 Link MATH-500: 500 problems spanning algebra, geometry, number theory, and probability, requiring step-by-step solutions.

This table shows the best performing models across all benchmarks.

1 c: Finding the balance


You can see that some models perform well on the benchmark at considerably less cost than the best performing model.

1 d: Exploring


Lots of actionable information on one graph.

Click on a point to reveal the model name and benchmark. The click also fixs all the datpoints for a given benchmark in focus. Click again to reset the view.

2 a: Human Judgement

Arena models_benchmarked best_model ranked_second_model
agent 10 Claude Opus 5 (High) Claude Opus 5 (Max)
code 50 claude-opus-5-max kimi-k3-max
document 39 claude-opus-5-high claude-opus-4-6
image-edit 53 gpt-image-2 (medium) grok-imagine-image-2.0 (low)
image-to-video 45 minimax-h3 dreamina-seedance-2.5-720p
search 34 gpt-5.6-sol-xhigh claude-opus-4-6-search
text-to-image 76 gpt-image-2 (medium) mai-image-2.6-preview
text-to-video 45 gemini-omni-flash flux-3-video
text 20 claude-fable-5 claude-opus-4-6-high
video-edit 9 dreamina-seedance-2.5-720p minimax-h3
vision 50 claude-fable-5 qwen3.8-max

Arena’s allow humans to choose (online) between different models on a given task. The table shows the results of a human judgement benchmarks.

2 b: Ranking


Opensource models are ranked on a leaderboard based on human judgements. Some of the Open Source models are near the top of the leaderboard and might be useful once we take into account the need for Digital Sovereignty or keeping your data within your own data center.

2 b: Ranking Safety


CISCO LLM Security Leadership board

Comprehensive model safety and security rankings, including single-turn score, multi-turn score, and detailed metrics.

See

Other available Security benchmarks:

2 c: Updating and Routing

id open_weight added_at url
deepseek-v4-flash-vision-exp TRUE 2026-08-21 17:51:18 DeepSeek
laguna-xs-2.1 TRUE 2026-08-14 16:26:20 Poolside
laguna-s-2.1 TRUE 2026-08-14 16:26:20 Poolside
qwen3.8-27b TRUE 2026-08-14 16:00:35 Alibaba Cloud / Qwen Team
glm-5.3 TRUE 2026-08-14 06:49:20 Zhipu AI
north-micro-vision-instruct TRUE 2026-08-13 19:06:17 Cohere
gemini-3.7-flash FALSE 2026-08-13 17:54:14 Google
deepseek-v4-pro-0813 TRUE 2026-08-13 17:29:21 DeepSeek
mai-code-1.1-flash FALSE 2026-08-13 17:05:20 Microsoft
sakana-namazu FALSE 2026-08-12 18:44:05 Sakana AI

LLM-Stats provides a live feed of model updates, including newly added models to the ecosystem. We can use this type of information to keep our benchmarks up to date and to identify promising new models for testing.

LLM-Stats also provides a routing service that automatically directs API calls to the best-performing model. We can use such a service to ensure that we always use the near best available model for specific capabilities. For a more complex community scorecard, it would require a discussion with a provider of routing services.

3 a: More information

⚠ Epoch table unavailable

This error is due to changes in the structure of the live data. The dashboard coding will need to be regularly updated to reflect the changes in the live data or a more robust form of data managling will need to be implemented.The Epoch best-model data frame was not produced by Explore_Benchmarks.R. Check that epoch.Rdata exists in DATA_DIR.


Here a list of best models for a given benchmark is shown. The dataset is from Epoch.AI. The dataset is updated regularly.

Certain models are robust across a series of benchmarks and therefore specific capabilities.

3 b: LLM input / Output


OpenRouter operates one of the largest AI inference platforms in the world, generating an authoritative empirical dataset relied upon by government agencies, academic researchers, major industry analysts, and global media outlets.

Here we use their dataset on the input and output modality of the models that they measure. For policymakers, by doing so, we can keep track of the expanding range of modalities and their popularity. This is an important consideration as you select benchmarks relevant to the LLMs in your organisation.

Conclusions

Choose AI models based on our values is possible via a set of current benchmarks. Much of the necessary benchmarks are available, but …

It is Nuanced:

AI alignment requires a nuanced understanding of the details for us to enact safely. However, the payoff is great.

We do need a community process. Here is a mockup of the dashboard that supports such a process.