Opus 5 has taken the top spot on the Artificial Analysis Intelligence Leaderboard, an aggregate benchmark that scores models across reasoning, math, coding, and knowledge tasks. The practical takeaway: if you rely on leaderboards to shortlist models, Opus 5 now sits at the front for general intelligence-style tasks—but a #1 ranking is a starting point, not a verdict for your workload.

Why this matters: the Artificial Analysis index blends multiple public evaluations into a single number, which makes it useful for a quick apples-to-apples comparison. It's a reasonable filter when you're deciding which two or three models to test seriously. It is not a substitute for measuring performance on your actual prompts, data, and latency budget.

Opus 5 Tops the Artificial Analysis Intelligence Leaderboard

The number to watch alongside the ranking is cost and speed. A top-scoring model that runs slower or costs several times more per token can be the wrong choice for high-volume production tasks. Check the leaderboard's companion metrics—output speed (tokens/sec), time to first token, and price per million tokens—before you commit. The best model for a chatbot at scale is often not the highest-ranked one.

What you can do now: pull your own representative test set—10 to 30 tasks that mirror real usage—and run Opus 5 against your current model. Score on correctness, format adherence, and refusal behavior, then divide by cost per successful completion. That single ratio usually tells you more than any leaderboard position.

Finally, treat rankings as perishable. Frontier models leapfrog each other on a monthly cadence, and index scores shift as new evaluations get added. Bookmark the source, re-check when you plan a migration, and keep your internal eval harness so you can re-benchmark in an afternoon rather than starting from scratch.