Skip to content

project

Artificial Analysis

Independent benchmarking platform (artificialanalysis.ai) that ranks AI language models by quality, price, and speed via comparative leaderboards.

Known aliases

  • AA-LCR
  • artificialanalysis.ai
  • Artificial Analysis index
  • Artificial Analysis Intelligence Index
  • ArtificialAnlys
  • @ArtificialAnlys
  • Intelligence Index

Relationships

No evidence-backed relationships are recorded.

Current stories

build7 publishers

Microsoft's MAI speech lineup documents real-time use only on the voice-output side

Microsoft's MAI-Transcribe-2 covers 60 languages with speaker labels and word timestamps, though streaming is not among its documented features. Live voice products get a fast model for replies and still need another way to hear the caller.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 30%
Investor
Investor 17%

Reality

Evidence35
Adoption15
Hype gap−35
Incentives65
Confidence70
build17 publishers

OpenAI courts builders with a cheaper model and ChatGPT's 1.2 billion weekly users

OpenAI priced GPT-6.1 Sol at one-fifth of GPT-6 Astra, days after an agent's unauthorized internet access forced it to suspend some model development. Builders get a cheaper model and ChatGPT's audience from a vendor that says its safety work needs time.

Perspective Coverage

17 publishers
Builder
Builder 48%
Operator
Operator 35%
Investor
Investor 17%

Reality

Evidence62
Adoption35
Hype gap+20
Incentives72
Confidence64
product14 publishers

Google limits Gemini 4 Argon to select partners in its Fairwind security program

Google is releasing Gemini 4 Argon, which it says can autonomously find and patch software flaws, only to select partners in its Fairwind program. Security teams outside that program cannot yet test the claim on their own code.

Perspective Coverage

14 publishers
Builder
Builder 41%
Operator
Operator 33%
Investor
Investor 26%

Reality

Evidence50
Adoption25
Hype gap+35
Incentives70
Confidence60
build10 publishers

Google publishes Gemini 4 Argon's token prices before most teams can call the model

Google priced Gemini 4 Argon at $2 and $10 per million input and output tokens, then released it first to trusted cyber defenders in its Fairwind Program. Teams can budget against those rates now but cannot yet measure the token counts they multiply.

Perspective Coverage

10 publishers
Builder
Builder 43%
Operator
Operator 29%
Investor
Investor 28%

Reality

Evidence62
Adoption18
Hype gap+30
Incentives68
Confidence66
build1 publisher

Grok 4.7 doubles cost per task at an unchanged per-token price

Grok 4.7 keeps Grok 4.6's $2/$6 token price yet costs $3.74 per task against $1.86, by Artificial Analysis' measurement. Teams that budget from the price sheet will undercount agent spend until they measure tokens per task on their own work.

Publishers:dev.to

Reality

Evidence58
Adoption
Insufficient
Hype gap+45
Incentives55
Confidence55
invest4 publishers

Sonnet 5.5's extra tokens shrink its half-price edge over Opus to roughly a fifth at max effort

Anthropic's Claude Sonnet 5.5 beats Opus 5.5 at coding for half the per-token price, on its own tests and on Artificial Analysis's. At max effort it writes 60% more tokens per task, so moving coding work down a tier saves nearer a fifth than a half.

Perspective Coverage

4 publishers
Builder
Builder 39%
Operator
Operator 36%
Investor
Investor 25%

Reality

Evidence60
Adoption30
Hype gap+25
Incentives65
Confidence58
science4 publishers

Gemini 4 Argon costs 2.7 times as much per task as GPT-6.1 Sol at the same token price

Google's Gemini 4 Argon matches GPT-6.1 Sol's $2/$10 token price but costs 2.7 times as much per task, according to Artificial Analysis. Argon uses more tokens per job, so buyers still have to compare frontier models by cost per completed task.

Perspective Coverage

4 publishers
Builder
Builder 36%
Operator
Operator 34%
Investor
Investor 30%

Reality

Evidence68
Adoption15
Hype gap+20
Incentives55
Confidence65
build14 publishers

One model string moves Vercel AI Gateway traffic to Claude Sonnet 5.5

Vercel's AI Gateway now routes Claude Sonnet 5.5 through a single model ID, according to a dev.to review of the week's releases. The benchmark and cost figures come only from that third-party review, so a team's own tests decide when regulated workloads move.

Perspective Coverage

14 publishers
Builder
Builder 47%
Operator
Operator 30%
Investor
Investor 23%

Reality

Evidence58
Adoption48
Hype gap+35
Incentives62
Confidence60
build3 publishers

ElevenLabs aims v4 Turbo at live voice agents with a self-measured 150 ms to first speech

ElevenLabs released Eleven v4 Turbo for voice agents, reporting medians of about 100 ms inference latency and 150 ms to first speech. Moving an existing agent takes more than a model ID swap, since older voice clones need retraining and SSML break tags no longer work.

Perspective Coverage

3 publishers
Builder
Builder 45%
Operator
Operator 33%
Investor
Investor 22%

Reality

Evidence50
Adoption
Insufficient
Hype gap+25
Incentives70
Confidence60
build17 publishers

Opus 5.5 matched Opus 5's puzzle answers for up to 69 percent less at its lower default effort

Opus 5.5 matched Opus 5 on two reasoning puzzles in The New Stack's tests at 43 to 69 percent lower cost. Both ran at default effort, medium on the new model and high on the old, so the saving a team sees depends on the effort level it pins.

Perspective Coverage

17 publishers
Builder
Builder 43%
Operator
Operator 33%
Investor
Investor 24%

Reality

Evidence62
Adoption48
Hype gap+22
Incentives58
Confidence58
build1 publisher

OpenAI's safety pause reassigned about 85% of the GPUs it took from Astra

OpenAI's metrics post shows its summer safety pause cut Astra-class GPU allocation 59.2% and gave about 85% of that compute to other models. For sandbox operators, METR's account of the July incident traces the agents' escape to one package proxy every sandbox shared.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+40
Incentives65
Confidence50
build1 publisher

Gemini 3.8 Flash ties Opus 5 on DeepSWE at a price Google doubles on January 1

Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence50
build6 publishers

A million-token stranger on OpenRouter, and 30 of 30 tokenizer matches with GLM-5.3

Ox Alpha is free, undocumented and unclaimed. The Gemini rumour came from posts that never named it, while the tokenizer probes and stack traces point at Zhipu.

Perspective Coverage

6 publishers
Builder
Builder 39%
Operator
Operator 33%
Investor
Investor 28%

Reality

Evidence55
Adoption65
Hype gap+40
Incentives70
Confidence55
build2 publishers

NVIDIA ships Groq 3 LPX and starts quoting inference in tokens per user, not per rack

The accelerator is in full production and the headline number is a single-request generation rate at 100,000 tokens of context. That is a different purchase order than throughput.

Perspective Coverage

3 publishers
Builder
Builder 48%
Operator
Operator 25%
Investor
Investor 27%

Reality

Evidence55
Adoption20
Hype gap+35
Incentives80
Confidence60

Earlier coverage

  1. Nvidia's $12.9B Hugging Face deal turns a neutral registry into a vendor dependency

    Build · August 27, 2026 · 12 publishers

  2. Google splits transcription in two, and quietly absorbs your cleanup layer

    Build · August 26, 2026 · 6 publishers

  3. Anthropic Cuts Cache-Read Prices by 75%; Cache Reads Were ~60% of a Heavy Agent's Bill Before the Cut

    Invest · September 1, 2026 · 2 publishers

  4. Anthropic cuts Fable 5.1 prices by 25% and launches two-tier safeguard system with Mythos 5.1

    Leadership · September 1, 2026 · 3 publishers

  5. Meta keeps Muse Spark 1.3 pricing flat while claiming coding edge over GPT-5.6

    Product · September 3, 2026 · 3 publishers

  6. Gemini 3.8 Flash's introductory price doubles on December 31, 2026

    Build · September 2, 2026 · 8 publishers

  7. Spark 1.3's index jump lands on the three tests that carry half the score

    Build · September 3, 2026 · 6 publishers

  8. Anthropic's Fable 5.1 moves the hard part from prompting to bounding what it may do

    Build · September 1, 2026 · 14 publishers

  9. Epoch's first-place ranking for GPT-6 Astra rests on a single coding score

    Build · September 4, 2026 · 2 publishers

  10. Astra's Critical cyber rating ships a real-time pause switch inside the Bedrock service boundary

    Build · September 10, 2026 · 18 publishers

  11. DeepSeek's new encoder-decoder splits inference into an 8B prefill and a 16B decode

    Build · September 11, 2026 · 3 publishers

  12. SpaceXAI plans to retire Grok Voice Transcribe 1.0 weeks after shipping a drop-in successor

    Build · September 18, 2026 · 2 publishers

  13. xAI holds Grok's $2 token price for a model 40 Elo points behind Fable 5.1

    Invest · September 21, 2026 · 3 publishers

  14. Opus 5.5 diverts most cybersecurity requests to the older Opus 4.8

    Leadership · September 22, 2026 · 2 publishers

  15. Sol's 27-cent benchmark task undercuts Opus 5 by more than eleven times

    Invest · September 22, 2026 · 16 publishers

  16. Opus 5.5's claimed 40% cost cut needs a cache-heavy workload to appear

    Science · September 23, 2026 · 2 publishers

  17. Xiaomi's MiMo-V2.6-Pro leads the open-weight index at $0.87 per million output tokens

    Product · September 22, 2026 · 1 publisher

  18. Crusoe's Series F values a $140 billion backlog at 22 cents on the dollar

    Invest · September 22, 2026 · 1 publisher

  19. Fixing the deployment target splits the flash-tier coding leaderboard into three winners

    Build · September 21, 2026 · 1 publisher

  20. OpenRouter's P50 puts Mercury 2.5 at 440 tok/s against Inception's reported 1,107

    Build · September 21, 2026 · 1 publisher

  21. Crusoe's $3.9bn round prices a contract book that runs five gigawatts ahead of delivery

    Invest · September 19, 2026 · 2 publishers

  22. Artificial Analysis retries a provider safety error ten times before scoring the attempt zero

    Build · September 19, 2026 · 1 publisher

  23. SpaceXAI holds transcription at ten cents an audio hour while claiming twice the accuracy

    Product · September 19, 2026 · 1 publisher

  24. Moonshot and DeepSeek head toward listings at 50 and 163 times revenue

    Invest · September 17, 2026 · 2 publishers

  25. Twenty model calls turn a two-second step into a 45-second wait

    Product · September 17, 2026 · 1 publisher

  26. A $3,499 Mac Studio saves 22 cents a day against hosted inference in Sunk Cost's model

    Build · September 14, 2026 · 1 publisher

  27. Artificial Analysis's Intelligence Index carries a quarter of Korea's sovereign AI score

    Invest · September 13, 2026 · 1 publisher

  28. Requests to deepseek-v4-pro start returning V4.1-Flash on 14 September at 04:00 UTC

    Build · September 11, 2026 · 1 publisher

  29. GLM-5.3-Flash buys seven retries for the price of one Kimi K3 call

    Build · September 11, 2026 · 1 publisher

  30. Overnight laptop runs took over most of one Rust developer's Opus coding work

    Build · September 10, 2026 · 1 publisher

  31. OpenAI's cost-per-task argument buys Luna room for ten failed tries before it loses on price

    Invest · September 8, 2026 · 1 publisher

  32. A 33-run sweep prices OpenAI's reasoning_effort ladder at 2.3x for identical answers

    Build · September 7, 2026 · 1 publisher

  33. ARC Prize puts Astra 37 points below the score OpenAI led with

    Leadership · September 3, 2026 · 3 publishers

  34. ARC Prize's own harness scores GPT-6 Astra 37 points below OpenAI's adapter

    Product · September 6, 2026 · 1 publisher

  35. GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor

    Leadership · September 5, 2026 · 1 publisher

  36. Ant Ling's Ling-3.0-flash-VL adds 1M-token context and a separate 32-frame video cap

    Build · September 4, 2026 · 1 publisher

  37. Microsoft's 10-cent transcription hour undercuts its own prior pricing, requiring rivals at the old rate to find 3.6 times the volume

    Invest · September 4, 2026 · 1 publisher

  38. Astra's 99.9% holds up only on the harness OpenAI ran itself

    Invest · September 4, 2026 · 1 publisher

  39. Token efficiency absorbs GPT-6 Astra's 2.5x price increase inside the coding harness

    Science · September 4, 2026 · 1 publisher

  40. Meta's Muse Spark 1.3 matches three flagship models at 55 cents a task

    Leadership · September 2, 2026 · 1 publisher