Skip to content

benchmark

Artificial Analysis Intelligence Index

A composite benchmark that scores large language models across multiple evaluations into one comparative intelligence rating from Artificial Analysis.

Known aliases

  • AAII
  • AA Index
  • AA Intelligence Index
  • Artificial Analysis
  • Artificial Analysis Intelligence Index v4.1.1
  • Intelligence Index v4.1.1

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Grok 4.7 doubles cost per task at an unchanged per-token price

Grok 4.7 keeps Grok 4.6's $2/$6 token price yet costs $3.74 per task against $1.86, by Artificial Analysis' measurement. Teams that budget from the price sheet will undercount agent spend until they measure tokens per task on their own work.

Publishers:dev.to

Reality

Evidence58
Adoption
Insufficient
Hype gap+45
Incentives55
Confidence55
leadership6 publishers

Gemini 4 Argon ties two rivals on a leading AI index while some Google staff doubt its coding

Google's Gemini 4 Argon tied GPT-6 Astra and Claude Fable 5.1 at 53 on the Artificial Analysis index, though some Google staff say its coding lags. Google disputes them, and until paid API access has a date, buyers cannot check the score on their own code.

Publishers:9to5google.comaibreakfast.beehiiv.combusinessinsider.comcsoonline.comimplicator.aitheguardian.com

Perspective Coverage

6 publishers
Builder
Builder 37%
Operator
Operator 35%
Investor
Investor 28%

Reality

Evidence55
Adoption15
Hype gap+25
Incentives65
Confidence55
science4 publishers

Gemini 4 Argon costs 2.7 times as much per task as GPT-6.1 Sol at the same token price

Google's Gemini 4 Argon matches GPT-6.1 Sol's $2/$10 token price but costs 2.7 times as much per task, according to Artificial Analysis. Argon uses more tokens per job, so buyers still have to compare frontier models by cost per completed task.

Perspective Coverage

4 publishers
Builder
Builder 36%
Operator
Operator 34%
Investor
Investor 30%

Reality

Evidence68
Adoption15
Hype gap+20
Incentives55
Confidence65
invest13 publishers

Google prices its third-ranked Gemini 4 Argon at half the cost of Anthropic's Claude Opus 5.5

Google priced Gemini 4 Argon at half Claude Opus 5.5's per-token rate for a model one composite index ranks third, behind two Anthropic models. On price alone it only matches OpenAI's newly discounted GPT-6.1 Sol, so the price edge Google is selling is against Anthropic.

Perspective Coverage

13 publishers
Builder
Builder 37%
Operator
Operator 31%
Investor
Investor 32%

Reality

Evidence55
Adoption20
Hype gap+30
Incentives65
Confidence55
build1 publisher

Gemini 3.8 Flash ties Opus 5 on DeepSWE at a price Google doubles on January 1

Google's Gemini 3.8 Flash ties Claude Opus 5 at 74% on DeepSWE for $2.36 a task, at an introductory price that doubles on January 1, 2027. For agent workloads, the comparison that holds up after January is cost per finished task, set by steps taken as much as by rate.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives60
Confidence50
product3 publishers

Meta keeps Muse Spark 1.3 pricing flat while claiming coding edge over GPT-5.6

Alexandr Wang says the new model costs developers no more than 1.2 and finishes the same work on about a quarter fewer tokens. That is a real saving on high-volume code generation and a rounding error most other places.

Perspective Coverage

3 publishers
Builder
Builder 40%
Operator
Operator 32%
Investor
Investor 28%

Reality

Evidence55
Adoption30
Hype gap+25
Incentives70
Confidence55
build6 publishers

Spark 1.3's index jump lands on the three tests that carry half the score

Meta's fourth Muse Spark in five months gains four points on Artificial Analysis' Intelligence Index, almost entirely in agentic tests, while two scores fall and the tier Meta benchmarked against rivals is still a limited preview.

Perspective Coverage

6 publishers
Builder
Builder 52%
Operator
Operator 26%
Investor
Investor 22%

Reality

Evidence68
Adoption25
Hype gap+30
Incentives65
Confidence70
build3 publishers

DeepSeek's new encoder-decoder splits inference into an 8B prefill and a 16B decode

V4.1-Flash retires the V4 Pro line and carries two active-parameter counts, 763B total with 8B on input tokens and 16B on output, so one sizing number no longer covers both phases of a request. Baseten had it running on day zero.

Publishers:businesstimes.com.sgdev.tolatent.space

Perspective Coverage

3 publishers
Builder
Builder 40%
Operator
Operator 28%
Investor
Investor 32%

Reality

Evidence60
Adoption35
Hype gap+25
Incentives40
Confidence58

Earlier coverage

  1. NVIDIA's edge-agent case rests on compact open models matching data-center capabilities

    Build · September 4, 2026 · 1 publisher

  2. Astra's 99.9% holds up only on the harness OpenAI ran itself

    Invest · September 4, 2026 · 1 publisher

  3. Gemini 3.8 Flash buys three index points for a 45% rise in cost per task

    Leadership · September 3, 2026 · 1 publisher

  4. Meta's Muse Spark 1.3 matches three flagship models at 55 cents a task

    Leadership · September 2, 2026 · 1 publisher

  5. Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8

    Build · August 31, 2026 · 1 publisher

  6. Z.ai's cost-parity claim on Chinese accelerators rests on model design as much as silicon

    Leadership · August 27, 2026 · 1 publisher

  7. GPT-5.6 Luna scores 52 against a peer median of 17. Its token count is what lands on your bill

    Science · August 20, 2026 · 1 publisher

  8. GLM-5.3 changed nothing but the training environments. That is the whole test.

    Build · August 19, 2026 · 3 publishers

  9. A 27B laptop model scores like a rented one, and thinks three times as hard to do it

    Product · August 19, 2026 · 1 publisher

  10. The AI bill nobody reconciles: cost per finished task, not per million tokens

    Leadership · August 18, 2026 · 1 publisher

  11. Korea cuts one of four sovereign AI teams, and usability did the cutting

    Invest · August 18, 2026 · 1 publisher

  12. GLM-5.3 says the quiet part: the base model did not change, the post-training did

    Science · August 16, 2026 · 1 publisher

  13. Three frontier launches in a day, all pitched on price. Open weights set the ceiling.

    Build · August 14, 2026 · 4 publishers

  14. Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem

    Build · August 14, 2026 · 1 publisher