Skip to content

model

GPT-OSS 120B

GPT-OSS 120B is an open-weight large language model in OpenAI's gpt-oss family, used for reasoning and agentic tasks and benchmarked against other models.

Known aliases

  • 120 B coordinator
  • GPT-OSS
  • gpt-oss 120B
  • gpt-oss-120b
  • OpenAI GPT-OSS 120B

Relationships

No evidence-backed relationships are recorded.

Current stories

build3 publishers

Typesafe's Jev API gains an open-weight rival in Cloudflare's Clef decision models

Cloudflare released two Jev-API-compatible decision models, Clef and Clef-flash, on Workers AI and as Apache 2.0 weights on Hugging Face. Typed classification steps in agent code can now move between providers or onto owned hardware, as long as they stay inside the text-only, 32k-context features Jev supports.

Perspective Coverage

3 publishers
Builder
Builder 53%
Operator
Operator 27%
Investor
Investor 20%

Reality

Evidence50
Adoption18
Hype gap+35
Incentives70
Confidence60
build1 publisher

Probes on a 27B open model match direct probes of a 397B model on deception

Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.

Publishers:lesswrong.com

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence30
build1 publisher

Blender 5.0 rejects three in ten scripts that ten LLMs wrote for it

Ten LLMs' Blender 5.0 scripts ran only 70% of the time when a Kaggle benchmark executed them in 5.0, against 91% for scripts targeting 3.6. Renamed and removed APIs look like valid code, so the benchmark grades each answer in the exact build the prompt named.

Publishers:dev.to

Reality

Evidence50
Adoption
Insufficient
Hype gap+5
Incentives30
Confidence45
invest7 publishers

OpenAI's first chip beat merchant silicon in 16 months. The moat was never the fab

Jalapeno's lead is measured against last generation and the volumes are tiny, but a first-pass ASIC clearing Nvidia, AMD and Google parts reprices the design barrier, not the supply chain.

Perspective Coverage

7 publishers
Builder
Builder 30%
Operator
Operator 24%
Investor
Investor 46%

Reality

Evidence50
Adoption8
Hype gap+40
Incentives65
Confidence55
build9 publishers

OpenAI's first Jalapeno numbers buy it leverage, not a procurement input

The 1.7x to 3.6x latency range is set by the baseline systems, not the chip, and the report's own publication date is unsettled. Read it as direction, not evidence.

Perspective Coverage

9 publishers
Builder
Builder 41%
Operator
Operator 31%
Investor
Investor 28%

Reality

Evidence52
Adoption14
Hype gap+38
Incentives82
Confidence68

Earlier coverage

  1. The conductor is the bottleneck: local agent stacks fail at orchestration, not at the workers

    Build · August 18, 2026 · 1 publisher

  2. Agent memory has a dose-response curve, and the cheapest dose won the biggest gain

    Build · August 18, 2026 · 1 publisher

  3. Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist

    Leadership · August 18, 2026 · 1 publisher

  4. 18x per joule in 16 months, and most of it was not your model choice

    Invest · August 16, 2026 · 1 publisher

  5. Seven local models, one prompt, one DGX Spark: the speed ranking decided nothing

    Build · August 15, 2026 · 1 publisher