Skip to content

Topic

LLM code generation against internal APIs

Natural-language-to-code assistants that emit client code for a specific in-house API, and the failure mode of confident, well-formed, non-functional output.

Current stories

build1 publisher

Blender 5.0 rejects three in ten scripts that ten LLMs wrote for it

Ten LLMs' Blender 5.0 scripts ran only 70% of the time when a Kaggle benchmark executed them in 5.0, against 91% for scripts targeting 3.6. Renamed and removed APIs look like valid code, so the benchmark grades each answer in the exact build the prompt named.

Publishers:dev.to

Reality

Evidence50
Adoption
Insufficient
Hype gap+5
Incentives30
Confidence45