Build1 distinct publisher2 min readUpdated
Four developers, 100 million requests a day, and a standing ban on outside pull requests: SpareBank 1 Utvikling's account team says what the model lacks is tacit context.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Where they drew the line is more useful than the verdict itself. Analysis, telemetry visualisation and kickstarting a task are all jobs where a human reads the output before anything reaches a customer [8]. Domain logic in a brownfield banking API is the other kind of job, and the team's argument for keeping it human is about unwritten constraints rather than syntax [9].
Consider what the team carries. Four developers and one product leader [4] own account APIs that take around 100 million requests a day on the most central endpoints [5]. Flatten that and it is roughly 1,160 requests a second, sustained, before any peak [1]. The bank behind those APIs has about 1.2 million customers in a country of five million people [6], somewhere near a quarter of the population [2]. Soderbom's description of the failure mode is not abstract: get it wrong and "we are newspaper headlines pretty fast" [7].
That is also why the cadence claim lands differently here than in a vendor deck. This is a team that already put mob programming on everything, including documentation and analysis, explicitly to spread domain knowledge and remove silos [17]. The tacit context an agent is missing is the same context the team spends its whole working day circulating between five people. A model joins that room only through the transcript, and there is no transcript.
The refusal of outside pull requests reads as an architectural decision rather than a cultural preference [10]. A team that will not accept an asynchronous patch from another human team has already answered the question about accepting one from a coding agent, because the artifact and the integration path are identical. Their onboarding runs on the same principle in reverse: shadow the mob, contribute to production on day one [11].
Discount it appropriately. This is one team, self-reported, in a conversation with an editor who had interviewed them a year earlier, when their line was that AI-generated code was not fast for them [16]. The published research paper is about how to introduce pair programming into a company, not about model output quality [12]. Nobody puts a number on the "subpar" code. What survives the discount is the criterion they use instead of code review taste: not good code versus bad code, but code that is easy to change, with high cohesion, low coupling and TDD following from that goal [14]. A generator that cannot see the constraints a change will hit cannot optimise for changeability, and speed of production is not a proxy for it.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Soderbom: the team does mob and pair programming on absolutely everything, no task is done in isolation, they use TDD, and they put things into production up to several times an hour.
Asgaut Mjolne Soderbom and Ola Hast discuss adopting Claude Code and the reasons why they consider it good for everything else, but not coding.
Podcast key takeaway: LLMs excel at analysis, telemetry visualisation and task kickstarting.
Podcast key takeaway: a strict 'no outside pull requests' policy means external changes require active pair programming rather than asynchronous, lower-quality integration.
Asgaut Mjolne Soderbom and Ola Hast are both senior developers at SpareBank 1 Utvikling.
SpareBank 1 Utvikling is a software company owned by 12 banks in an alliance and builds the digital bank solutions for those banks, private and company sector.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source practitioner self-report
All material rests on one InfoQ podcast in which the two subjects describe their own team. The operational specifics are concrete and internally consistent, and a published SINTEF paper is cited, but no telemetry, evaluation method, incident data, or third-party confirmation is supplied, and the LLM verdict has no disclosed task sample or model detail.
One four-developer team, scoped tool use
Adoption evidence is real but narrow: a single four-developer team at one bank-alliance software company running Claude Code in a restricted, non-coding role, alongside a multi-year internal pair-working programme covering seven team experiments. There is no evidence of uptake beyond this organisation.
Restrained interview, generalised headline
The interview itself is sober and specific, which limits overstatement. The mild positive gap comes from generalising claims — 'LLMs lack tacit context in brownfield codebases', day-one productivity from mob programming, pairing as the foundation of high-velocity delivery — that are presented as takeaways while resting on one small team's unmeasured experience.
Practitioner advocacy, no vendor stake
The speakers are employees promoting their own way of working and are the subjects of research they helped run, and this is a repeat appearance on the same outlet's podcast, so there is reputational incentive to present the practice favourably. There is no disclosed commercial relationship with Anthropic or any tooling vendor, and the AI conclusion is unflattering to the tool rather than promotional.
Clear on what was said, thin on verification
What the team asserts is unambiguous and directly quoted, so the reporting of positions is reliable. Confidence in the underlying operational and LLM-capability claims is limited by the single-publisher, single-team, self-reported evidence base and by the absence of any measurement behind the generalised takeaways.
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
build
Superpowers makes spec-driven work a precondition, then ships it to twelve harnesses1 distinct publisher
build
NVIDIA put a number on agent skills: 300+ verified, two harnesses, baselines under 50/1001 distinct publisher
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 24, 2026