Skip to content

Leadership1 publisher3 min readPublished

Swapping LLMs barely moved Orient Software's document extraction quality

Son Nguyen says his team ran the same prompt and the same use case through different models, found the difference surprisingly small, and rebuilt the preprocessing, context and validation around the model instead. He calls that surrounding system the agent harness.

The Board Room · Leadership desk

What happened

  • Son Nguyen, Orient Software's CTO and a Neurond AI co-founder, says clients open almost every meeting by asking which model to use, then pull up benchmarks, debate context windows and pass around pricing sheets.
  • On a document intelligence project his team ran the same prompt and the same use case through different LLMs, expecting a noticeable jump in extraction quality, and found the difference surprisingly small.
  • The problems sat in document layout, tables and OCR, the context supplied and how results were validated, so the team improved preprocessing and added validation to catch errors before they reached the client.
  • He wrote that he has seen teams pick the best model on paper and stall for six months, while others built something transformative with a model that never topped a leaderboard.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • constraint An application coupled tightly to one model turns each vendor release into a reengineering bill, on Nguyen's account, so the harness is what decides whether next year's model change is a configuration edit or a project.
  • decision The choice a buyer controls this quarter is who owns business data, validation and approval rules, because those three layers are not on any model vendor's price list.
  • exposure Permissions are layer seven of the buyer's own design, so when an agent can delete as well as read, it is the customer who built that reach.
  • cost The one figure in the column is the six-month stall, with no team count behind it, so a board asking what harness work pays back is being handed an anecdote.

The evidence here is one practitioner's account of one project. Nguyen is CTO of Orient Software and co-founded Neurond AI, firms that sell software development, AI and data science services [1]. Orient Software sells the integration work this argument locates the hard part in, and integration work is billable. His own experiment cuts the other way on cost: the model swap was the cheap thing to try, and he reports it bought almost nothing [6]. The column does not publish the extraction scores, so "surprisingly small" is his characterisation [20].

Same prompt, same use case, different LLMs [6]. That design varies the model and holds everything else fixed, so it tells you what a model change is worth inside one existing pipeline, and leaves the worth of the rebuild unmeasured. The preprocessing, context and validation changes came afterwards, with no controlled comparison against the old configuration [8][18].

"Give the same model different tools, memory, context and guardrails, and you get dramatically different results," Nguyen wrote [5]. The document project tested the first clause and reported the answer. The second clause rests on the client work he describes.

Nguyen lists seven layers around the model: context management, memory, tools and integrations, orchestration, business data, verification, and security and permissions [10]. Three of those are defined by the customer. Business data means your pricing rules, policies and customer records, which he says no general-purpose model knows [13]. Verification means checking the work instead of trusting the first answer [10]. Permissions decide what an agent may do alone, and, as he puts it, "Reading a document is very different from deleting a file" [12]. Business data, verification and permissions all sit outside a model vendor's pricing sheet [19].

His forward case is about swap cost. If an application is tightly coupled to one model, he writes, every release cycle may become an expensive reengineering project, while a well-designed harness lets business logic, data pipelines, integrations and evaluations stay intact as the model underneath changes [15]. He describes a Claude Fable 5 generation built for long-horizon autonomous work, Gemini 3.7 Flash with adjustable reasoning levels and a 1-million-token context window, and DeepSeek as an open-weight family of V4 Pro, Flash and a vision variant [16]. "Once the system was designed properly, the LLM became a component we could swap, not the foundation everything depended on," he wrote [17].

The six-month stall is the line most likely to be quoted from this column: "I've seen teams pick the best model on paper and stall for six months," Nguyen wrote [3], with no team count and no period attached [20]. A buyer can run the cheaper half of his experiment in-house in days, holding the current pipeline and prompt and putting a second model behind them [6]. If output quality barely moves, the benchmark comparison has answered itself, and the argument moves to who is going to own the OCR, the validation and the approval rules [7].

What to watch

  • A before-and-after measurement of a harness rebuild, with published scores, would move this from account to evidence.
  • Whether model vendors absorb memory, tools and verification into their own platforms, shrinking what a buyer has to build.
  • Whether buyers start requiring a documented model-swap test on the existing pipeline before sign-off.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories