Skip to content

Build1 publisher3 min readPublished

Orchestra builds its cheaper-model switch gate from the buyer's own corrected traces

The San Francisco startup records production work and expert corrections, turns them into task-specific tests, and reroutes traffic only when a cheaper candidate passes; the studies behind that are its own.

The Engineer · Build desk

Illustration accompanying Orchestra builds its cheaper-model switch gate from the buyer's own corrected traces

What happened

  • Orchestra sits between an application and providers such as OpenAI or Anthropic, testing cheaper routes against the model the customer already trusts.
  • The company says traffic moves to an improved prompt, a different existing model or a trained open-weight specialist only after the candidate meets the customer's quality, cost and latency requirements.
  • Orchestra's own studies report cost and latency gains on narrow tasks, while named customers, production savings, pricing and independent validation are all undisclosed.
  • Two former Instacart colleagues, Luis Manrique and Aamir Poonawalla, entered Y Combinator's Summer 2026 batch as Understudy and introduced the Orchestra brand ahead of September Demo Day.
  • Dealroom reports a $125,000 Y Combinator investment, and the supplied company and YC materials contain no financing announcement.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision With validation still first-party, the due diligence that pays is running the vendor's evaluation loop against your own traffic and reading its pass criteria before any requests move.
  • exposure A gateway on the request path ties the buyer's live latency and availability to a company whose own production record is months long.
  • constraint Every saving is gated on the buyer's own experts finding time to correct enough outputs to produce a grader that discriminates.
  • precedent If trace-to-specialist conversion holds up on narrow tasks, frontier API spend starts getting budgeted as data acquisition for a model the buyer keeps.

The loop closes only when someone labels. Orchestra records selected production work and expert corrections, then turns those examples into task-specific evaluations [1]. Those corrections come from people inside the buyer. runtimewire.com reports that the technical work matters only once Orchestra has convinced a buyer to route live traffic through its gateway and to provide enough examples to define what a successful response looks like [19].

Sample size on that graded set decides what the switch gate can see. A grader built from a hundred corrections will pass candidates that a thousand corrections would have failed, and a buyer cannot tell which of the two it holds without the eval and its counts in front of it. In my view the eval and those counts are what to ask for before the pilot.

For gains measured on a narrow task to carry into another workload, three conditions have to hold. The input distribution has to be stable enough that last quarter's traces predict next quarter's. A correct response has to be definable tightly enough for a small model to learn it from the available corrections. And the failure modes the buyer cares about have to appear in the graded set. Classification and extraction usually clear that bar; open-ended drafting usually does not.

The saving a buyer books is baseline frontier spend, minus the cost of serving the specialist, minus what Orchestra charges to sit on the request path [17]. A pilot produces the first two numbers, and Orchestra sets the third.

The founders wrote in their YC company profile that customers should be "accumulating intelligence, not accumulating API bills" [5]. runtimewire.com frames the company as testing whether production AI spending can create an owned technical asset instead of a recurring API dependency [21]. Two assets come out of the loop, and ownership of the second one is not disclosed. The distilled specialist decays when the traffic mix moves, while the evaluation set can be re-run against any later candidate, including the frontier model it was built to replace.

Two people intend to own evaluation, optimization, training and production routing without building frontier models [7]. YC lists the company as active, founded in 2026 and staffed by its two founders [8]. Poonawalla spent about a decade at Instacart, most recently as a senior software-development manager leading its ads-serving and infrastructure group [11]. YC credits him for work on the auction platform and experimentation systems [12]. Ads serving is a latency-budgeted routing problem with a measurement loop attached, so that experience is in roughly the right shape for this product. Orchestra also credits Poonawalla with leading Instacart's Curbside Pickup as it grew from zero to approximately $4 billion in gross transaction value [13]. Manrique was employee number two at Gumloop, where YC says he helped close approximately $2 million in the first year [14]. Before that he was a principal product manager at Instacart leading advertising machine-learning work [15].

What to watch

  • A named customer publishing before-and-after production numbers would move the cost claim off first-party studies.
  • Published gateway pricing would fill in the third term of the savings calculation.
  • Contract language on who owns the evaluation set decides whether a buyer keeps anything after dropping the gateway.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories