Skip to content

Build1 publisher3 min readPublished

Interfaze's open-weight interfaze-1-lite attaches confidence scores and bounding boxes to extracted data

Interfaze released interfaze-1-lite, an Apache 2.0 open-weight model for OCR and speech that returns confidence scores and bounding boxes with its answers. Teams can send doubtful rows to a reviewer once they have tested how the scores behave on their own documents.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Interfaze's open-weight interfaze-1-lite attaches confidence scores and bounding boxes to extracted data
Generated illustration

What happened

  • Interfaze calls the design Mixture of Architecture: a reasoning core interprets each request and hands work to specialist models that read documents, recognize speech and locate objects.
  • The model accepts text, images, audio and files, including PDFs and Word documents, and returns either text or JSON.
  • Interfaze says the full model runs on a single 80GB GPU such as an H100, and its setup instructions require a GPU with compute capability 8.9 or newer.
  • Teams that do not self-host can use the hosted API at $0.85 per million input tokens and $1.50 per million output tokens.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability With Apache 2.0 weights that load through Transformers, a team can run extraction on its own GPU and keep invoices and recordings away from a third-party API.
  • constraint Teams whose GPUs are older than the documented minimum, or lack 80GB of memory, must either use the metered API or buy hardware before they can self-host.
  • decision The lite model scores below the larger interfaze-1 on text-to-SQL and multilingual Q&A, so teams whose workloads lean on those tasks give up accuracy by choosing it.

Interfaze's own document example shows how the check works. The model extracts transactions into a fixed JSON schema. Beside each value it returns the OCR text, the value's position on the page and a confidence score [6]. A threshold sits on that score. Rows that fall below it go to a human reviewer instead of being posted automatically [7]. The bounding box shows the reviewer which part of the page to check the value against [2].

The specialist models are there to produce that evidence. Interfaze says the combination is meant to pair schema-following and contextual reasoning with the evidence metadata that specialized vision and audio systems return [5]. I think the split is right for extraction work, where a reviewer's first question is where on the page a number came from.

The capability list goes well past documents and speech. It has nine items, including GUI detection, translation, forecasting and guardrails [10] [23]. Forecasting is a lot to ask of a model called lite.

A threshold works only if the score is calibrated on the documents a team actually processes. Runtimewire, which reported the launch, says a confidence score alone does not establish that a system is reliable. It adds that developers still have to check the scores against their own documents and pick the point at which a row goes to review [8]. The failure case is a wrong value with a high score. The threshold sends that row straight to posting. Catching it takes a labelled sample from real traffic. Plot correctness against score, then put the cut where errors start to appear above it.

The benchmark table needs the same treatment. Interfaze reports interfaze-1-lite ahead of four commercial models on five tests, including olmOCR document processing and structured-output value accuracy [15]. Two of those leads are thin. On olmOCR it scores 83.8% to Claude-Sonnet-5's 83.5% [17], a gap of 0.3 points [20]. On structured output it scores 81.5% to Gemini-3.7-Flash's 80.2% [18], a gap of 1.3 points [21]. The company's leaderboard mixes evaluations it ran independently with data from the model providers, and some scores are self-reported [19]. The olmOCR lead carries over only if two things hold. A team's documents have to look like the benchmark's. And the 0.3 points have to survive one harness running every model.

On the hosted API, an output token costs about 1.76 times as much as an input token [22]. Sending back text, position and score with every value makes a response larger than bare JSON. Runtimewire's account does not say whether that metadata is billed as output tokens. A call accepts up to 128,000 tokens of context and returns up to 32,000 [13].

What to watch

  • Independent runs of olmOCR and structured-output accuracy that put interfaze-1-lite and the compared commercial models through one harness.
  • Calibration data, from Interfaze or its users, showing how often high-confidence extractions turn out wrong on real documents.
  • A smaller or quantized build of interfaze-1-lite that runs on GPUs below compute capability 8.9 or with less than 80GB of memory.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories