Product1 distinct publisher3 min readPublished
Nvidia's Apache 2.0 beta seizes the port Ollama or LM Studio was using and forwards each call to whichever machine is free, which makes local throughput a function of how many PCs you own. Hands-on testing found it serving an engine Nvidia does not list.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
At 9pm, a lead agent chops a coding job into eight subagent calls [5]; the second PC in the house is mid-game, and the router passes it over [7]. That behaviour is the thing being bought, but the pitch and the product part company on the details. Nvidia's framing counts households with two or more PCs [3]. The router's own per-request checklist counts something much narrower, because it asks whether the node is online and ready, whether a supported engine is enabled there, whether that exact model is present, how much work the node and engine already have, and what the GPU is currently doing [6].
The "exact model present" test is the one that shapes your weekend. A pool only helps for the models you have already pulled onto more than one machine, so the real price of a second node includes a second copy of the weights. And because PAIR polls nodes and lets them join or leave as they please [8], eligibility is a number that moves through the day rather than capacity you can plan against.
The compatibility trick is small and worth understanding before you install anything. PAIR's proxy claims 11434 for Ollama or 1234 for LM Studio, and the real engine, started after PAIR, moves one port up to 11435 or 1235 [10][1]. The harness keeps the config it already had [4]. The answer may come from a different computer. The Jobs view is where you find out which one [11].
XDA's hands-on went further than the design intends. The writer was in a different country from the hardware, aimed Tailscale at a LAN-first tool [14], and still ended up steering requests to two models on two machines through one endpoint, using an engine PAIR does not support out of the box [15]. The sensible read is that the supported-engine list records what Nvidia tested, not what the proxy will refuse. The same session produced the beta scars you would expect, including a desktop client that trails the terminal one and a hand-written config file that PAIR deleted [13].
So, two axes for your own house. First, whether the pain is queue depth or a model that will not load on your best card. Second, whether the other machines are genuinely idle at the hour you actually work, or occupied by people who live there. This one quadrant, queue depth plus idle peers, is what the router actually serves, and cheaply [12]. Queue depth plus busy peers leaves the router with nothing to hand work to [7]. A ceiling problem is a VRAM purchase in either column [2]. Anyone landing in the good quadrant should expect to live in the terminal client and to keep their own copy of any config they write by hand [13].
Ranked by verification strength, evidence, and original report placement.
As part of its IFA announcements, Nvidia unveiled a free, open-source tool called Personal AI Router (PAIR) that takes every machine on a local network and puts them behind a single address that AI apps already know how to talk to.
PAIR is licensed Apache 2.0 and runs on Windows, Linux and macOS.
Nvidia's pitch for PAIR is that more than half of US households have two or more PCs and that most of that hardware sits idle most of the day.
Nvidia describes PAIR as a virtual inference router rather than a new inference engine: Ollama or LM Studio still runs the model, PAIR only decides which machine gets each request, and the agent keeps talking to the endpoint it already knows so no harness needs rewriting.
PAIR is pitched at subagent fan-out, where one lead agent splits a research or coding job into smaller pieces handed to specialist subagents, so a single request arrives at the inference layer as dozens of independent model calls that can all target the same local engine.
For each request PAIR evaluates whether a paired node is online and ready, whether a supported engine is enabled there, whether that exact model is present, how much work the node and engine already have, and current GPU utilisation.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
build
Qwen 3.8's Apache-licensed 27B is the one you can actually own, and its KV cache is why1 distinct publisher
build
One DGX Spark, four Macs, and an attempt to turn token billing into a capital purchase1 distinct publisher
build
The flash_attn error in llama.cpp is a layout constraint, and it decides your context window1 distinct publisher
build
A 27B model reportedly beat a license check in 30 minutes. Nobody has seen the binary.1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Verified where it could be typed, vendor-fed where it matters
The reviewer at XDA Developers actually installed this and watched two machines answer one endpoint, which beats a press release — and the checkable specifics, the port shift and the config file that vanished, read like someone who was there. What nobody has tested is the part that decides everything: the readiness, model-presence and GPU-utilisation weighing is Nvidia's description of its own scheduler, repeated once, with no second installation anywhere to hold it against.
Day one, one tester, two computers
Everything on the record is a beta that just shipped and a single hands-on — and that hands-on was run the wrong way round, over Tailscale from another country, because the reviewer's machines were elsewhere. Nvidia's own number for how far this scales, eighteen GPUs in one cluster, is Nvidia's; no user has reported a pool remotely that size. There is no deployment story yet, only an availability story.
The premise promises more than the router delivers
Nvidia's framing invites you to add your household's idle silicon together; the tool hands each request whole to one machine, so the biggest model you can run is still whatever your best box holds and extra PCs buy concurrency only. What keeps the gap narrow is that Nvidia says so itself, repeatedly, and XDA Developers passes the correction on rather than burying it. The overreach lives in the pitch, not in the hands-on.
Free software that makes the second GPU worth owning
Nvidia is publishing, at no charge and under Apache 2.0, precisely the missing piece that turns a spare gaming PC into inference capacity — and the pitch conveniently counts how many homes already have one. It is also the piece that never lets small cards add up to a big one, which keeps the reason to buy upmarket intact. Cutting the other way: the reviewer notes AMD and Intel hardware works fine in practice, and the reporting rests on review access Nvidia granted.
Credible first look, single vantage point
We are confident about what the reviewer touched and much less about what he was told. One publisher, a beta days old, an unsupported network path standing in for the intended LAN, and a scheduler whose behaviour is described rather than measured — enough to report accurately on what PAIR is, not enough to say how it holds up when four machines are contending for the same job.