Skip to content

Invest1 publisher3 min readPublished

Touchmark opens a forwards market for tokens because finance cannot forecast them

A two-person YC company is selling prepaid inference contracts at up to 30 percent off. The pitch works because enterprise AI spend doubled to about $1.2M per organization and 78 percent of IT leaders got surprised.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Touchmark opens a forwards market for tokens because finance cannot forecast them
Photo: techfundingnews.com

What happened

  • Touchmark is a two-person company, part of Y Combinator's Summer 2026 batch.
  • Touchmark is a forwards market that opens to commercial buyers and sellers this week.
  • On Touchmark, a billion tokens on a named model is a single prepaid contract with a delivery window attached, drawn down through an API when the window arrives.
  • The buyer pays in full up front and consumes the capacity through an API like any other endpoint.
  • A company that knows it will spend heavily on AI in September can pay for that month now and get its tokens at a discount of up to 30 percent, or more for larger orders via direct quotes.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

Touchmark, a two-person company in Y Combinator's Summer 2026 batch, opens a forwards market for AI inference to commercial buyers and sellers this week, in which a billion tokens on a named model is a single prepaid contract with a delivery window attached [1][2][3]. Buyers pay in full up front and draw the capacity down through an API when the window arrives, at a discount of up to 30 percent, or more on larger orders quoted directly [3][4][5].

The interesting part is not the market design. It is the budgeting hole the pitch is aimed at: enterprise AI spending roughly doubled year over year in 2026, to an average of about $1.2 million per organization, and 78 percent of IT leaders reported charges they had not budgeted for [6][7]. That doubling implies an average of roughly $600,000 the prior year [8]. Agents made the forecasting problem worse, because a single autonomous workflow can consume many times the tokens of a one-shot query, quietly, over hours [9]. Finance teams that are comfortable forecasting seat licenses have spent the year discovering they cannot forecast this [10].

The product is therefore sold to the buyer's finance function, not its platform team. A buyer selects a model, a token volume, a preferred provider and a delivery window, and the earlier the purchase, the deeper the discount [11][12]. Alongside the listed contracts sits a request-for-quote board, live at launch, where a buyer posts volume, delivery window and throughput and latency floors, denominated in tokens, dedicated GPUs or reserved throughput, and providers compete to quote [13][14].

Supply skews toward open-weight families: Kimi, GLM, Qwen [15]. That follows demand rather than shaping it. OpenRouter's routing data showed Chinese open-weight models crossing into a majority of the tokens it processes by the middle of this year, most priced well beneath the U.S. frontier [16]. The commoditization thesis lives or dies on that substitutability; a forward on a single proprietary frontier endpoint is a bilateral supply deal wearing a ticker.

On the sell side, the first provider is Wafer, described in the announcement as one of the leading inference providers by throughput [17]. The appeal there is timing: cash today against capacity that will not be served for months, in a business where capacity sits underused one month and runs short the next, and where the industry's current answer is a series of private one-off deals [18][19][20]. Co-founder Roman Yanushevskyi frames the same gap from the provider's side, arguing that thin demand wastes GPU hours while excess demand churns customers [21]. Chief executive Ilia Bolgov says pricing for reserved capacity is bilateral and buyers have very little visibility into a fair forward price, and that the market needs price discovery, forward purchasing and the ability to resell unused capacity [22][23].

Resale is the part that is still an ambition rather than a feature, and the source material does not describe how a contract settles if a provider fails to deliver, whether prepayments are escrowed, or who carries the risk that a named model is deprecated inside the delivery window. A buyer paying in full is extending unsecured credit against a two-person intermediary's counterparty list; the discount is compensation for that, plus the time value of money. Bolgov trained in mathematics at Cardiff and Imperial and worked in product at Revolut's wealth and trading division; Yanushevskyi took gold at the 2022 International Olympiad in Informatics and interned twice on Citadel Securities' options team [24][25]. Neither resume underwrites delivery.

Watch three things. Whether any secondary transfer of unused capacity actually clears, since that is what separates a forwards market from a prepaid voucher scheme. Whether any closed-weight frontier lab lists capacity, or whether the book stays open-weight. And whether the discount survives contact with falling spot prices: 30 percent off today is only a saving if the spot price at delivery has not dropped more than 30 percent [26].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories