Skip to content

Build1 publisher3 min readPublished

Interruptible GPU capacity bills at 18 percent of on-demand on one operator's own account

A dev.to post prices a $3.0421 single-card GPU instance at 55 to 70 cents an hour on interruptible capacity. The same account's interruptible GPU quota is 64 vCPUs in every region checked, and a partial approval capped it there.

The Engineer · Build desk

Illustration accompanying Interruptible GPU capacity bills at 18 percent of on-demand on one operator's own account

What happened

  • A dev.to post puts one single-card GPU instance at $3.0421 an hour on demand and 55 to 70 cents an hour on interruptible capacity, taken from the author's own bills rather than a vendor page.
  • He asked for a 128-vCPU interruptible GPU quota, got a same-day partial approval at 64, and was told larger increases have left self-service and now go through an account team.
  • The separately set on-demand GPU quota differed widely across the same account: 64 vCPUs in one region, 16 in another, and 8 in a third.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Price does not unlock the next rung up: two of his 48-vCPU two-card instances would need 96 vCPUs against 64 approved, so capacity arrives only in whole shapes.
  • decision Teams with a person waiting on a session now have to price whether a roughly 80 percent cut covers building the controller and fallback ladder that hide reclaims.
  • contradiction The habit of dismissing AWS's "up to 90 percent" line as marketing comes from CPU fleets at 55 to 75 percent off, and on this GPU family the author's bills land near the top of the range instead.

Before anyone budgets against 18 percent of on-demand, look at what the evidence is: one account, one instance family, and a survey taken across three regions on a single afternoon [1][5]. For the number to transfer, your family has to be one of the GPU families where the discount runs near 80 percent; the author reports 55 to 75 percent for CPU workhorse families [4][23]. And you have to be in the right zone, because on-demand pricing was identical in every region he checked while interruptible pricing was not: a 27 percent spread between availability zones in one region, 1.2 percent in another [6].

The cheapest zone was also the one AWS itself rated lowest for capacity, and it dropped him; four cents an hour of savings bought an outage [8]. "The price you do not pay is not free. It is quoted somewhere else, in interruptions," he wrote [9]. Four cents is small against the spread it sits inside: 27 percent of the 62.5-cent midpoint of his observed band is about 17 cents an hour, four times the saving he chased [10].

The two-minute reclaim notice is the part most teams already design for, and he calls it the least interesting of the three problems, because it is the one you can engineer around [7]. The binding one is quota. His account is allowed 64 vCPUs of interruptible GPU capacity, and it was 64 in every region he checked, so moving does not help [11]. He asked for 128, got a same-day partial approval at 64, then learned that increases past that threshold have been removed from self-service and now require a review reachable only through an account team [12].

His two-card instance is 48 vCPUs [13]. Two of them would be 96, which is 32 vCPUs over the approved quota, about 50 percent more than he has [14]. One instance leaves 16 vCPUs stranded, well short of a second [14]. "I do not scale; I choose a shape and fit inside it," he wrote [15].

The fallback is quoted separately. On-demand GPU quota on the same account was 64 vCPUs in one region, 16 in another, and 8 in a third [16]. A 48-vCPU shape fits in none but the first, so in two of the three regions the on-demand escape hatch cannot host the machine it is meant to replace [17].

The design follows from that. A controller holds a target number of healthy slots and never names a machine or a zone; which zone he is in is an outcome of that, and it changes with every replacement [20]. Under it sits a ladder of shapes tried in order, starting with two cards on one machine [22]. "Four pieces, none of them clever, all of them necessary," he wrote [21]. For batch-shaped work the entry cost is a retry, against a price cut of roughly 4.3x to 5.5x on this family [18][2]. For a live session, the 80 percent has to fund the controller and the ladder, and it cannot buy a larger quota at any price [19][12].

What to watch

  • Whether AWS restores self-service quota increases above 64 vCPUs for interruptible GPU instances, or keeps them behind an account team.
  • A second account, in a different GPU family and region, reproducing the 77 to 82 percent discount on its own bills.
  • Whether the 27 percent intra-region spread between availability zones holds from one week to the next.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories