Build1 publisher2 min readPublished
A resident Whisper model plus embedding model together burned 2 euros of GPU electricity across 30 days
An RTX 3090 holding transcription and embedding models in VRAM averaged 25 W over the month at a cost of 2 euros, while the same post puts hardware amortisation at about 25 euros a month, twelve times the power bill.
The Engineer · Build desk

What happened
- A power meter on an RTX 3090 recorded EUR 2.00 of electricity over 30 days, measured at the card, with a transcription service and an embedding model resident around the clock.
- The resident WhisperX service held an average of 7.9 GB of VRAM with a 10.3 GB peak, waiting for audio from the author's own tools.
- The nomic-embed-text model, permanently loaded under Ollama at 308 MB of VRAM, accounted for 23 cents of the month's electricity.
- Spikes reached about 120 W for seconds to minutes at a time, with one 2-hour bucket averaging 121 W after a batch of long audio files arrived.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Whoever is deciding whether to buy or keep the card is making a bet on purchase price and resale value, because the power side of the question is settled at a couple of euros a month.
- contradiction The same post carries both the claim that local runs at under a tenth of API spend and the admission that hardware is the real cost, and those two sentences point at opposite purchasing decisions.
- cost At 23 cents a month for a permanently resident embedding model, writing lazy-load and unload logic to reclaim idle power costs more engineering time than it can ever return.
An average of 25 W held for 720 hours is 18 kWh. The post reports EUR 2.00 for the month, so the implied price is about 0.11 EUR per kWh. Its hypothetical sustained-load case, 187 kWh for about EUR 23, implies about 0.12. The two agree closely enough to trust that one tariff produced both. That tariff is quoted in BGN, 0.30 per kWh by day and 0.18 between 22:00 and 06:00, and the post does not state the conversion it used to report euros.
One pair of figures does not close. The 2-hour power history gives a baseline of about 35 W around the clock, with everything above it recorded as spikes, while the stated average for the Whisper service is 22 W. A 35 W floor with upward excursions cannot average 22 W. Take 35 W as the real floor and the month is 25.2 kWh, about EUR 2.80 at the implied price, some 80 cents above the published total.
Hardware dominates the monthly cost in the post's own accounting: a used 3090 at EUR 800 to 1,200, or about EUR 25 a month over three years, twelve times the power bill. "The electricity is not the cost. The hardware is the cost," the author wrote. The claim that the local stack runs at less than a tenth of the API equivalent holds only while the card is treated as sunk. Add the EUR 25 and the stack is about EUR 27 a month, against the post's own cloud estimate of EUR 14 to 29. At the top of the hardware range, EUR 1,200 over 36 months is EUR 33, and with electricity that is EUR 35.
For the EUR 2.00 to describe someone else's box, the duty cycle has to match. Average GPU utilisation here is 0%, the work arrives as seconds of transcription at a time, and a card held at its 260 W power limit is the separate case the post prices at 187 kWh and about EUR 23 a month. The bill also scales directly with the local rate, so a reader paying three times 0.11 EUR per kWh pays three times EUR 2.00. And the figure is the GPU alone, not the whole machine, nor the second box, where two RTX 3090s run Qwen3.8-27B under vLLM on demand.
What to watch
- Used 3090 resale prices: the EUR 800 to 1,200 spread is worth about EUR 11 a month in amortisation, more than five times the whole electricity bill.
- A month dominated by queued batch transcription. That workload would test whether the average holds near the idle floor.
- An itemised basis for the EUR 14 to 29 cloud comparison, given in the post only as a range.