Build2 publishers3 min readPublished
Kimi K3 on its cheapest host undercuts Fireworks' Ember-1 despite a 23% cut in reasoning tokens
Fireworks' Ember-1 used 23% fewer reasoning tokens than Kimi K3 in The New Stack's tests, yet Kimi on the cheapest host would cost $1.96 to Ember's $2.48. Ember beats Fireworks' own Kimi rate and loses at the cheapest, so buyers have to price the host before the model.
The Engineer · Build desk

What happened
- Fireworks launched Ember-1 on September 23 as a research preview built on Moonshot's open-weight Kimi K3.
- Fireworks bills Ember-1 at $3 per million input tokens and $15 per million output, the rate it charges for Kimi K3, while other hosts sell Kimi for as little as $1 and $9.
- Across three tests of rising difficulty, run five times each, Kimi K3 was right on all 15 runs and Ember-1 on 14.
- With both models served by Fireworks, Ember-1 finished each set of tests 3.4 times faster than Kimi K3.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability Agent pipelines that wait on wall-clock time can get a large latency cut from Ember on Fireworks now, an advantage not yet tested against the cheaper Kimi hosts.
- exposure A buyer who adopts Ember on Fireworks' benchmark and customer figures carries the risk that its own tasks resemble neither, since the evaluations are vendor-run and the customers unnamed.
- precedent Fireworks calls Ember-1 the first of a series of specialized models from its Serverless Training service, so more house-trained derivatives will likely arrive at its own rates for the base model.
Reasoning tokens bill as output [4]. In The New Stack's runs, output was nearly the whole bill [2]. Kimi K3 cost $3.26 for the full suite on Fireworks and would have cost $1.96 at the cheapest listed price [16][17]. The cheaper bill is 0.60 of the Fireworks one [1]. The output prices, $9 and $15, sit in the same ratio, while the input prices differ by a factor of three [2]. Moving Kimi to the cheapest host saves $1.30 on the suite. Moving to Ember on Fireworks saves $0.78 [5].
A bill made of output tokens has a clean break-even. At $15 against $9, Ember-1 has to emit 40% fewer output tokens than Kimi just to tie the cheaper host [3]. Fireworks advertises roughly 40% fewer tokens as Ember's headline result [2], a coincidence I would not put in a budget. The measured reasoning-token cuts were 18% on the logic puzzles, 16% on deploy scheduling and 35.5% on the probability questions [7][8][6].
The tester built the suite on a stated hypothesis: if Ember cut reasoning it needed, accuracy would slip first on the hardest problems [13]. The probability test produced both Ember's largest cut and its only error [6][9]. On the fifth run it answered the first question 0.94619 instead of 0.94629, and the slip carried into two other answers [9]. The tester wrote that the miss "doesn't seem too significant to me since it was one out of 15" [12].
According to runtimewire, Fireworks found that turning down Kimi K3's reasoning-effort setting weakened its results [20]. Its researchers instead post-trained a K3-based model to shorten reasoning while keeping the parts that help it recover from mistakes and adapt to feedback [20]. Training the shorter trace into the model, once the setting had failed, is the sound choice. The work ran to more than 50 training experiments and over 200 evaluations [20]. Fireworks said the model "learned to cut unnecessary reasoning while keeping the thinking that matters" [3].
Fireworks' own evaluations put Ember-1 at 92.2% on SWE-bench Verified against 93.2% for K3 at maximum effort, and at 82.0% against 80.9% on Terminal-Bench 2.1 [21]. The benchmark sets run from 50 to 500 examples [21]. In a live test with one of two customers, Fireworks reported 39% fewer total tokens and 71.3% fewer reasoning tokens, at a task score of 0.753 against 0.751 [22]. If that customer's other tokens held flat, reasoning was about 55% of its baseline [8]. A team whose traffic is less reasoning-heavy would get a smaller total cut from the same model [8].
Token savings do not account for Ember's speed [15][14]. On the logic puzzles, Ember produced about 60 reasoning tokens per second of wall time and Kimi about 22 [7]. Fireworks was generating Ember's tokens faster [7]. Dmytro Dzhulgakov, a Fireworks co-founder, wrote in a September 27 post on X that Ember-1 was "hot on HN" and said it was 40% faster and cheaper, according to runtimewire [18]. The outlet reports that the launch materials back the token claim but do not establish a general 40% speed gain [19].
The cost case is strongest for multi-turn agents, according to runtimewire. Fireworks says those agents may send earlier reasoning back to the model on each turn, so a long trace is processed repeatedly [23]. Its listed rate for cached input is $0.30 per million tokens [24]. In my view the comparison that holds is cost per correct task on a team's hardest workload, priced at every host it would buy from. On The New Stack's suite, Ember cost about 18 cents per correct run and Kimi on the cheapest host about 13 [9].
What to watch
- Speed and cost runs of Kimi K3 on the $1/$9 hosts, to show whether Ember's lead on time holds outside Fireworks' serving stack.
- Independent multi-turn agent tests that resend reasoning each turn, the setting where Fireworks says shorter traces cut cumulative token use.
- Any change to Ember-1's $3/$15 pricing now that its model page lists it as ready for serverless use after the research-preview window.