Build1 publisherNot yet confirmed elsewhere2 min readPublished
GMI Cloud offers a free week of Alibaba's Qwen and Wan models to sell its inference API
GMI Cloud is giving developers seven days of free access to Alibaba's Qwen3.8-Max, Qwen3.8-Flash and Wan3.0 via an OpenAI-compatible API. For a company built on renting GPU capacity, the week tests whether developers will take its inference service as an API.
The Engineer · Build desk

What happened
- Alibaba's Qwen team announced in an October 9 post that the free week was live on GMI Cloud and pointed developers there.
- New users must hold at least $10 in their GMI account; GMI says the money is not spent on the free access and stays available for later use.
- A separate enterprise-credit program, open by application through November 5, offers up to $10,000 per organization.
- GMI's 2024 Series A raised $82 million, $15 million in equity and $67 million in debt, led by Headline Asia with Banpu and Wistron, TechCrunch reported.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams that want Qwen3.8-Flash for latency-sensitive traffic have to measure it on their own workload inside the free window, since the only role labels on offer are GMI's.
- cost The free trial still needs a funded account, so trying the models means getting a payment method onto a GMI account before the first call.
- constraint Until GMI or Alibaba report sign-ups or conversion, outsiders cannot tell whether the move from GPU rental to API sales is producing paying customers.
GMI's first data centers supported Bitcoin mining [6]. The company moved into AI cloud infrastructure after clients and investors asked for GPU capacity, according to RuntimeWire [6]. In a 2025 interview, founder and CEO Alex Yeh described what came next as a sequence [7]:
1. Provide GPUs. 2. Build out the machine-learning stack. 3. Offer API access to customers who do not want to manage the hardware themselves.
Yeh has said customers increasingly want a simple API rather than direct interaction with GPUs, RuntimeWire reported [8].
Integration should be the cheap part of trying GMI. The models sit behind an OpenAI-compatible API [2], and I'd expect most teams to reach them by changing configuration in a client they already run. Evaluation is where the effort goes. GMI pitches Qwen3.8-Max at jobs with long contexts, Qwen3.8-Flash at work where latency matters and Wan3.0 at making video [3]. Those descriptions are GMI's own, and RuntimeWire notes that the promotion does not establish independent performance results [3][12]. For "latency-sensitive" to hold for a given product, that product's prompt lengths, concurrency and output sizes would have to resemble whatever traffic GMI had in mind. Seven days of free use [2] covers one weekly traffic cycle exactly once. A team can run a harness against all three models in that time, provided the harness is written before the clock starts.
The round behind GMI's infrastructure was mostly borrowed money. Debt made up about 82% of the 2024 Series A [16]. Banpu was to provide power, and Wistron planned to co-develop products with GMI [10]. In my view, that capital structure explains the move up the stack better than the campaign page does. Debt-heavy financing rewards keeping GPUs busy, and an API can fill them with traffic from developers who would not first arrange a dedicated GPU cluster [15].
Alibaba gets another route to developers out of the arrangement [14]. RuntimeWire previously reported that Nebius added an Alibaba Qwen model to its hosted inference menu [13]. For organizations the clock is longer: GMI's enterprise-credit applications stay open until 27 days after the date of Qwen's post [17].
What to watch
- GMI or Alibaba publishing sign-up, usage or paid-conversion figures from the free week.
- Independent latency and long-context measurements of Qwen3.8-Max and Qwen3.8-Flash on GMI's endpoint.
- Whether GMI extends the enterprise-credit program past November 5 or changes the $10 balance requirement.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+5
- Incentives70
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Alibaba's Qwen team said in an October 9 post that a week of free access to Qwen3.8-Max, Qwen3.8-Flash and Wan3.0 was live on GMI Cloud; the post points users to GMI.
- [2]
GMI's campaign page advertises seven days of free use and describes the models as available through an OpenAI-compatible API.
- [3]
GMI's lineup presents Qwen3.8-Max for long-context work, Qwen3.8-Flash for latency-sensitive workloads and Wan3.0 for video generation; those capability descriptions come from GMI.
- [4]
New users must have at least $10 in their GMI account balance, according to the campaign terms; GMI says the money is not used for the free model access and remains available for later use on its platform.
- [5]
GMI's offer page provides a separate application path for enterprise credits; GMI says that program runs through November 5 and offers up to $10,000 per organization.
- [6]
GMI started with data centers supporting Bitcoin mining, then shifted toward AI cloud infrastructure after clients and investors asked for GPU capacity.
- [7]
In a 2025 interview, GMI founder and CEO Alex Yeh described the next steps as moving from providing GPUs, to building out the machine-learning stack, and then to offering API access for customers who did not want to manage the hardware themselves.
- [8]
Yeh has said customers increasingly want a simple API rather than direct interaction with GPUs.
- [9]
GMI's 2024 Series A totaled $82 million, comprising $15 million in equity and $67 million in debt, led by Headline Asia with participation from Banpu and Wistron; it was announced in October 2024.
- [10]
In the Series A, Banpu was to provide power, while Wistron planned to co-develop products with GMI.
- [11]
GMI and Alibaba have not reported results from the promotion, including sign-ups, usage or conversion to paid workloads.
- [12]
The promotion does not establish independent performance results for the models.
- [13]
RuntimeWire previously reported that Nebius added an Alibaba Qwen model to its hosted inference menu.
- [14]
The arrangement gives Alibaba's models another distribution route to developers.
- [15]
Developers can test the models without first arranging a dedicated GPU cluster, while GMI gets a chance to introduce its inference service to users who may later need more sustained capacity.
- [16]
Debt was about 82% of GMI's 2024 Series A.
- [17]
GMI's enterprise-credit program closes 27 days after Qwen's October 9 post.
Sources
1 independent publisher whose own reporting we read for this story.
- runtimewire.comGMI Cloud offers developers a free week with three Alibaba models
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- AI inference APIsFollow
- GPU cloudFollow
- Open-Weight Model DistributionFollow