Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

GMI Cloud offers a free week of Alibaba's Qwen and Wan models to sell its inference API

GMI Cloud is giving developers seven days of free access to Alibaba's Qwen3.8-Max, Qwen3.8-Flash and Wan3.0 via an OpenAI-compatible API. For a company built on renting GPU capacity, the week tests whether developers will take its inference service as an API.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying GMI Cloud offers a free week of Alibaba's Qwen and Wan models to sell its inference API
Generated illustration

What happened

  • Alibaba's Qwen team announced in an October 9 post that the free week was live on GMI Cloud and pointed developers there.
  • New users must hold at least $10 in their GMI account; GMI says the money is not spent on the free access and stays available for later use.
  • A separate enterprise-credit program, open by application through November 5, offers up to $10,000 per organization.
  • GMI's 2024 Series A raised $82 million, $15 million in equity and $67 million in debt, led by Headline Asia with Banpu and Wistron, TechCrunch reported.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams that want Qwen3.8-Flash for latency-sensitive traffic have to measure it on their own workload inside the free window, since the only role labels on offer are GMI's.
  • cost The free trial still needs a funded account, so trying the models means getting a payment method onto a GMI account before the first call.
  • constraint Until GMI or Alibaba report sign-ups or conversion, outsiders cannot tell whether the move from GPU rental to API sales is producing paying customers.

GMI's first data centers supported Bitcoin mining [6]. The company moved into AI cloud infrastructure after clients and investors asked for GPU capacity, according to RuntimeWire [6]. In a 2025 interview, founder and CEO Alex Yeh described what came next as a sequence [7]:

1. Provide GPUs. 2. Build out the machine-learning stack. 3. Offer API access to customers who do not want to manage the hardware themselves.

Yeh has said customers increasingly want a simple API rather than direct interaction with GPUs, RuntimeWire reported [8].

Integration should be the cheap part of trying GMI. The models sit behind an OpenAI-compatible API [2], and I'd expect most teams to reach them by changing configuration in a client they already run. Evaluation is where the effort goes. GMI pitches Qwen3.8-Max at jobs with long contexts, Qwen3.8-Flash at work where latency matters and Wan3.0 at making video [3]. Those descriptions are GMI's own, and RuntimeWire notes that the promotion does not establish independent performance results [3][12]. For "latency-sensitive" to hold for a given product, that product's prompt lengths, concurrency and output sizes would have to resemble whatever traffic GMI had in mind. Seven days of free use [2] covers one weekly traffic cycle exactly once. A team can run a harness against all three models in that time, provided the harness is written before the clock starts.

The round behind GMI's infrastructure was mostly borrowed money. Debt made up about 82% of the 2024 Series A [16]. Banpu was to provide power, and Wistron planned to co-develop products with GMI [10]. In my view, that capital structure explains the move up the stack better than the campaign page does. Debt-heavy financing rewards keeping GPUs busy, and an API can fill them with traffic from developers who would not first arrange a dedicated GPU cluster [15].

Alibaba gets another route to developers out of the arrangement [14]. RuntimeWire previously reported that Nebius added an Alibaba Qwen model to its hosted inference menu [13]. For organizations the clock is longer: GMI's enterprise-credit applications stay open until 27 days after the date of Qwen's post [17].

What to watch

  • GMI or Alibaba publishing sign-up, usage or paid-conversion figures from the free week.
  • Independent latency and long-context measurements of Qwen3.8-Max and Qwen3.8-Flash on GMI's endpoint.
  • Whether GMI extends the enterprise-credit program past November 5 or changes the $10 balance requirement.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence50
Adoption
Insufficient
Hype gap+5
Incentives70
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Alibaba's Qwen team said in an October 9 post that a week of free access to Qwen3.8-Max, Qwen3.8-Flash and Wan3.0 was live on GMI Cloud; the post points users to GMI.

    ReportedSupportedSource: RuntimeWire, citing Qwen's October 9 postView cited source
  2. [2]

    GMI's campaign page advertises seven days of free use and describes the models as available through an OpenAI-compatible API.

    ReportedSupportedSource: RuntimeWire, describing GMI's campaign pageView cited source
  3. [3]

    GMI's lineup presents Qwen3.8-Max for long-context work, Qwen3.8-Flash for latency-sensitive workloads and Wan3.0 for video generation; those capability descriptions come from GMI.

    ReportedSupportedSource: RuntimeWireView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. runtimewire.com

    1 article · October 8, 2026

    GMI Cloud offers developers a free week with three Alibaba models

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories