Skip to content

Build6 publishers3 min readPublished Updated

Meta's real announcement is the split: 30B on your GPU, everything else behind the API

Muse Glimmer ships as Apache 2.0 weights sized for a 24GB card. Muse Spark 1.2 stays on Muse Code and the Meta Model API. Plan capacity for two tiers, not one.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Meta's real announcement is the split: 30B on your GPU, everything else behind the API
Photo: abcnews.com

What happened

  • Meta released Muse Glimmer on August 10 and released the model weights on Hugging Face under an Apache 2.0 license.
  • Muse Glimmer is a roughly 30-billion-parameter open-weight model.
  • Zuckerberg paired the release with an August 10 manifesto titled "The Future is for Everyone", described as a 6,500-word argument for distributing superintelligence as widely as possible.
  • Meta's research brief identifies the strategic split as local, downloadable weights versus controlled product and API access, a structure that puts the access thesis into practice while preserving Meta's control over its compute-intensive models.
  • Meta says quantization reduces the language model to less than 20 GB, leaving enough memory for image processing, working context and a speculative-decoding model within a 24 GB or 32 GB memory envelope.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Meta released Muse Glimmer on August 10, a roughly 30-billion-parameter agent model published under an Apache 2.0 license on Hugging Face, and paired it with a 6,500-word Mark Zuckerberg essay titled "The Future is for Everyone" arguing for widely distributed personal AI [1][2][3]. The essay will absorb the commentary; the thing that changes a capacity plan is that Meta now operates two tiers with a deliberate line between them, downloadable local weights on one side and metered hosted access to its compute-intensive models on the other [4].

The local tier is defined by a memory budget. Meta says quantization brings the language model under 20GB, leaving room for image processing, working context and a speculative-decoding model inside a 24GB or 32GB envelope [5]. That means roughly 4GB of non-weight headroom on a 24GB card and roughly 12GB on a 32GB card [6]. Unsloth says a quantized build runs in about 18GB on a single GPU, and shipped a GGUF download, a local guide and support in its open training stack [7]. Latent.space describes the result plainly: it fits on a single RTX 3090 [8]. Meta published its own quantized build and developer documentation [9][10], and says integrations are planned for llama.cpp, MLX, Ollama and other local runtimes [11].

Capability-wise, Glimmer is dense, includes a perception encoder, and is tuned for tool calls, code generation and debugging, recovery from failed tool calls and multi-step workflows, with training data from more than 100 languages [12][13]. It is distilled from Muse Spark's outputs, which is worth noting: the local tier is downstream of the hosted one [14]. Meta's own evaluation passage compares Glimmer with Gemma4-31B and Qwen3.6-27B [15]. That is vendor-produced evidence and still needs independent replication [16]. NVIDIA, meanwhile, states a 120K-plus context window and claims more than 20,000 tokens per second on a single Blackwell Ultra GPU, and published deployment recipes for GeForce RTX 5090, DGX Spark/Station, Jetson, SGLang, vLLM and NIM containers [17][18].

The hosted tier is where the harder capability sits. Meta introduced Muse Spark on April 8 as the first model from Meta Superintelligence Labs, built for its own products and later used to add planning, research and calendar connections to Meta AI [19]. Spark 1.2 is a coding-focused update covering code generation, complex debugging, codebase understanding and end-to-end developer workflows, available through Muse Code and the Meta Model API [20]. Axios reported on August 10 that Meta planned to open the Spark 1.2 weights in the coming weeks, and CNBC reported the same plan [21][22].

Read the manifesto as a pricing document and it is consistent with the architecture: free or affordable access for billions, private agent modes, paid auctions for scarce compute, and a $1B fund for communities hosting Meta data centers [23]. Auctions are what you build when compute is rationed, and rationing is the reason the strong models stay on Meta's side of the wire. Aaron Scher has challenged the essay point by point, including its treatment of bioweapon offense versus defense [24]. Zuckerberg controls a majority of Meta's voting power according to the company's latest annual filing, so the strategy is unlikely to be moderated internally [25].

What to watch: whether the Spark 1.2 weights actually land, and on what license; whether independent evaluations reproduce Meta's size-class claims against Gemma4-31B and Qwen3.6-27B [15][16]; whether the llama.cpp, MLX and Ollama integrations arrive with usable throughput rather than just loading [11]; and whether an auction price for hosted compute ever appears [23]. Until then, budget for two lines: 24GB to 32GB of local VRAM per agent seat, and API spend for anything Glimmer cannot finish [5].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories