Skip to content

Build2 publishers2 min readPublished

Bedrock serves Z.ai's 753B-parameter GLM 5.3 to eligible enterprise customers

Amazon Bedrock now offers Z.ai's 753B-parameter mixture-of-experts GLM 5.3 as a managed API, available only to eligible enterprise customers. Coding-agent teams can call a model that size without running inference hardware of their own.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Bedrock serves Z.ai's 753B-parameter GLM 5.3 to eligible enterprise customers
Generated illustration

What happened

  • Requests go to a source Region the caller picks, using the us.zai.glm-5.3 or global.zai.glm-5.3 cross-Region profile, and Bedrock routes them for processing.
  • Prompt caching is implicit by default, and AWS recommends explicit cache controls, available on the Responses and Chat Completions APIs.
  • Z.ai reports a 50% improvement over GLM 5.2 on its own internal coding benchmark.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Smaller teams cannot build on GLM 5.3 through Bedrock until AWS confirms they qualify as eligible enterprise customers.
  • decision Picking the us. or global. profile ID sets how widely Bedrock may route each request, so the choice belongs with whoever owns data-residency policy.
  • cost Moving production agents from GLM 5 to GLM 5.3 needs an in-house evaluation budget, since Z.ai's published gains are measured against GLM 5.2 on revised tests.

An IAM policy for GLM 5.3 has to cover two resources: the base model and the target inference profile [17]. The actions listed in the prerequisites are bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream and bedrock:CallWithBearerToken [17].

Caching matters most for agent loops. An agent resends its system prompt and repository context on every turn, and AWS says caching that repeated input reduces both latency and input cost [5]. For a coding agent with a large, stable repository prefix, I would set the explicit controls AWS recommends, so the team decides in code which prefix gets cached [4]. The part of this release I'd credit most is API parity. AWS says the OpenAI-compatible Responses and Chat Completions APIs now match Invoke and Converse more closely than they did for GLM 5 [16].

On public suites, Z.ai claims competitive performance on DeepSWE, Terminal Bench 3.0 and FrontierSWE [9]. The post says Z.ai's benchmark tests were themselves updated after the GLM 5.1 announcement because the improvements were so large [11]. Replacing the test because the model outgrew it is an unusual way to lose a baseline. For the 50% figure to carry over, a team's repositories and task lengths would have to resemble Z.ai's internal benchmark [10]. GLM 5 has been on Bedrock since earlier this year [12], so teams already running it can compare the two models on their own agent traces.

The CyberGym number is also Z.ai's own. It measured a leading 84.5 at release [13]. AWS's worked example runs an authorized security test of the reader's own application with Strix, an open-source AI penetration testing agent [14]. That demo needs Docker and Strix installed with its bedrock extra [14].

Service tiers set price against latency. Flex optimizes cost for less time-sensitive work, Priority costs more for latency-critical requests, and Standard is the default balance [15]. The post does not say which customers count as eligible enterprise customers, and it does not include prices for any tier [2].

What to watch

  • AWS stating what makes a customer eligible for GLM 5.3, or widening access beyond eligible enterprise customers.
  • Published per-token prices for the Flex, Standard and Priority tiers, including any separate rate for cached input.
  • DeepSWE, Terminal Bench 3.0 or CyberGym results for GLM 5.3 from someone other than Z.ai.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories