Runtime as featured inForbesRead the article
All open source models
Z.ai logo

GLM-5.3 Flash

Z.ai · #4 in our ranking · Best value

GLM-5.3 Flash is an open-weight mixture-of-experts model from Z.ai, released Aug 2026, with 320B parameters (18B active per token), a 1M-token context window, and the MIT.

GLM-5.3 Flash starts from a new base model with hybrid sparse and linear attention, which cuts long-context serving costs. With 18B active parameters it is cheap to run, and Z.ai reports it approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM-5.3 Flash specs

Context1M tokens
Parameters320B
Provider
Z.ai
Architecture
Mixture-of-experts
Parameters
320B total · 18B active
Context
1M tokens | 786K words
License
MIT Commercial use
Input
Text, Image
Released
Aug 2026

Sourced from the model card on Hugging Face. Last checked 2026-10-01.

Highlights

  • 18B active parameters
  • MIT license
  • Image input and 1M context

Where it fits for payment teams

  • Partner-bank RFIs

    #2 pick. Image input at 18B active parameters and an MIT license, a cost-effective default for high RFI volume.

  • Transaction monitoring

    #2 pick. 18B active parameters with strong agentic skills for the enrichment queries behind each alert.

  • KYC and KYB checks

    #2 pick. Image input, MIT license, and low cost for high onboarding volume.

  • Payments customer support

    #1 pick. Low cost at 18B active parameters with strong tool use for account lookups.

Harnesses in Runtime

Run GLM-5.3 Flash through any of these harnesses on Runtime.

Claude CodeOpenCodeCline

How to run GLM-5.3 Flash

01

Download the weights

Pull zai-org/GLM-5.3-Flash from Hugging Face. Check the MIT before you deploy.

02

Serve it in your cloud

Run it with vLLM or SGLang on your own GPUs, or use a managed provider that hosts the model.

03

Put it to work in Runtime

Point an agent or a single skill at the model through OpenCode, then compare it to your current model with evals.

GLM-5.3 Flash FAQ

What is the context window of GLM-5.3 Flash?

+

GLM-5.3 Flash supports a 1M-token context window (1,048,576 tokens), roughly 786K words of English text.

How many parameters does GLM-5.3 Flash have?

+

GLM-5.3 Flash is a mixture-of-experts model with 320B total parameters, of which 18B are active per token.

Is GLM-5.3 Flash free for commercial use?

+

Yes. GLM-5.3 Flash is released under the MIT license, which is permissive and allows commercial use and fine-tuning.

What can GLM-5.3 Flash take as input?

+

GLM-5.3 Flash accepts text, image input and generates text.

Is GLM-5.3 Flash good for partner-bank RFIs?

+

Yes. GLM-5.3 Flash is our #2 pick for partner-bank RFIs. Image input at 18B active parameters and an MIT license, a cost-effective default for high RFI volume.

Is GLM-5.3 Flash good for transaction monitoring?

+

Yes. GLM-5.3 Flash is our #2 pick for transaction monitoring. 18B active parameters with strong agentic skills for the enrichment queries behind each alert.

Is GLM-5.3 Flash good for KYC and KYB checks?

+

Yes. GLM-5.3 Flash is our #2 pick for KYC and KYB checks. Image input, MIT license, and low cost for high onboarding volume.

Is GLM-5.3 Flash good for payments customer support?

+

Yes. GLM-5.3 Flash is our #1 pick for payments customer support. Low cost at 18B active parameters with strong tool use for account lookups.

Can I run GLM-5.3 Flash in my own cloud with Runtime?

+

Yes. Download the weights from Hugging Face (zai-org/GLM-5.3-Flash) and serve them with vLLM or SGLang in your cloud, or use a managed provider that hosts the model. Runtime agents can then use GLM-5.3 Flash through a harness like OpenCode, with your data staying in your environment.