Z.ai describes GLM-5.3-Flash as the first natively multimodal model in the GLM-5 family. Its announcement lists 320 billion total parameters and about 18 billion active parameters per inference step, with training across text and visual data. The aim is stronger coding, tool use, visual understanding, and professional work with controlled inference cost. Read the AutoClaw technical overview.

Z.ai / AutoClaw official mark
Z.ai / AutoClaw official mark; click to open the announcement.

What the architecture claims mean

Total parameters describe the overall model; active parameters describe the portion used for a request. GLM-5.3-Flash combines sparse attention for retrieving relevant context with linear attention for local dependencies. Z.ai also describes long-context index compression and publishes comparisons for attention compute and KV-cache use. These are architectural metrics, not end-to-end latency or total cost; results depend on framework, quantization, batching, and hardware.

What native vision enables

The announcement says the model handles text, images, video, and files, including charts, document layout, and interface state. Test it on screenshots or chart-heavy documents before asking it to act on what it sees. Vision support does not guarantee perfect OCR, detail recognition, or temporal reasoning. Verify amounts, medical content, and critical interface states against the original.

Steps for evaluating a self-hosted setup

  1. Check the official model page for downloadable weights, license, precision, and hardware requirements.
  2. Build an isolated test with a listed framework such as SGLang or vLLM; record versions, quantization, and context length.
  3. Test code, charts, screenshots, long documents, and tool calls. Measure answer quality, memory, throughput, and latency separately.
  4. If used in an agent, test screen reading, tool selection, and permissions independently rather than treating a benchmark as launch acceptance.

18 billion active parameters does not mean a device needs memory for only 18 billion weights. Inference still stores or distributes the full weights and needs room for KV cache, visual tokens, and concurrent requests. Z.ai says the weights use an MIT license; review the current model repository, weights, and third-party terms before commercial use.