FLAGSHIP OPEN-WEIGHT
BILINGUAL REASONING.
GLM is Z.ai’s open-weight family of large language models, led by GLM-5.2 — a 753B-parameter multimodal flagship with 1M-token context, FP8 inference, native tool calling, code generation, explicit reasoning, and vision inputs. Procure GLM inference through Compute Exchange Token Forwards or Reserved GPU Rental.
GLM-5 AND BEYOND
The GLM-5 line is the current flagship; GLM-4.7-Flash covers latency-sensitive workloads; GLM-OCR specializes in document and vision extraction. All ship under MIT license as open-weight releases from Z.ai.
GLM-5.2
FLAGSHIP · CURRENTZ.ai's latest flagship multimodal model — strong bilingual (Chinese-English) reasoning, long-context understanding (up to 1M tokens), vision inputs, advanced tool use, code generation, and agent-oriented behavior.
GLM-5.1
PREVIOUS FLAGSHIPPrevious-generation flagship in the GLM-5 line — strong bilingual reasoning, long-horizon agentic tasks (sustains thousands of tool calls), and SWE-Bench Pro state-of-the-art code generation.
GLM-4.7-Flash
FAST / LIGHTWEIGHT MOE30B-parameter MoE with 3B active per token — preserved thinking mode for multi-turn agentic tasks, with speculative decoding and multi-token prediction for low-latency, high-throughput inference.
GLM-OCR
VISION / OCR SPECIALISTCogViT visual encoder + GLM-0.5B language decoder for OCR, document parsing, formula and table recognition. #1 on OmniDocBench V1.5 (94.62). *PP-DocLayoutV3 sub-component under Apache 2.0.
WHAT GLM DOES WELL
BILINGUAL REASONING
Strong Chinese-English reasoning across long-form text, code, and structured tasks. Competitive with Western flagships on English benchmarks and class-leading on Chinese.
1M LONG CONTEXT
Up to 1M tokens on GLM-5.2 (with sparse-attention IndexShare reducing per-token FLOPs ~2.9× at full context) — absorb full legal filings, codebases, or research corpora without chunking.
NATIVE TOOL USE
First-class function calling, structured outputs, and Responses API support. Agent-oriented architecture handles multi-step plans and tool composition.
EXPLICIT REASONING
Reasoning mode surfaces chain-of-thought scratchpads for complex math, code, and analytical tasks. Tunable depth at the API boundary.
MULTIMODAL INPUTS
GLM-5.2 accepts image and text inputs natively. Pair with the specialized GLM-OCR (0.9B, CogViT + GLM-0.5B; #1 on OmniDocBench V1.5) for high-volume document extraction pipelines.
FP8 EFFICIENCY
Native FP8 quantization keeps cost-per-token competitive and inference fast on H100-class hardware across the open-weight provider network.
REPRESENTATIVE WORKLOADS
BILINGUAL ENTREPRISE ASSISTANTS
Customer-facing or internal assistants serving Chinese-English markets with consistent reasoning quality across both languages.
LONG DOCUMENT ANALYSIS
Legal filings, financial disclosures, technical specifications — 432K context absorbs full documents without retrieval chunking.
AGENTIC WORKFLOWS
Tool-calling backbone for multi-step agents — research loops, code generation pipelines, structured action sequences.
RAG WITH REDUCED RETRIEVAL
Long-context tolerance lets you pack more context per query and reduce the brittleness of retrieval recall.
MULTIMODAL DOCUMENT PIPELINES
GLM-OCR for visual extraction → GLM-5.2 for reasoning over extracted content. End-to-end open-weight document understanding.
OPEN-WEIGHT PRODUCTION INFERENCE
MIT-licensed alternative to closed flagships, with deployable weights for sovereign and on-prem buyers.
HOW TO ACCESS GLM
Two procurement paths through Compute Exchange. Choose by who you want operating the model.
COMMITTED INFERENCE
Lock GLM inference capacity in advance, denominated in Standardized Token Units. Provider operates the model; you tap tokens against a committed balance over terms up to six months.
- · Provider operates and scales GLM endpoint
- · Per-STU rate locked at commitment
- · Realtime, batch, or mixed latency
- · Quotes against the published STU index
RUN YOUR OWN
Reserve H100-class capacity from the neocloud network and deploy GLM weights yourself. Full operational control — sovereign data, custom serving stack, fine-tuned variants.
- · MIT-licensed weights — deploy anywhere
- · Custom serving stack (vLLM, SGLang, TensorRT-LLM)
- · Sovereign / on-prem / air-gapped deployments
- · Terms from 1 month to 24+ months
GLM, EXPLAINED
Who builds GLM?+
What is the difference between GLM-5.2 and GLM-5.1?+
Is GLM open-weight?+
How does GLM compare to Western flagship models?+
How do I procure GLM inference through Compute Exchange?+
How does GLM map to the STU methodology?+
PROCURE GLM CAPACITY.
Submit a commitment request and Compute Exchange returns GLM inference quotes across the verified open-model provider network.
GLM is an open-weight model family developed by Z.ai (Zhipu AI), released under MIT license. Compute Exchange facilitates quotes from verified third-party inference providers serving the GLM family and does not operate inference infrastructure or guarantee model availability, performance, or SLA. All commitment terms — including pricing, tap, rollover, and settlement — are negotiated directly between buyer and provider.