Overview
- ZhipuAI published GLM-5.2 as an open‑source model with code and weights available on GitHub, Hugging Face and ModelScope so researchers and operators can download, run and inspect the model.
- ZhipuAI says GLM-5.2 supports a solid 1 million token context, which lets the model process very long prompts by holding large key‑value caches and prefilling long input sequences before generation.
- The company reports GLM-5.2 ranks as the top globally usable model in a Code Arena blind test and positions its long‑context performance between Claude Opus 4.7 and 4.8 while claiming open‑source state‑of‑the‑art coding scores.
- ZhipuAI and vendors highlight infrastructure work that lowers per‑token compute by about 2.9× at long contexts and enables Day‑0 inference on domestic accelerators, with Moore Threads saying its MTT S5000 card achieved full‑stack adaptation for GLM-5.2.
- Immediate hardware support and open weights could speed enterprise and academic use inside China by allowing local deployment, auditability and tailored optimizations while shifting demand toward high‑memory, high‑bandwidth GPU stacks for million‑token workloads.