Z.ai GLM-5.3-Flash: China's AI Chip Independence
Summary
Z.ai's GLM-5.3-Flash model can now run on Chinese AI chips, not just Nvidia's. This matters because it shows a Chinese lab can serve a frontier-class coding model on domestic hardware. The model, which appeared on developer tools around August 20, was confirmed by Beijing-based Z.ai on August 26. Z.ai stated the model was served on a large cluster of Chinese AI chips. GLM-5.3-Flash is a significant model, described as having 320 billion parameters with 18 billion active parameters and a one-million-token context window. It also supports text, image, and video. The company's API prices it cheaply, at $0.15 per million input tokens. Z.ai claims their setup improved serving performance by three times on domestic hardware. This brings the per-token cost close to mainstream Nvidia GPUs. This development suggests a growing independence in AI chip technology.
This is an AI-generated audio summary. Always check the original source for complete reporting.