Open-weight & owned stack · Quantization efficiency
Zhipu Reports Fast GLM Inference On Chinese Chips
Zhipu said GLM-5.3-FlashX reached nearly 200 tokens per second across roughly 100,000 domestic AI accelerators.
Read the original at finance.biggo.comOpens the publisher's site in a new tab