A few days ago, we reported on an anonymous model with no company name, no technical whitepaper, and not even any pricing. Ox Alpha (jokingly called by netizens牛來), in just six days it consumed 42 trillion tokens of usage, climbed to the top of OpenRouter’s usage leaderboard, and pushed DeepSeek down. The identity of this mysterious model was finally revealed on August 26: it was indeed a new version of the GLM series from China’s Zhipu AI (Zai). GLM-5.3 FlashBloomberg reported that Zhipu AI confirmed the model’s identity in response to inquiries and said it will open-source the model weights tonight.

Ox Alpha stays anonymous for six days, usage tops OpenRouter
August 20th, a man named stealth/ox-alpha A model quietly appeared on the OpenRouter platform. No company announcement, no technical documentation—just a striking set of specs: a 1.05 million token context window, a 130,000 token maximum output, support for text/image/video input, and one week of free access.
During the free period, Ox Alpha’s usage grew explosively. According to OpenCode’s statistics, it processed approximately 42 trillion tokens in six days, making it the most heavily used model since DeepSeek Flash’s 56-day run. The founder of OpenRouter also confirmed on social media that Ox Alpha’s usage at one point exceeded DeepSeek’s by more than double. Early community testing showed that Ox Alpha beat GPT-5.6 Sol and Fable 5 on code writing and agentic tasks. Its 1.05 million token context window is on par with Gemini 3.7 Flash, but it adds video input capability and is currently free. In comparison, DeepSeek V4 Flash’s context window is only 128,000 tokens.
42T tokens of Ox Alpha in 6 days
Most used model after DeepSeek Flash's 56 day run
Reveal in a few hours pic.twitter.com/pcyubGwm3E
— OpenCode (@opencode) August 26, 2026
Technical Fingerprint: The DNA of the GLM Family
Before Zhipu AI officially acknowledged it, independent researchers had already identified Ox Alpha’s identity through “technical fingerprint analysis.” OrcaRouter’s research team conducted a tokenizer comparison on Ox Alpha, and all 11 indicators matched Zhipu’s GLM family. In numeric probing tests, Ox Alpha consumed 29 tokens, exactly matching GLM (DeepSeek uses 98). Additionally, Ox Alpha’s deployment configuration was highly similar to GLM-5.3, and these settings were adopted just two days after the release of GLM-5.3’s new configuration.
Zhipu AI founder Tang Jie hinted three months ago that a multimodal mode was about to be launched. Previous Stealth anonymous models were ultimately revealed to be releases of GLM-5 or MiMo; this strategy of “warming up anonymously first, collecting high-concurrency data, then officially unveiling” has precedent.
Specifications and positioning of GLM-5.3 Flash
GLM-5.3 is the flagship model released by Zhipu AI on August 14, 2026. It uses the same 743B-parameter base model as GLM-5.2, with all improvements coming from post-training. Compared to the full version of GLM-5.3, the Flash version focuses on speed and cost efficiency, making it a lightweight variant optimized for high-concurrency scenarios.
According to Zhipu AI’s official documentation, GLM-5.3 delivers comprehensive improvements in complex software engineering and agentic capabilities. In code generation, it is currently one of the strongest open-weight models. The Flash version inherits these capabilities while pushing inference costs extremely low. Ox Alpha specifications include: a 1.05 million token context window, support for text/image/video multimodal inputs, knowledge cutoff extended to November 2025, and a cache hit rate of approximately 77.9%. These specifications make it one of the strongest multimodal models currently available for free.
Ox Alpha’s release strategy has sparked widespread discussion. Zhipu AI chose to list the model on OpenRouter without revealing its identity, letting the model’s performance speak for itself. This approach carries high risk—if the model had performed mediocrely, the anonymous release would have drawn no attention. However, Ox Alpha far exceeded expectations, defeating multiple top-tier models on coding and agentic tasks, quickly igniting community discussion.
Open-source weights and their downstream impact
Zhipu AI has confirmed it will open-source the model weights for GLM-5.3 Flash tonight. This means other developers can integrate the model into their own products. For independent developers and small teams, a free, open-source model that supports multimodal input and performs excellently on coding tasks will significantly lower the barrier to AI applications. However, there are a few points to note. The service configuration during the anonymous testing period may differ from the official release version. After the one-week free usage period ends, Zhipu AI has not yet announced subsequent pricing. Additionally, the anonymous testing model may be taken down or repriced at any time.
Ox Alpha’s success also validates the strategy of “first going free and fast, then officially announcing open source.” This approach let the model face market scrutiny without a brand halo, winning attention through actual performance. For other Chinese AI companies, this may become a release model worth referencing. From an industry perspective, Ox Alpha’s emergence once again proves the competitiveness of China’s open-source models. Following DeepSeek, Zhipu AI has made waves in the global market with another open-source model. This tactic of “testing anonymously first, then revealing identity” also reflects how Chinese AI companies’ marketing strategies are evolving in the international market—no longer just publishing technical reports, but building reputation through the product’s own performance.
Source: KOCPC Chinese