
The Zhitong Finance App learned that on the evening of August 26, Qianwen Office launched the newly released Qwen 3.8-Flash model for the first time, and launched the standard model at the same time. From now on, all users can experience Qwen3.8-Flash in a new standard mode. Based on the latest model, users can complete tasks with less credit consumption and faster token throughput. In the future, the model supply for Qianwen Office will be limited to the standard and advanced models. 95% of daily tasks can be completed through the Qianwen Office standard model, and only 5% of complex tasks will need to use the advanced model.

Improved user experience comes from model upgrades and agent collaborative optimization. The newly architected Qwen3.8-Flash achieved performance that surpassed Claude Opus 4.6 with 100 billion total parameters. At the same time, the Qianwen Big Model Team and the Qianwen Office Team also jointly launched an office-specific version of Qwen3.8-Flash, which performs special training and tuning for scenarios such as multi-step planning, tool selection, and context compression, and further improves throughput efficiency through inference optimization and customized Harness architectures. In the actual office scenario test, the single-task generation speed of the Qianqu Office Standard Mode was increased by about 100%, and token consumption was reduced by an average of 75%.
In real AI application scenarios, high performance usually means high cost and high latency, while low cost requires sacrificing intelligence. Deep collaborative optimization of agents and models is breaking the “impossible triangle” of performance, cost, and speed. As the intelligence density of models continues to increase, and the two-way optimization of the Qianwen Office and model, Agents will soon bid farewell to Token anxiety and usher in an era of “big volume management”.