№220|02:27 PM ET
Independent reporting on technology, markets & policy
TechEchelon
№01 / Anchor·ARTIFICIAL INTELLIGENCE

ByteDance Trains 10 Trillion-Parameter AI Model in Push to Surpass Anthropic

ByteDance is training an AI model with as many as 10 trillion parameters — three times the size of the largest Chinese model released to date — in a push to rival Anthropic's most advanced systems without relying on model distillation.

SM
Sara Montes de Oca
AUG 8, 2026 · 01:01 PM ET · 3 MIN READ
via Wikipedia (ByteDance)

ByteDance is training an AI model with as many as 10 trillion parameters, a scale that would position the TikTok parent among the most ambitious AI developers in the world and place it within reach of Anthropic's most advanced systems, according to three people with knowledge of the matter.

The development, first reported by Ars Technica, places the Chinese tech giant at an early stage of pre-training — a phase that typically spans three to six months before a model is fine-tuned and, if development proceeds as planned, released.

The proposed model would be roughly three times the size of Moonshot's Kimi K3, currently the largest Chinese model released to date, according to the sources.

Anthropic does not publicly disclose the size of its models. Industry estimates, however, put its most advanced Mythos 5 at approximately 8 trillion parameters and its Fable 5 at roughly 5 trillion. The exact final size of ByteDance's model would only be determined at a later stage of development, one of the people said.

Parameter count sets the fundamental capacity of a model to store information, though overall capability also hinges on factors such as data quality and training methods.

ByteDance's effort signals how aggressively Chinese AI labs are pressing to close — and potentially overturn — the gap with leading American developers. In recent weeks, models from Moonshot and Alibaba have shown strong performance on benchmarks, lagging behind only Anthropic's Fable 5 in certain areas. Industry insiders say multiple Chinese labs are currently training models at the scale of Fable 5, while ByteDance is pushing for the largest footprint of all.

A distinctive element of ByteDance's approach is its deliberate avoidance of model distillation — the process of training a smaller model to replicate the outputs of a larger one — which many Chinese competitors have used to accelerate development. That stance has been in place for more than a year and is credited by some with contributing to a slower pace of releases relative to rivals.

ByteDance founder Zhang Yiming reiterated his preference for independent development in an internal meeting two weeks ago, telling the company's Seed model development team to target "world-leading model capabilities" in the long run without becoming overly focused on near-term rankings, according to one of the people. Chinese media outlet Latepost and The Information first reported Zhang's comments from that meeting.

Seed, the ByteDance division responsible for model development, is led by former Google DeepMind scientist Wu Yonghui and employs approximately 2,000 people across China and overseas, including core researchers, infrastructure engineers, data labeling staff, and translators.

ByteDance has invested more aggressively in AI over the past three years than any other major Chinese tech company, building out data centers and expanding its cloud unit, Volcano Engine, which sells AI services to enterprises. The company is also developing custom AI chips.

Its existing AI portfolio is already substantial. The SeeDance model ranks among the top performers globally in video generation, and Doubao — its flagship consumer-facing model — counts 324 million monthly active users, making it the most popular AI product in China.

Mythos 5, Anthropic's most capable model, has been restricted to approved organizations since a temporary ban in June on security grounds, underscoring the broader geopolitical backdrop against which ByteDance is racing to build independent frontier capabilities.

ByteDance did not respond to a request for comment.

Disclaimer

SM
━ ABOUT THE REPORTER
Sara Montes de Oca

Sara Montes de Oca is the Editor in Chief of TechEchelon. Previously a correspondent and producer in Washington, D.C., covering business, finance, and politics.

More from Sara
● THE BRIEF · DAILY NEWSLETTER

Five stories every morning. Before the opening bell.

Written for readers who already know the basics — markets, AI, and the policy decisions that shape both.

Mon — Fri · 06:30 ET · Free

No spam · Unsubscribe anytime