ByteDance Establishes AI Data & Safety as a First-Level Department
Nashnova编辑部
ByteDance has merged several scattered data teams into its third top-level AI division — AI Data & Safety — sitting alongside Seed and Flow. This means data supply has been elevated from a support function to a front-line weapon in the foundation-model race.
What does the new division actually do?
It merges Global Data (formerly under TikTok), the group-wide data platform DMC, and Flow's AI Data Platform (AIDP) into a single organization.
Its core mandate: provide cross-modal data services to every ByteDance foundation model, covering the full pipeline — standard-setting, sourcing, synthetic cleaning, and quality evaluation.
This means → ByteDance is pulling data capabilities out of individual business lines into one entity that spans from base models to applications — one gateway serving all models.
Who is leading it, and why him?
The head is Adam Wang (王赢磊), previously responsible for TikTok platform integrity and TikTok livestreaming.
According to people close to ByteDance, the livestreaming business he ran was one of TikTok's largest revenue sources.
In plain terms = ByteDance picked someone with large-scale commercial operations experience, not a pure technologist. This reflects a division positioned beyond "collect and label" — it also owns budget management and vendor relationships.
Why elevate data to this level?
The direct driver: ByteDance's commitment to "absolutely no distillation" — building its own models from scratch. Zhang Yiming stated this at the late-July Seed all-hands; CEO Liang Rubo echoed it on August 5: "We will build in-house, accept short-term lag, and optimize for the long term."
This means → without distillation — extracting capabilities from others' models as a shortcut — the only lever is the quality and scale of your own data. That pushed the data team to top-level status.
The spending is already heavy: data budgets for world models and coding models exceeded tens of millions of dollars in early 2026, with standing authorization to "top up at any time."
How large is the investment, and where does it go?
The team doing data evaluation for the video-generation model Seedance alone numbers over 1,000 people; each algorithm engineer is backed by roughly a dozen data colleagues.
Multiple industry practitioners call Seedance 2.0's success "a victory of data."
The data organization now runs a horse-race mechanism — parallel teams competing across world models, code, and advanced academic domains — with every project required to account for its input-output ratio.
How intense is the industry-wide data arms race?
Tencent has been poaching from ByteDance's data teams at up to 3× salary over the past six months; Alibaba and Tencent have also visibly increased data-procurement budgets, with some giants setting exclusivity windows on datasets or locking up key vendor staff.
Globally, top-tier model companies' external data budgets are expanding by billions of dollars per year: Anthropic alone earmarked over $1 billion for reinforcement-learning data in 2025.
Data-services firm Mercor grew annualized revenue from $500 million last year to $2 billion by mid-year, with 91% coming from OpenAI, Anthropic, and peers; its valuation has surged to $20 billion.
What is the biggest challenge after the merger?
The new division involves merging multiple legacy teams with overlapping mandates; ByteDance is still sorting out the org chart and optimizing headcount.
This reflects a real tension: the data organization has reached a thousand-person scale, yet the foundation-model frontier shifts fast — the larger the org, the harder it is to stay nimble.
In plain terms = whether this restructuring translates into model competitiveness depends on whether a thousand-person team can pivot as quickly as a hundred-person one.
Content is for reference only, not financial advice.