Multimodal AI growth forces companies to rethink computing costs

As creation of content expressed in multiple formats grows, video generation costs are reshaping companies’ AI computing budgets
By LeadLeo Research Institute
Multimodal content creation is moving from a niche use of large models toward the forefront of AI applications. LeadLeo research shows that its share of nine core business application categories rose from 0.6% to 11.9% in six months, making it the fastest-growing category. The change goes beyond a shift in the rankings: Generating a 10-second, 1080p video consumes about 350,000 tokens, hundreds of times as many as a text task. The implications extend to content platforms, marketing departments and anyone planning computing capacity purchases.
Fast-changing rankings
The report examined a sample of about 700 large-model applications. In the first half of 2025, text content creation ranked first with a 23.7% share, followed by question-and-answer assistants, data processing and analysis. Multimodal content creation accounted for just 0.6%, placing last among the nine categories.
Six months later, the rankings began to change. Multimodal content creation rose to 11.9%, while intelligent customer service climbed from 5% to 9.4%, and AI search from 3.1% to 4.7%. The report identified these as the three fastest growing categories, with multimodal content creation posting the largest gain.
The core applications remained strong. Text content creation still ranked first in the second half, with a 19.5% share, while question-and-answer assistants and data processing also continued to account for substantial share. Multimodal applications captured a larger share of the growth without replacing text-based uses.
Why did multimodal applications grow? The report attributes the increase to continuing improvements in multimodal technology in 2025, which accelerated demand in content production, information retrieval and customer interactions. Content-heavy applications were among the first to gain traction: The number of games on the Steam platform using generative AI rose from more than 1,000 in 2024 to more than 7,818 in 2025, an increase the report puts at 681%.
Revenue from video generation services is also growing. Kuaishou (1024.HK) said its Kling AI service generated more than 850 million yuan ($127 million) in revenue in the second quarter of 2026, up more than 200% year-over-year. Spending on AI-generated short-video marketing materials on its platform rose more than 70% in the same period. From video production to advertising, multimodal applications are entering more commercial settings.
Changing unit of account
The shift in application share also changes how computing demand is measured. The report estimates that generating a 10-second, 1080p video consumes about 350,000 tokens, hundreds of times as many as a text task. When an application moves from writing copy to making video, the computing requirements for a single task are on a different scale.
The way models are used adds to that demand. The report observed a shift from one-off generation toward continuous inference and end-to-end tasks. As models become embedded in frequently used workflows such as search, marketing, customer service and office work, the usage volume, frequency of interactions and inference load all rise.
Token usage nationwide is growing as well. China’s National Data Administration said average daily token use exceeded 140 trillion in March this year. The National Bureau of Statistics later said the daily figure had reached “several hundred trillion,” without providing a more precise number. The measure covers all types of AI applications and reflects overall growth in model usage.
Taken together, these figures change how companies calculate their computing budgets. While the initial question was whether a model could perform a task, costs now need to be measured in terms of token usage volume multiplied by consumption per task. As work shifts from text to video and from one-off generation to continuous use, costs may grow faster than the user base.
Who pays first?
The cost pressure won’t be shared equally. Multimodal growth is concentrated in content-heavy applications such as games, leaving internet content platforms and AI-native companies among the first to bear the costs. The report says these customers need high concurrency, elastic capacity and low-cost usage models, making model-as-a-service (MaaS) and cloud-based inference services a fit for usage-based pricing.
The scope of the figures matters. The application shares are based on a sample of about 700 and reflect changes in that sample’s usage mix, not the distribution of industry revenue. Whether multimodal applications continue to grow also depends on how quickly video generation costs fall. The report identifies pricing and the maturity of available services as key variables affecting computing demand.
For companies, the implications come down to three considerations: including multimodal production costs in annual budgets; specifying how model usage will be billed in procurement contracts; and assessing, task by task, which processes warrant an upgrade from text to video generation. In planning an AI budget, companies first need to determine how many tasks will move from text to video and from one-off generation to continuous inference.
Multimodal AI is more than just another feature option: It can multiply the computing required for a single task hundreds of times. Rather than focusing only on model rankings, companies should calculate how many tasks in their business will shift from text to video and from one-off generation to continuous inference. The first shift affects the experience, while the second affects the cost.
LeadLeo Research Institute is an original content platform for research on banks and companies and an innovative digital research service provider with nearly 100 senior analysts. You can contact the platform at CS@leadleo.com
This commentary is the views of the writer and does not necessarily reflect the views of Bamboo Works
To subscribe to Bamboo Works weekly free newsletter, click here