AIGC

Text-to-video has widened the scope of AI-generated content, while world models are pushing the technology into gaming, robotics and manufacturing

  

By CLS Marketwatch

AI has achieved many astonishing leaps in content generation over the last few years, across images, videos and audio, blurring the boundary between the digital and real worlds. This trajectory includes the fast-emerging field leveraging AI to generate digital content (AIGC). Yet, as visuals grow sufficiently lifelike and audio increasingly natural, a pivotal question arises: where should AI head next? As the industry diversifies and iterates, an increasingly clear answer is emerging: AI needs to move beyond just “generating materials” toward “understanding the world.”

The next frontier: World model

If the past two years have seen AIGC master content “imitation,” then the core challenge for the next phase is enabling AI to step beyond the screen and acquire the ability to interact in real-time and solve practical problems within real-world scenarios.

Here, versatility of the underlying technology infrastructure lays the groundwork for cross-domain transfer, as computing clusters, massive-scale cleansed datasets, and model training capabilities required for training video models are highly compatible with the technical underpinnings of world models. Nearly all leading text-to-video developers are venturing into world model development. And marginal costs remain manageable while value remains significant, since substantial infrastructure investments already in place can be easily redeployed.

In contrast, gains from text-based large language models are approaching their ceiling, while video and game generation models remain firmly on a clear growth track, which offers a well-defined return profile for investment in world model technologies. In view of this, industry players agree that the next phase calls for AI to be able to “understand the world”.

From pixels to physics-grounded authenticity

With this goal in mind, the world model has achieved a systematic leap in capabilities compared to traditional AIGC tools. Most importantly, the world model’s training data is largely derived from interaction data in the physical real world, rather than simply video footage. This enables the model to internalize fundamental physical principles such as gravity, inertia, and the attenuation of light and shadow. In highly dynamic and complex scenarios, such as explosions, particle systems and fluid effects, the visuals generated by world models exhibit greater physical coherence, fundamentally reducing the incidence of visual anomalies. This not only elevates visual quality but also underscores that generated content possesses an “authentic” foundation that can be verified against the physical world.

Furthermore, traditional video generation tools require users to wait for the final output to be produced after imputing a prompt, which follows a linear, one-time generation process where users passively receive the output. In contrast, world models support full-cycle real-time feedback, adjustment and rendering. Users can intervene and optimize throughout the entire generation process without waiting, thus advancing the creative workflow from an iterative loop of “generate-modify-regenerate” to a continuous interactive stream.

We can take Shengshu AI’s product portfolio as an example. Its three product lines – the Vidu S1 real-time interactive model, Vidu generative world model, and Motubrain world action model for embodied intelligent robots – correspond to the progressive capabilities of world models, from understanding to interaction to action, forming a complete technological ecosystem that spans both digital content and physical entities.

However, the commercial value of world models ultimately needs to be validated in specific real-life scenarios, with the gaming industry serving as a primary application arena. As one of the most advanced forms of digital content, games inherently demand physical simulation, real-time interaction, and three-dimensional spatial understanding, which align closely with what world models can offer.

Gaming heritage powers physical intelligence

Notably, technical capabilities accumulated from the gaming realm are now being reverse-engineered to empower the physical world. For instance, NetEase’s (NTES.US; 9999.HK) SmartEase embodied intelligence brand has migrated 3D virtual technologies and underlying technical frameworks originally developed for the games “Sword of Justice” and “Naraka: Bladepoint” into construction machinery, enabling the pre-training and troubleshooting of models for equipment like excavators and loaders within virtual 3D environments. These AI capabilities, cultivated through game R&D, have been applied to the construction of intelligent systems for construction machinery, producing intelligent excavators and loaders.

At the same time, the fusion of embodied intelligence and physical AI is improving. Examples include Giga AI, which has developed a full-stack world model product matrix covering a wide range of application scenarios, from content creation and driving simulation to embodied intelligence. Mature world models, equipped with general-purpose understanding capabilities, are expected to enable future robots to observe environments and anticipate actions much like humans do. Rather than requiring task-specific training, these robots will rely on vast stores of world knowledge to autonomously execute complex operations.

AIGC is undoubtedly at a critical turning point of evolving from a “producer of digital content” to an “interpreter of the physical world.” As the core technological foundation of this paradigm shift, world models are redefining the way AI understands and generates materials, empowering machines with the ability to learn about space, anticipate motion, interact in real time, and execute complex tasks just as humans do. Applications have already taken off in fields such as gaming, film and TV drama, embodied intelligence and industrial manufacturing. At this stage, AI’s development mirrors the early days of mobile internet, with fusion across a wide range of industries only just beginning.

CLS Marketwatch provides insights and analysis on China’s industries. You can contact the author at liujingyi@cls.cn

To subscribe to Bamboo Works weekly free newsletter, click here

Recent Articles

So-Young does healthcare

So-Young International names new CFO

Cosmetic surgery center operator So-Young International Inc. (SY.US) on Tuesday announced its appointment of Shen Nan as its new CFO effective Aug. 3. She replaces Jin Xing, the company’s CEO,…