Last week, Noetra, a Japanese AI company, announced the launch of their national research effort with the goal of having “Real-World Native AI” by 2030. Noetra will be building out a hyperscale-level AI factory with NVIDIA and using it to train their multimodal foundation model on proprietary Japanese manufacturing/logistics data. This multi-billion-dollar project funded by players such as Sony Group Corporation, SoftBank Corp., NEC Corporation, Honda Motor Co., Ltd, etc. forces us to think on how it will change the worth of the industrial data needed to underpin it.
Parallels with LLM Companies’ Appetite for Text
The value of data, in the context of training AI, is derived from its scarcity, exclusive rights, and difficulty for competitors to reproduce. This equation is clear with frontier labs now willing to pay for textual content. Notable instances include Google and Reddit’s deal in 2024 valued at $60M, OpenAI and News Corp’s deal in 2024 valued at over $250M across 5 years, and Amazon and New York Times’s deal in 2025 worth $25M per year. Recognizing that models are differentiated from the data which they’re trained on is one of the main modes of competition in the AI industry.
Physical AI/real-world models now extend that competition into the industrial world, where the data may be more scarce and ultimately more valuable. Junichi Miyakawa, the CEO of Softbank, even commented for Noetra’s launch: “In a society that coexists with AI, the data held by Japan’s industries and businesses will be a key source of competitive strength”.
Potential Data Pipelines
Warehouse and manufacturing data are typically held within private operations and generated through years of interaction among workers, machines, software systems, etc. This wealth of data already exists for automation/mobility vendors such as Covariant who are using data generated by deployed warehouse robots to improve its robotics models. Machine-vision companies such as Cognex provide tools for collecting, labeling, and retraining production images. Zebra’s scanners, RFID readers, mobile computers, and vision systems capture the identity, location, and human context surrounding warehouse workflows. Granted, all this data is still at the customer level and not yet a deliverable for AI labs. However, with the demand for industrial data rising, the automation/mobility players are sitting on valuable assets that have the potential to be monetized and support real-world model training.
Automation Is Expanding the Industrial Data Supply
The value of industrial data will also depend on the number of connected systems capable of generating it. It’s clear through large initiatives such as Noetra’s that demand is abundant. However, do industrial settings have enough automation, or at least care enough to adopt so there’s enough supply? VDC’s research suggests that warehouses are already prioritizing the technologies that create and contextualize real-world operational information.
44% of Warehouse Organizations View Autonomous Mobile Robots as a Critical Requirement
“What are your views on the following technologies in warehousing environments?“

In the 2025 Buyer Behavior Guide, VDC Strategy captured strong automation sentiment across AMRs and other technologies within warehouses. With over 70% of industry decision-makers viewing AMRs, RFID, wearables, AS/RS, and computer vision as critical requirements or important, it is safe to say captured data will continue to expand. The next challenge will be determining who can turn that growing volume into a usable/scalable training asset. As physical AI advances, data ownership, interoperability, and the ability to capture real-world outcomes may become just as important as the automation systems themselves.
To learn more about trends among key-decision makers in manufacturing, warehousing, and other industries, look for VDC’s 2026 Research Outline and upcoming 2026 Buyer Behavior Guide.