AiGENTiA InsightsPhysical Limits

Physical Data and AI

30 August 2026


Models have consumed the text of the world. The physical world is not text, and it is not digitised — which is where the next constraint bites.

This essay is still being written. The outline below is the argument it will make.

The scaling trajectory of artificial intelligence has encountered a structural ceiling. For a decade, scaling meant ingesting more web text, crawling public depositories, and converting human discourse into high-dimensional embeddings. That strategy has reached its natural limit. Models have consumed the text of the world. The physical world is not text, and it is not digitised — which is where the next constraint bites.

Public text data will be exhausted within the decade

Empirical research confirms the arrival of this wall. According to the Epoch AI Data Limits Study, detailed in their arXiv:2211.04325 Paper Entry, the total effective stock of public human-generated text data spans between 300 and 400 trillion tokens. Frontier large language models will fully deplete this public corpus between 2026 and 2032. Synthetic text generation offers marginal optimization, but self-referential training runs into degenerated representations. The internet is empty of new unstructured text at the scale required for the next order-of-magnitude leap.

What software did to document workflows, foundation models are attempting to do to physical operations. Yet, writing code to summarize prose is fundamentally different from controlling a robotic arm, driving a haul truck, or conducting an ultrasound. Text is an abstracted, low-bandwidth shadow of reality. The physical world operates on continuous spatial fields, dynamic friction, force vectors, and material degradation. The digital corpus is running dry because it captured human thought, not physical causality.

Physical data is hard to capture, expensive to label, and impossible to scrape

Web text was scraped for the cost of bandwidth. Physical data cannot be copied from a server; it must be gathered by hardware interacting directly with dynamic environments. The instrumentation, cost, spatial coverage, and operational consent required to capture physical ground truth introduce cost structures that software companies have never faced.

The unit economics of physical data collection are brutal. As documented in the NeuralChain AI Data Collection Pricing Guide, capturing physical robot action data costs between $2.00 and $150.00 per usable demonstration episode depending on task complexity, such as transitioning from single-arm pick-and-place tasks to bimanual fine manipulation with integrated force capture. Operator labor costs alone run between $25.00 and $90.00 per hour.

The resulting scarcity is stark. While language models train on trillions of tokens, the multi-institutional Open X-Embodiment Project Paper, also archived in the arXiv:2310.08864 Paper Entry, aggregates datasets across 21 institutions and 22 robot embodiments into only ~1 million trajectories, representing 527 skills across ~160,000 tasks. Across the entire global research ecosystem, the stock of usable, high-fidelity real-robot data is estimated at only a few hundred thousand hours. Specialized efforts like the Hugging Face NVIDIA Open-H-Embodiment Data Repository, using the LeRobot format for healthcare and ultrasound applications, demonstrate the precision required when pairing kinematics with video trajectories, but such efforts remain isolated islands in a vast physical sea.

Shortcut solutions fail under scrutiny. A common assumption is that passive internet video—such as millions of hours of YouTube uploads—can substitute for dedicated hardware sensors. As detailed in the Neuracore Data Collection Architecture Breakdown, passive video records visual observations, but completely omits motor control signals, action labels, joint kinematics, force-torque vectors, and tactile resistance. Visual representations can be pre-trained on video, but real-world execution requires closed-loop telemetry. Without direct force and kinematic feedback, a model cannot calibrate pressure or anticipate resistance.

To bridge this coverage deficit, companies are forced into physical crowdsourcing. As reported by Gasgoo regarding Robotics Data Crowdsourcing, Daimon Robotics partnered with China Mobile to deploy embodied data collection across thousands of offline telecom retail locations, utilizing distributed physical sites to capture diverse spatial environments. Gathering physical data requires boots on the ground, sensors in the field, and capital on the balance sheet.

Physical telemetry is concentrated in proprietary industrial fleets

Public text belonged to the open web. Physical data belongs to the asset owner. Because capturing real-world telemetry requires deployed hardware fleets, ownership of physical data is far more concentrated than text data ever was.

Consider autonomous mobility. According to the Teslarati FSD Fleet Mileage Benchmark and official Tesla Safety Report Telemetry Disclosures, Tesla’s Full Self-Driving (Supervised) fleet logged over 10 to 13 billion cumulative real-world driving miles by mid-2026. In Q3 2025 alone, Tesla received approximately 2.5 billion telemetry packages from its global fleet, generating between 20 and 30 million new driving miles every single day. No academic institution or software startup can replicate this telemetry pipeline through web crawlers or synthetic generation.

A parallel dynamic controls heavy industry. As highlighted in the DigitalDefynd Caterpillar AI Case Study, Caterpillar embedded proprietary telemetry modules across its heavy excavation and mining fleets via Cat® Product Link™, streaming real-time engine load, hydraulic pressure, and thermal metrics into its Cat® Equipment Care Advisor platform. This telemetry powers autonomous haulers and predictive maintenance engines that digital-native competitors cannot observe.

This telemetry creates structural pricing power. The DataIntelo Heavy Equipment Telematics Report reveals that proprietary telematics platforms like Caterpillar Cat Connect and John Deere Operations Center achieve 94% to 97% failure prediction accuracy. This exclusive, non-scrapable machine telemetry stream allows these OEMs to command 18% to 22% pricing premiums on equipment and service contracts. Web crawlers can index public manuals, but they cannot index the thermal telemetry of a hydraulic pump operating under load in a remote quarry.

Embodied execution requires physical ground truth

When physical interaction data is captured at scale, AI transitions from text generation to real-world agency. When it is absent, spatial systems remain trapped in static visual prediction.

Capital markets recognize this inflection point. According to the Forbes Analysis on World Labs, Dr. Fei-Fei Li’s World Labs raised $1 billion in funding at a $5 billion valuation in February 2026, supported by a $200 million strategic check from Autodesk. Their product, Marble, focuses on generating persistent 3D world models to teach AI systems spatial awareness and physical causality beyond text tokens. Similarly, a report from The Elec on Physical AI Valuations outlines how Physical Intelligence ($\pi_0$) raised $600 million in Series B funding at a $5.6 billion valuation—with subsequent 2026 funding talks targeting an $11 billion valuation—despite owning zero hardware chassis or commercial end-products. Their valuation rests entirely on generalist Vision-Language-Action (VLA) foundation models trained on physical interaction data.

Conversely, platforms that rely exclusively on synthetic physics engines hit a hard capability wall. Advocates argue that high-fidelity physics simulators such as NVIDIA Isaac Sim or MuJoCo eliminate the need for real physical data by generating infinite synthetic scenarios. But as analyzed in the Avala AI Analysis of Sim-to-Real Data Leaks and detailed in Weighty Thoughts on World Models & Physics Simulation, pure simulation suffers from the persistent reality gap. Physics engines struggle to accurately compute micro-friction, non-rigid deformable materials, complex hydraulic fluid dynamics, and unpredictable sensor noise.

Simulation acts as a compute multiplier for real data, not a replacement. Models trained purely in simulated environments stumble when confronted with the unmodeled dynamics of physical reality. High-level spatial reasoning stalls without the ground truth of continuous physical feedback.

Proprietary sensor streams create unassailable competitive moats

The shift from text tokens to physical telemetry changes the structure of corporate advantage. In the first wave of enterprise AI, advantages accrued to companies with large text databases, customer support logs, and document repositories. In the agentic physical economy, advantage shifts entirely to operators who own real-time sensor streams.

If an enterprise operates industrial machinery, manages vehicle fleets, monitors logistics hubs, or runs physical facilities, its sensor telemetry is no longer operational exhaust. It is the primary asset required to build autonomous agency. Software startups can clone model architectures overnight, but they cannot simulate ten billion miles of physical road conditions or millions of hours of excavator hydraulic stress.

We are witnessing a fundamental shift in asset valuation. The competitive moat of the next decade will not belong to those who build the foundation models, but to those who control the sensors feeding them. The internet was built on text. The physical economy will be built on telemetry.

Keep reading