Section 7: Commercial Deployment (2025-2026)
The autonomous driving industry has transitioned from research demonstrations to commercial-scale operations. By the end of 2025, multiple companies are running paid robotaxi services across dozens of cities worldwide, autonomous trucks are hauling freight on public highways, and new business models are emerging as companies pivot and consolidate. This section surveys the current state of commercial AD deployment across robotaxis, trucking, delivery, and Japan's domestic programs.
7.1 Robotaxi Services
Robotaxi services represent the most visible and commercially significant application of autonomous driving technology. As of late 2025, several companies have moved beyond pilot programs to sustained, revenue-generating operations with plans for rapid expansion in 2026.
| Company | Cities (End 2025) | 2026 Plan | Weekly Rides |
|---|---|---|---|
| Waymo | 5 (US: San Francisco, Phoenix, Los Angeles, Austin, Atlanta) | 15+ cities including London (first international market) | 450K, targeting 1M |
| Baidu Apollo Go | 10 (China: Beijing, Shanghai, Guangzhou, Shenzhen, Wuhan, Chongqing, etc.) | Middle East and European expansion | ~250K |
| Tesla | 2 (limited deployment in Austin, San Francisco Bay Area) | Multiple city expansion planned | Undisclosed |
| Zoox (Amazon) | 2 (San Francisco, Las Vegas) | Paid public service launch and geographic expansion | Undisclosed (fleet of ~50 purpose-built vehicles) |
| Pony.ai | Multiple cities in China (Beijing, Shanghai, Guangzhou, Shenzhen) | Scale to 3,000 vehicle fleet | Undisclosed |
Waymo remains the clear global leader in robotaxi deployment. Operating across five US cities by the end of 2025, the Alphabet subsidiary is delivering approximately 450,000 weekly rides with plans to surpass 1 million weekly rides as it expands to 15 or more cities in 2026. The planned London launch marks Waymo's first foray outside the United States, signaling confidence in the transferability of its technology stack to international road environments with different driving conventions (left-hand traffic), road markings, and regulatory requirements.
Baidu Apollo Go dominates the Chinese market, operating in 10 cities with approximately 250,000 weekly rides. Baidu's 2026 strategy includes ambitious international expansion into the Middle East and Europe, leveraging partnerships with local transportation authorities. The company benefits from China's regulatory environment, which has been comparatively supportive of AD testing and deployment.
Tesla is pursuing a fundamentally different approach to robotaxi operations, relying on its existing consumer vehicle fleet equipped with camera-only sensor suites. While deployment remains limited as of end 2025, Tesla's scale advantage (millions of vehicles already on roads collecting data) represents a unique asset if the vision-only approach achieves sufficient reliability.
Zoox, acquired by Amazon, is notable for its purpose-built, bidirectional robotaxi vehicle designed without a traditional driver's seat. With approximately 50 vehicles in operation across San Francisco and Las Vegas, Zoox plans to launch paid public service and expand its geographic footprint in 2026.
Pony.ai, backed by Toyota, operates across multiple Chinese cities and has announced plans to scale to a fleet of 3,000 vehicles. The company completed a successful IPO on NASDAQ in late 2024, providing capital for its aggressive expansion plans.
7.2 Autonomous Trucking
Autonomous trucking has emerged as a compelling commercial application, driven by persistent driver shortages, predictable highway environments, and the economic value of long-haul freight. Several companies achieved significant milestones in 2025.
Aurora Innovation
Aurora achieved a landmark milestone by launching the first US public road driverless freight service. Aurora's autonomous trucks have logged over 20,000 miles of driverless highway operation, primarily along the Interstate 45 corridor in Texas. The company has established key commercial partnerships with FedEx and Werner Enterprises, two of the largest freight carriers in the United States, providing a clear path to revenue generation and fleet scale-up. Aurora's technology stack, anchored by the Aurora Driver platform, integrates LiDAR, radar, and cameras with a sophisticated planning and control system optimized for highway driving.
Kodiak Robotics (Kodiak AI)
Kodiak achieved another first by launching the first US commercial driverless trucking operations on private roads in West Texas. While the initial operations are geographically limited to controlled environments, the commercial nature of these operations (hauling real freight for paying customers) marks a significant step beyond demonstration runs. Kodiak also pursued a SPAC listing, targeting a valuation of approximately $2.5 billion, reflecting investor confidence in the autonomous trucking market.
Plus
Plus (formerly Plus.ai) is pursuing a SPAC listing with a planned valuation of approximately $1.2 billion. The company's strategy focuses on factory-produced autonomous truck commercialization, with a target of bringing production vehicles to market by 2027. Plus differentiates itself through OEM partnerships that integrate its AD technology directly into new truck platforms during manufacturing, rather than retrofitting existing vehicles.
TuSimple (now CreateAI)
In one of the most dramatic pivots in the AD industry, TuSimple, formerly a leading autonomous trucking company, renamed itself to CreateAI and pivoted entirely away from autonomous driving into the AI animation business. This decision followed a series of operational challenges, regulatory scrutiny, and leadership changes. The pivot serves as a cautionary tale about the capital intensity and technical difficulty of bringing autonomous trucking to commercial scale.
7.3 Autonomous Delivery
The autonomous delivery sector has undergone significant consolidation and strategic pivots, reflecting the challenges of building a sustainable business model around last-mile delivery robots.
Nuro
Nuro, once the leading autonomous delivery robot company, executed a major strategic pivot away from its original delivery robot business model toward AD technology licensing. Rather than operating its own fleet of delivery vehicles, Nuro now licenses its autonomous driving technology stack to partners. The most significant partnership is with Lucid Motors and Uber, targeting a 2026 robotaxi launch. This partnership envisions a six-year program with a fleet of up to 20,000 vehicles, representing a substantial commercial commitment. Additionally, Nuro has begun data collection operations in Japan, signaling interest in the Japanese market where regulatory support for autonomous mobility services is growing.
Amazon Scout
Amazon discontinued its Scout autonomous delivery robot program in 2022, consolidating its autonomous vehicle efforts around Zoox, which Amazon acquired in 2020 for approximately $1.3 billion. This consolidation reflects Amazon's strategic decision to pursue higher-value robotaxi applications rather than sidewalk delivery robots, which faced challenges around payload capacity, sidewalk navigation complexity, and unit economics.
7.4 Japan Domestic Deployment
Japan has adopted a proactive regulatory and policy approach to autonomous driving, motivated by the country's aging population, declining workforce, and the need to maintain transportation services in rural and depopulated areas.
Government Targets
- FY2025: Driverless autonomous driving services operating at 50 locations nationwide
- FY2027: Expansion to 100 or more locations
- 2030: 10,000 Level 4 autonomous vehicles in operation across Japan
Key Deployments
Eiheiji Town, Fukui Prefecture: In May 2023, Eiheiji became the site of Japan's first Level 4 autonomous driving on public roads. The service operates a 7-passenger electric cart along a 2-kilometer route, providing transportation for elderly residents. This deployment, while modest in scale, established the regulatory and operational precedent for subsequent L4 deployments across Japan.
Shiojiri City, Nagano Prefecture: TIER IV (a leading Japanese AD technology company) achieved Japan's first Level 4 certification for an autonomous bus operating at up to 35 km/h. This certification under Japan's revised Road Traffic Act represents a significant regulatory milestone, demonstrating that Japan's legal framework can accommodate autonomous public transportation.
Tokyo National Diet Area: TIER IV operates an autonomous shuttle service covering a 3.5-kilometer route in the government district of central Tokyo. This high-visibility deployment near Japan's Parliament building serves both as a public demonstration and a practical transportation service for the area.
Yokohama, Minato Mirai: A collaborative demonstration by Nissan and BOLDLY (a SoftBank subsidiary) in Yokohama's Minato Mirai waterfront district. This deployment showcases the potential for autonomous mobility in Japan's urban commercial districts.
Expressway Operations: Japan has conducted autonomous truck demonstrations on the Shin-Tomei Expressway, with plans to launch regular autonomous trunk transport operations between the Kanto (Tokyo) and Kansai (Osaka) regions starting July 2025. This expressway deployment targets one of Japan's most critical freight corridors and addresses the country's acute truck driver shortage.
Section 8: AI/ML Cutting Edge
The AI and machine learning techniques powering autonomous driving are undergoing a paradigm shift. Traditional modular pipelines (perception, prediction, planning as separate components) are giving way to End-to-End (E2E) neural network architectures that learn the entire driving task from sensor input to vehicle control output. Simultaneously, Foundation Models, World Models, and Vision-Language-Action (VLA) architectures are bringing the generalization capabilities of large-scale AI to autonomous driving. This section covers the frontiers of AD AI research and development as of early 2026.
8.1 End-to-End Autonomous Driving
End-to-End autonomous driving replaces the traditional modular pipeline with a single (or loosely coupled) neural network that maps raw sensor inputs directly to driving actions. The potential advantages include eliminating information loss at module boundaries, enabling the system to learn features that are directly relevant to driving performance, and simplifying the overall system architecture.
Tesla FSD V13/V14
Tesla's Full Self-Driving (FSD) system underwent a complete End-to-End neural network overhaul with versions V13 and V14, representing the most aggressive production deployment of E2E autonomous driving technology.
Architecture: The FSD V13 system processes multimodal inputs including all camera feeds (8 cameras), navigation data, vehicle state information (speed, steering angle, acceleration), and even audio signals from the vehicle's microphones. All of these inputs are fused within a unified neural network that outputs driving trajectories directly.
Scale increases from prior versions:
- Model size: 3x larger than previous FSD versions
- Data input: 4.2x more data ingested per inference step
- Training compute: 5x increase in training compute requirements
- Hardware: Training conducted on a cluster of 29,000 NVIDIA H100 GPUs
Lane Connectivity Network: A key innovation in FSD V13 is the Lane Connectivity Network, a Transformer-based autoregressive model that generates an understanding of road layout in real time. Rather than relying on pre-mapped lane information, this network predicts lane structure, connectivity, and topology from camera observations, enabling the system to navigate unmapped roads and construction zones.
FSD V14: Tesla has announced plans for FSD V14, which is expected to feature a neural network model approximately 10x larger than V13. This dramatic scale increase reflects Tesla's belief that model capacity remains a primary bottleneck for driving performance, and that their massive training infrastructure (expanding beyond 100,000 GPUs) can support training at this scale.
NVIDIA Alpamayo (CES 2026)
NVIDIA unveiled Alpamayo at CES 2026, a 10-billion parameter Chain-of-Thought Vision-Language-Action (VLA) model purpose-built for autonomous driving. Alpamayo represents a new paradigm in AD AI by incorporating explicit language-based reasoning into the driving decision process.
Architecture: Alpamayo comprises two main components:
- Cosmos Reason (8.2 billion parameters): A vision-language model that processes video and sensor inputs, generating natural language descriptions of the driving scene and causal reasoning about the situation (e.g., "The pedestrian is looking at their phone and stepping off the curb, so I should slow down and prepare to stop.")
- Action Expert (2.3 billion parameters): A specialized model that translates the language-based reasoning into concrete driving trajectory outputs
Pipeline: Video and sensor data flow into Cosmos Reason, which produces a chain-of-thought reasoning trace in natural language. This reasoning is then consumed by the Action Expert, which generates the final driving trajectory. This two-stage approach enables interpretable decision-making, a significant advantage for safety validation and regulatory approval.
Training Infrastructure:
- AlpaSim: A dedicated simulation framework for generating training scenarios
- Dataset: Over 100 TB of real-world driving data collected from 25 countries, comprising 1,727 hours of driving with 7 cameras, LiDAR, and radar per vehicle
Adoption: Alpamayo has been adopted by several major automotive partners including JLR (Jaguar Land Rover), Lucid Motors, Uber, and Berkeley DeepDrive (academic research). The Mercedes-Benz CLA is announced as the first production vehicle to integrate the Alpamayo platform.
Wayve
UK-based Wayve is pursuing a data-driven End-to-End approach, training its models on thousands of GPUs on Microsoft Azure cloud infrastructure. Wayve's approach is notable for its emphasis on learning to drive from data alone, without hand-coded rules or HD maps. The company has announced a partnership with Nissan for deployment in production vehicles targeting 2027, representing one of the first OEM commitments to a purely data-driven E2E driving system.
End-to-End Pipeline Architecture
The following diagram illustrates the general architecture of modern End-to-End autonomous driving systems, showing how raw sensor inputs are processed through a unified neural network to produce driving actions:
(Multi-view)"] LID["LiDAR
(Point Cloud)"] RAD["Radar
(Velocity)"] NAV["Navigation
(Route)"] VEH["Vehicle State
(Speed, Steering)"] AUD["Audio
(Ambient)"] end subgraph Backbone["E2E Neural Network"] ENC["Multi-Modal
Encoder"] BEV["BEV / 3D
Feature Space"] TRF["Transformer
Backbone"] DEC["Trajectory
Decoder"] end subgraph Reasoning["Optional: Chain-of-Thought"] VLM["Vision-Language
Model"] COT["Language-based
Causal Reasoning"] end subgraph Output["Driving Output"] TRJ["Planned
Trajectory"] CTL["Vehicle
Control"] end CAM --> ENC LID --> ENC RAD --> ENC NAV --> ENC VEH --> ENC AUD --> ENC ENC --> BEV BEV --> TRF TRF --> VLM VLM --> COT COT --> DEC TRF --> DEC DEC --> TRJ TRJ --> CTL style Inputs fill:#e3f2fd,stroke:#1565c0 style Backbone fill:#fff3e0,stroke:#ef6c00 style Reasoning fill:#f3e5f5,stroke:#7b1fa2 style Output fill:#e8f5e9,stroke:#2e7d32
In this architecture, the multi-modal encoder fuses all sensor inputs into a shared representation space (often a Bird's Eye View or 3D volumetric space). A Transformer backbone processes these features, optionally generating chain-of-thought reasoning via a Vision-Language Model (as in NVIDIA's Alpamayo). The trajectory decoder then produces the planned path, which is converted to low-level vehicle control commands (steering, throttle, brake).
8.2 Foundation Models and World Models
Foundation Models (large-scale pretrained models adaptable to diverse tasks) and World Models (models that predict how the environment evolves over time) are becoming central components of next-generation autonomous driving systems. These models bring generalization, scene understanding, and predictive capabilities that complement or replace traditional rule-based components.
World Models
World Models learn to simulate the dynamics of the driving environment, enabling autonomous vehicles to "imagine" future scenarios and plan accordingly. This capability is valuable for both training (generating synthetic scenarios) and inference (predicting the consequences of potential actions).
GAIA-1 / GAIA-2 (Wayve): Wayve's Generative AI for Autonomy (GAIA) models treat scene evolution prediction as a next-token prediction task, analogous to how large language models predict the next word in a sequence. GAIA-1 demonstrated that driving scene prediction could be framed as a generative modeling problem. GAIA-2, released in March 2025, extends this to a multi-view controllable generative world model capable of generating consistent multi-camera views of predicted future driving scenes. This enables more realistic simulation and better training data generation.
DriveDreamer Series: The DriveDreamer family of models has produced a rapid succession of innovations:
- DriveDreamer (ECCV 2024): The original real-world driving data world model, capable of predicting future driving scenes from current observations
- DriveDreamer-2 (AAAI 2025): Enhanced with LLM integration for better scene understanding and more diverse scenario generation through natural language conditioning
- DriveDreamer4D (CVPR 2025): Extends scene prediction to full 4D (3D space plus time), generating spatiotemporally consistent future driving scenes
OccWorld (ECCV 2024): OccWorld employs a Spatiotemporal Transformer architecture for predicting future 3D occupancy grids and ego-vehicle poses. By operating in 3D occupancy space rather than 2D image space, OccWorld provides geometrically accurate predictions that are more directly useful for planning and collision avoidance.
Cosmos (NVIDIA, 2025): NVIDIA's Cosmos is a comprehensive Physical AI world model platform designed to support the development of AI systems that understand and interact with the physical world. Cosmos provides pretrained world models, training infrastructure, and APIs that enable AD developers to build on NVIDIA's massive-scale foundation rather than training world models from scratch.
2025 New Entries (CVPR 2025): The field continues to advance rapidly, with several notable new world models presented at CVPR 2025:
- GaussianWorld: Leverages 3D Gaussian Splatting for efficient, high-fidelity world model rendering
- ReconDreamer: Focuses on reconstructing and dreaming new driving scenarios from limited observations
- FUTURIST: A future-prediction model with strong temporal consistency
- MaskGWM: A masked generative world model with improved training efficiency
LLM/VLM/VLA Applications in Autonomous Driving
Large Language Models (LLMs), Vision-Language Models (VLMs), and Vision-Language-Action (VLA) models are being applied to autonomous driving in increasingly sophisticated ways, moving from auxiliary tools to core planning components.
DriveMLM: A multimodal LLM specifically designed for autonomous driving behavior planning. DriveMLM processes visual inputs alongside natural language descriptions of the driving context to generate high-level behavioral decisions (e.g., "change to left lane," "yield to pedestrian"). This approach bridges the gap between human-like reasoning and machine execution.
DiffVLA++: Presented at ICCV 2025, DiffVLA++ bridges the gap between cognitive reasoning and End-to-End planning. The model combines a vision-language understanding component (which "thinks" about the scene in natural language) with a diffusion-based action generation component (which produces smooth, physically plausible driving trajectories). This fusion enables both interpretable reasoning and high-quality motion planning in a single framework.
Research Scale: The intersection of Foundation Models and autonomous driving has attracted enormous research attention, with over 348 published papers on Foundation Models for AD as of early 2026. The TUM-AVS repository (maintained by the Technical University of Munich's Autonomous Vehicle Systems group) provides a comprehensive catalog of AD Foundation Model research, serving as a valuable reference for researchers and practitioners.
8.3 3D Occupancy Prediction
3D Occupancy Prediction has emerged as a critical complement to Bird's Eye View (BEV) representations in autonomous driving perception. While BEV provides an excellent 2D overhead view of the driving scene, it fundamentally lacks height information, making it unable to represent overhanging structures (bridges, overpasses), varying-height obstacles, or the vertical extent of objects.
3D Occupancy Networks discretize the 3D space around the vehicle into a voxel grid, predicting whether each voxel is occupied and, optionally, what semantic class occupies it. This provides a complete volumetric understanding of the environment.
Tesla FSD V13 Occupancy Networks 2.0: Tesla's FSD V13 includes an upgraded Occupancy Networks module (version 2.0) that predicts fine-grained 3D occupancy from camera-only inputs. This is particularly important for Tesla's vision-only approach, as 3D occupancy prediction from cameras alone requires sophisticated geometric reasoning that compensates for the lack of direct depth measurements from LiDAR.
New Methods in 2025:
- MambaOcc: Applies the Mamba state-space model architecture to 3D occupancy prediction, achieving competitive accuracy with significantly lower computational cost compared to Transformer-based approaches. The linear scaling of Mamba (versus quadratic for Transformers) is particularly advantageous for high-resolution 3D voxel grids.
- GaussianFormer3D: Uses 3D Gaussian representations for efficient and accurate occupancy prediction, leveraging the spatial efficiency of Gaussian primitives to represent occupied and free space.
- OccCylindrical: Proposes a cylindrical coordinate system for occupancy prediction, which better matches the sensor characteristics of rotating LiDAR systems and provides higher resolution at closer ranges where it matters most for driving safety.
8.4 Reinforcement Learning and Sim-to-Real Transfer
Reinforcement Learning (RL) for autonomous driving has received renewed attention following the success of DeepSeek-R1, which demonstrated that RL-only training paths (without supervised fine-tuning) can produce high-quality reasoning capabilities. This has sparked interest in applying similar pure-RL approaches to AD, where an agent learns to drive entirely through trial-and-error in simulation.
Sim-to-Real Transfer Challenges
The fundamental challenge of applying RL to autonomous driving lies in Sim-to-Real transfer: policies trained in simulation must perform reliably in the real world, where conditions differ from the simulator in numerous ways:
- Tire characteristics: Real tire behavior involves complex nonlinear dynamics (slip angles, temperature-dependent friction, wear patterns) that are difficult to model accurately in simulation
- Road surface conditions: Wet, icy, gravel, and uneven surfaces create friction variations that simple simulation models may not capture
- Aerodynamic disturbances: Wind gusts, drafting effects from other vehicles, and tunnel transitions create forces that are computationally expensive to simulate accurately
- Vehicle loading: Cargo weight and distribution affects vehicle dynamics significantly, especially for trucks, and varies between trips
Latest Advances
Dynamics-Decoupled Trajectory Alignment: A recent approach that achieves zero-shot transfer success by decoupling the high-level trajectory planning (which transfers well between sim and real) from low-level dynamics control (which is domain-specific). By training a trajectory planner in simulation and adapting only the low-level controller to real vehicle dynamics, this method avoids the need for the simulator to perfectly replicate real-world physics.
Latent Space Modeling: Rather than attempting to simulate the real world at full fidelity, latent space approaches learn a compressed representation of the driving environment in which both simulated and real observations can be mapped. RL policies trained in this shared latent space transfer more robustly because they operate on abstract features rather than raw sensory inputs that differ between simulation and reality.
Digital Twin-Based Transfer: Using high-fidelity digital twins of specific real-world environments (built from detailed 3D scans and sensor recordings) to create simulation environments that closely match the actual deployment context. This approach reduces the domain gap by making the simulation as close to reality as possible for a specific target environment, enabling more reliable sim-to-real transfer for localized deployments.
8.5 Generative AI Applications
Generative AI is transforming how autonomous driving systems are trained and validated by enabling the creation of synthetic training data, particularly for rare and dangerous scenarios that are difficult or impossible to capture in real-world data collection.
SynAD (ICCV 2025)
SynAD (Synthetic data for Autonomous Driving), presented at ICCV 2025, demonstrates that synthetic data can significantly enhance End-to-End autonomous driving models. By generating photorealistic synthetic driving scenarios and using them to augment real-world training data, SynAD achieved the lowest collision rate among compared methods in benchmark evaluations. This result validates the potential of synthetic data to improve safety-critical metrics, particularly for scenarios that are underrepresented in real-world datasets (e.g., pedestrian darting into traffic, multi-vehicle pileups, extreme weather).
Structured Scenario Generation
A systematic approach to generating training scenarios uses a 5-layer model (road network, traffic infrastructure, temporal patterns, agent behaviors, environmental conditions) combined with Foundation Models to produce rare and challenging scenarios. By specifying conditions at each layer and using generative models to fill in the details, this approach can systematically create edge cases that stress-test AD systems. For example, combining an unusual road geometry with adverse weather, aggressive nearby driver behavior, and sensor degradation creates a compound scenario that is extremely rare in real-world data but critical for safety validation.
Industry Adoption
Major AD companies have embraced synthetic data generation as a core part of their development workflow:
- Waymo uses generative models to create synthetic scenarios for adverse weather conditions (heavy rain, fog, snow) and rare edge cases that its fleet rarely encounters in its primarily sunny deployment cities
- Waabi (founded by Raquel Urtasun, formerly of Uber ATG) has built its entire development approach around simulation-first training, using highly realistic generative models to create training data that reduces the need for expensive real-world data collection
- Multiple other companies are incorporating synthetic data for night driving, construction zones, emergency vehicle interactions, and other challenging scenarios that require robust handling but occur infrequently in standard driving data