The autonomous vehicle industry spent a decade chasing the wrong problem: hand-engineering rules for a single, simplified environment. That approach produced brittle demos, not scalable products. Alex Kendall and Wayve took a fundamentally different path—treating driving as a general-purpose embodied intelligence problem that must work across any city, any vehicle, and any condition. In this talk, Kendall traces Wayve’s evolution from Cambridge deep learning research to production systems now being deployed with Nissan and Uber across 500+ cities. He makes a practitioner’s case for why end-to-end learned systems outperform modular stacks, how vision-language-action models like LINGO enable driving with natural language reasoning, and what it actually takes to make embodied AI economically viable. The conversation moves beyond the usual autonomy hype into the hard trade-offs around data scale, simulation fidelity, safety assurance, and the unit economics of fleet deployment. For anyone building systems that must act in the physical world—not just predict—this is a grounded, technically rich look at what works, what still breaks, and where the frontier actually sits.

Key Takeaways

  • Why geofenced autonomy in simplified environments like Phoenix created a false sense of progress, and how Wayve’s general-purpose approach—driving across 500+ cities—forces the model to learn robust, transferable representations rather than overfitting to a single domain.
  • How LINGO, a vision-language-action model, jointly reasons over driving and language modalities, enabling the vehicle to drive while providing natural language commentary about its decisions—a step toward interpretable, auditable embodied AI.
  • The economic argument for end-to-end learned systems: modular stacks accumulate compounding errors and require per-sensor, per-environment tuning, while learned systems amortize engineering cost across data and compute scale.
  • Why large-scale, diverse camera data from existing vehicle fleets is the critical moat—and how Wayve leverages this to train models that generalize across vehicle platforms, sensor configurations, and geographic regions without per-location calibration.
  • The open safety assurance problem: how to validate a learned driving policy when the operational design domain is effectively unbounded, and why traditional deterministic verification methods break down for end-to-end neural systems.

Who should watch: Robotics and autonomy engineers, ML researchers working on vision-language-action models, and technical leaders evaluating the build-vs-buy economics of learned systems for physical-world deployment.

Why This Matters

Wayve’s trajectory mirrors a broader shift we’re tracking: the collapse of modular robotics stacks into end-to-end learned systems. As foundation models eat into perception, planning, and control, the unit of competition moves from sensor hardware and hand-tuned planners to data flywheels and training infrastructure—reshaping the economics of every physical AI vertical.

Watch the full video →