Nathan Benaich and Ali Madani bypass hype for a practitioner's deep dive into building a frontier AI company for biology. The core of the conversation is a critical technical trade-off: sequence-first versus structure-first protein design. Madani explains why Profluent chose a sequence-first approach, using a language model architecture that treats amino acid sequences as text, and how this decision impacts everything from data efficiency to the functional validation of designed proteins. He breaks down the scaling laws they've observed for protein language models, drawing direct parallels to LLMs but highlighting key divergences. The discussion then shifts to the business of AI-driven biology, with Madani unpacking the $2.25 billion deal with Eli Lilly. He clarifies the deal's structure, explaining how it's designed to align incentives for both discovery and the subsequent, expensive work of clinical development, moving beyond a simple licensing model. The conversation frames the entire field's evolution from 'discovery'—finding what exists in nature—to 'design'—engineering bespoke molecules with specific therapeutic functions, and the new moats this creates.
Key Insights
- The 'sequence-first' vs. 'structure-first' debate is a fundamental architectural choice: Profluent's sequence-first approach treats proteins as language, enabling scaling on massive, unlabeled sequence data before incorporating structural constraints, which Madani argues is more data-efficient.
- Protein language models exhibit their own scaling laws, but the 'compute-optimal' frontier differs from LLMs because the cost of generating and validating a single high-quality functional data point in a wet lab is astronomically higher than scraping text from the web.
- The $2.25B Eli Lilly deal is not a standard license; it is structured as a multi-stage partnership with milestone payments that specifically de-risks the capital-intensive process of taking an AI-designed molecule through IND filings and clinical trials.
- The shift from 'discovery' to 'design' means the key metric changes from identifying a binding target to engineering a complete therapeutic profile—including stability, immunogenicity, and manufacturability—all simultaneously optimized by the generative model.
- Functional validation remains the primary bottleneck. Madani emphasizes that generating a plausible protein sequence is easy; the real moat is building a high-throughput, iterative 'design-build-test' loop where wet-lab results directly and rapidly inform the next generation of models.
- Profluent's open-source release of a gene editor designed by AI was a strategic move to demonstrate a new paradigm, proving that a functional protein can be created entirely in silico and given away, shifting the value capture downstream to the therapeutic application.
Who should listen: ML researchers and biotech founders building generative models for scientific domains where data is scarce and validation is expensive.
Why This Matters
This episode is a masterclass in translating a frontier AI trend—generative design—into a defensible business. We track the shift from horizontal AI platforms to vertically integrated companies where the model is necessary but insufficient, and the real moat is the proprietary data flywheel from a tightly coupled wet-lab, a playbook Profluent is executing explicitly.