Google DeepMind researcher Pushmeet Kohli and Biohub researcher Sal Candido said AlphaFold did not solve protein folding, in a panel moderated by Brandon Anderson. The discussion framed the breakthrough as only the beginning for AI in biology, with static structure prediction leaving dynamics, disorder and whole systems unresolved. The central claim was that scaling compute and data alone will not produce models that understand cells and disease.

AlphaFold Was a Start: DeepMind and Biohub on What Biology AI Still Lacks

Scaling, data quality and AlphaFold limits

Candido argued that scaling laws do not exist everywhere by default and finding them is the main work. A scaling law appears when more compute and more data reliably improve results, turning research into an engineering task. That outcome depends on architecture and on data that contains the right information and statistics, since a model can only extract and generalize what its training data holds.

Kohli reframed the bitter lesson as a warning against rigid professional identity, recalling the original discussion at DeepMind. Modelers should not assume that a dataset is fixed, and data teams should not assume that modeling is fixed. The problem comes first, and the team should invest in modeling, data generation or scientific expertise according to what the problem requires.

The panel said researchers often optimize for available data rather than the most important scientific questions. Candido described training protein language models on metagenomic sequences, much of which may not represent real whole proteins, yet performance on designing and understanding real proteins improves. The risk is simply scaling data that is easy to generate, instead of asking what data is needed and building it with the community.

What this means for applied biological AI

AlphaFold relied on handcrafted architecture and scientific intuition, and the speakers said that approach still matters where data is limited. Good data was presented as more important than simply having more data, with inductive biases helping models use biological structure efficiently. For companies, this means biological AI projects need joint planning of measurement, modeling and domain expertise rather than buying compute alone.

Moving from individual proteins to predictive models of living systems, including a virtual cell, will require fundamentally different datasets. The discussion pointed to protein dynamics and disorder, protein design, richer representations from cryo-EM micrographs, and hidden knowledge inside protein language models that has not been unlocked. For drug discovery, the relevant question was when AI can deliver 10x to 100x acceleration, not incremental improvement.

The marker to watch is whether teams demonstrate scaling behavior on newly collected biological data and convert it into system-level prediction. That includes calibrated uncertainty and trustworthiness rather than full interpretability, and possibly AI systems that interpret other models better than humans. Biohub frames its mission as curing disease through 10x breakthroughs, so confirmation will be new datasets, community use and measurable gains in design and discovery.