At BIO-Europe Spring 2026’s, a panel assembled under the title “The Art of Wayfinding with AI in Drug Discovery and Development” delivered one of the conference’s most data-dense examinations of how artificial intelligence is restructuring the operational and economic architecture of drug development. Moderated by Barnaby Pickering, Director at 59 North Communications, the session featured Aliza Apple, VP Catalyze360 AI and Global Head of Lilly TuneLab at Eli Lilly; Karen Akinsanya, Executive VP, Chief Biomedical Scientist, and Head of Discovery R&D at Schrödinger; Gabriele Corso, Co-Founder and CEO of Boltz; Jean-Philippe Vert, Co-Founder and CEO of Bioptimus; and Javier Nunez-Vicandi, Principal of Digital Medicine Strategy at Sofinnova Partners. The discussion arrived at a precise inflection point: eighteen months ago, the viability of generalized foundation models for life science remained genuinely uncertain — by March 2026, pharma traction with those models had exceeded even internal projections, and the question had shifted from whether AI works to how fast it will eliminate the structural assumptions that have governed drug development for decades.

The Cost Collapse Reshaping Innovation Economics

The discussion opened with a structural diagnosis of where the AI-in-drug-discovery field stands after twelve months of acceleration that most participants described as faster than their own forecasts. The Sofinnova representative framed the dominant dynamic in explicitly economic terms: the industry is experiencing a cost collapse in the unit of innovation for life sciences — a compression that, if it follows the trajectory observed elsewhere in technology, does not reduce total activity but rather triggers a massive expansion of it. The mechanism invoked was Jevons Paradox — the principle that as the cost of a productive input falls, consumption of that input increases rather than contracts — applied to the economics of biological hypothesis generation and early-stage asset validation. Healthcare representing approximately 18% of U.S. GDP operates today as a scarcity-constrained market; the Sofinnova view is that the AI-driven efficiency curve will shift it toward a demand-constrained market, representing what the firm characterizes as the greatest expansion of healthcare markets in history.

Within this framing, the past twelve months produced two concrete categories of progress. The first was the demonstrated capability of generalized foundation models — architectures trained on massive, diverse biological datasets rather than narrow single-task pipelines — to achieve performance levels across multiple life science applications simultaneously. Eighteen months prior, this generalizability remained unproven; by early 2026, multiple companies on the panel had moved beyond benchmark results into commercial partnerships translating model outputs into diagnostic tools, pathology decision-support systems, and patient stratification workflows. The second category was infrastructural: the Bioptimus co-founder identified LLM-driven code generation as a qualitative change in the operational efficiency of founding and running AI-native companies, compressing development cycles and reducing the human capital overhead required to build at frontier scale.

The Lilly representative anchored the pharma perspective in a philosophy the company has held for some time but is now executing with new structural models: buy what accelerates, build what differentiates. The $1 billion co-innovation lab established with Nvidia in the Bay Area embodies a departure from the binary build-versus-buy calculus — instead, it instantiates a dedicated co-development model where talent from both organizations works jointly on shared problems rather than exchanging deliverables across an arm’s-length collaboration boundary. The outputs of that joint effort, spanning new model development and robotics applications, are designed to feed directly into Lilly’s TuneLab platform and extend its capabilities to the broader ecosystem rather than remaining proprietary to Lilly alone.

Platform Architectures and the Federated Learning Infrastructure

Lilly TuneLab emerged as the session’s most fully-described platform case study — a model that inverts the conventional pharma data-sharing dilemma by making Lilly’s proprietary machine learning models available to external partners in exchange for training data contributions, without requiring any underlying data to move, commingle, or become visible across the network. The mechanism is federated learning: rather than centralizing data and training a model in one location, the model is distributed to each partner’s local node, trained on that node’s proprietary data, and only the updated model weights — not the data itself — are returned and aggregated across the network. The result is a continuously improving shared model trained on the collective signal of the entire network, with each participant’s raw data remaining entirely within their own infrastructure. The technique was originally developed by Google more than a decade ago and has been deployed at scale in industries including consumer technology and finance, but had not achieved broad adoption in pharmaceutical R&D prior to TuneLab’s architecture.

The Lilly representative made the underlying strategic logic explicit: even a 150-year-old company with substantial historical datasets must accept the mathematical reality that there will always be more data outside Lilly than inside it. TuneLab is the operational response to that asymmetry — a mechanism to systematically internalize external signal without surrendering the privacy, competitive sensitivity, or regulatory standing of any participant’s proprietary data. The platform currently hosts over 30 models predicting ADME properties, with compute requirements modest enough to run on a standard CPU in two to three minutes across several hundred molecules simultaneously — making broad chemical space screening feasible before concentrating computational investment in high-accuracy physics-based tools for candidate-level decisions.

The Schrödinger representative confirmed an active partnership with Lilly through TuneLab, specifically targeting the integration of federated learning models into agentic, cross-functional workflows where scientists across disciplines — not just computational chemists — can access multimodal predictions without specialized infrastructure overhead. The underlying premise is that the barrier to running complex physics-based and ML-based predictions has already collapsed within Schrödinger’s own platform over the past year: tasks that previously required domain expertise to configure and execute are now accessible to essentially the full scientific team. The next phase — which the Schrödinger representative forecasted within twelve months rather than five years — is multimodal model access layering chemistry predictions, ADME data, structural biology, and clinical context into unified, agentic workflows.

Boltz presented a complementary architecture. The company has built large-scale biomolecular open-source models used across all major pharmaceutical companies and thousands of biotech organizations — with Lilly and Schrödinger among confirmed users — and has recently layered a commercial platform, Boltz Lab, on top of that open-source foundation, operating on usage-based pricing. The Boltz co-founder challenged the assumption that open-source equals free: even accessing the open-source models requires substantial infrastructure to configure them for scientific use and meaningful compute to run them. Boltz Lab’s value proposition is that centralized, optimized deployment at scale makes model access cheaper than the compute cost a company would incur running the open-source versions independently, while eliminating the integration and configuration burden for scientists who are not ML engineers.

The clearest proof-of-concept the Boltz representative cited was in de novo antibody design. Twelve months prior, initiating an antibody design campaign against a novel target would default to immunization or yeast display — wet-lab experimental approaches with defined throughput ceilings. Boltz and others have since demonstrated large-scale computational pipelines capable of designing novel antibodies with superior binding affinities and developability profiles compared to immunization and yeast display outputs, validated across multiple targets. The same trajectory is beginning to emerge for small molecules, where de novo design and hit-discovery-to-optimization pipelines are generating measurable traction with users at scale.

Bioptimus occupies the highest-abstraction layer of the stack — training foundation models on biology itself rather than on specific molecular or chemical tasks. The company’s models are trained on datasets spanning millions of patients, with histopathology images generating thousands of image patches per patient and total training scale reaching the order of trillions of tokens — a volume the Bioptimus co-founder placed as not far from frontier LLMs. The data weighting methodology mirrors empirical findings from LLM training: rather than filtering aggressively for high-quality data, the models benefit from including as much data as possible with differential weighting by indication, institution, and scanner — acknowledging that heterogeneous data from thousands of hospitals produces substantially more robust and generalizable models than data from a single high-quality source. The company’s AGE model, focused on histopathology, is already translating from benchmark performance into commercial partnerships with pathology tool developers.

The Physics Versus Foundation Model Debate and Its Real Implications

The session’s most technically substantive exchange centered on whether physics-based simulation and data-driven foundation models are complementary or fundamentally competitive — a debate with direct consequences for capital allocation, platform strategy, and the projected timeline to meaningfully improved clinical success rates.

The Schrödinger position, grounded in thirty-six years of physics-based computational chemistry, is that the two approaches are not in competition but are structurally interdependent. Physics-based methods — specifically free energy perturbation and related first-principles techniques — deliver accuracy from the ground up but are computationally expensive and slow. Machine learning delivers speed and scale but requires training data large enough to generalize across the chemical and biological contexts relevant to drug design, which does not yet exist comprehensively. The current optimal workflow uses physics-based methods to generate high-accuracy synthetic training data within specific protein structural contexts, which then feeds ML models capable of screening much larger chemical spaces at low compute cost — with the physics-based tools reserved for the final candidate-level interrogation before committing to synthesis expenditures of several thousand dollars per molecule. The Lilly representative confirmed this is precisely the logic underlying the Lilly-Schrödinger TuneLab integration: ML models scan chemical space broadly, physics-based tools validate the specific molecules selected for synthesis.

The Boltz co-founder took a harder position: physics-based simulation operates on assumptions — simplifications of quantum mechanical reality that introduce systematic error — while foundation models trained on sufficient data break those assumptions entirely and are guided by empirical signal rather than theoretical priors. On specific targets and modalities where training data is now adequate, foundation models already outperform physics-based approaches on binding affinity and developability prediction. The trajectory, in the Boltz view, is toward foundation models replacing simulation-based free energy calculations as the primary accuracy tool, with simulation retained primarily for synthetic data generation to bootstrap models in data-sparse contexts.

The Schrödinger representative’s counterpoint was structural rather than technical: the approximately one thousand discrete decision points between target nomination and a molecule entering human trials — a figure offered as a deliberate simplification — cannot be addressed by any single model architecture. The clinical success rate of assets entering and completing Phase 1 has not yet measurably shifted despite years of AI-enabled discovery, a point the panel acknowledged as a shared source of impatience. The implication is that the bottleneck is not in any single step where AI has demonstrated capability, but in the compounding complexity across the full development chain.

Investment Frameworks, Stack Disaggregation, and the Road Ahead

Looking ahead, the Sofinnova framework for evaluating AI-native life science companies resolves to three primary variables: data, talent, and compute. The logic is that the current paradigm — one foundation model serving many applications, displacing the prior one-model-one-application architecture — concentrates value in teams with unfair access to all three simultaneously. Bioptimus was cited as a direct application of this framework: a world-class research foundation in frontier AI methodology, combined with proprietary data access through institutional partnerships and preferential compute access through hyperscaler relationships. The investment thesis is that these inputs, combined with scaling laws demonstrating predictable performance improvements as data and compute increase, allow a forecast of model capability trajectories rather than requiring a point-in-time capability assessment.

The Sofinnova representative identified two distinct company archetypes succeeding in the current environment. The first is the frontier research lab — a talent-dense team with unfair access to data and compute, building capabilities that scale to seven- and eight-figure partnership agreements rapidly after capability milestones are reached, but requiring substantial capital injection before those milestones arrive. The second is the fast commercial company — one that leverages capabilities built by others, prioritizes immediate go-to-market traction, and stays exactly one step ahead of customer demand rather than two, building business model and channel expertise rather than foundational model capability. The portfolio logic is to hold positions across both archetypes rather than selecting one.

The broader structural shift the panel converged on is the disaggregation of what has historically been a vertically integrated stack in pharma and biotech. The monolithic company that owns data, models, orchestration, and asset development simultaneously is giving way to specialized layers — data infrastructure, foundation model development, application-layer orchestration, and asset-holding entities — with value creation concentrated in the companies that achieve genuine differentiation at one layer rather than average capability across all of them. The Lilly representative identified biological insight, novel chemistry, and modality innovation — including the combinatorial complexity of ADC design, where tools like Boltz’s remain modality-specific rather than multi-component — as durable sources of differentiation that the current wave of AI tools does not eliminate. Schrödinger and similar platforms did not end biotech; the current wave will accelerate capital efficiency without removing the requirement for a fundamentally differentiated scientific hypothesis.

The panel’s final convergent message for organizations navigating adoption was precise: AI capability is not a point-in-time assessment but a wave with a documented acceleration curve — protein structure models that failed to achieve adequate performance six months before the session had, by the time of the panel, demonstrated capabilities that would have been dismissed as aspirational at their prior benchmark. The operational implication is that capability evaluations must be continuous rather than episodic, and that therapeutic areas or modalities where AI tools appeared inadequate on last assessment may have crossed performance thresholds since. The underlying driver of the entire landscape remains unchanged: ten million people die of cancer annually, most complex multifactorial diseases remain mechanistically undercharacterized, and the models being built in 2026 are being trained on data volumes — millions of patients, trillions of tokens, thousands of institutions — that no individual human expert will ever approach.

Website |  + posts

Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.