ML6 • Blog

AI In Life Sciences Is Struggling. The Issue: Lack Of Connected Data.

Geschrieben von Jan Rommel | 18.08.2026, 13:52:55

Key takeaways

1. The real bottleneck in AI-driven life sciences is disconnected, unharmonized data and not model capability. No model can reliably reason over fragments.

2. A connected data layer achieved by self-healing pipelines, harmonization, knowledge graphs, and multi-agent synthesis is durable infrastructure that compounds in value, while models get swapped out.


3. Today's best model is tomorrow's baseline. The connected data layer underneath is the lasting asset, growing sharper as more experimental, clinical, and real-world data flows in.


4. Once the data layer is connected and trustworthy, an agentic layer on top turns fragmented knowledge into decisions people can act on.

Everyone is watching the wrong part of the drug-discovery pipeline

Data silos, not a shortage of intelligence, hold AI in life sciences back. Current headlines focus on models that fold proteins, design molecules, and write papers in weeks. People are right to be excited. But the part that decides whether any of this reaches a patient is quieter, less glamorous, and almost entirely unsolved: the data layer underneath. What the field lacks is connected, harmonized, queryable biological and clinical data. Fix that, and everything above it moves faster. Ignore it, and the smartest model in the world keeps reasoning over fragments.

The industry needs this challenge solved before anything else moves. Structured, queryable data is an important resource that autonomous research agents and groundbreaking models depend on. It is not an adjacent task. It is the foundation on which AI success in life sciences gets built.

Why is the data flow the real bottleneck, not the model?

Model architecture gets the headlines, but a growing body of evidence shows that the quality, curation, and relevance of the data matter at least as much to performance, and often more. A model trained on data that is biased, mislabelled, or biologically uninformative produces biased, mislabelled, or biologically uninformative outputs, no matter how sophisticated its architecture.

A growing body of research formalizes this as a shift from a model-centric to a data-centric approach, where the systematic engineering of data, rather than the next clever model tweak, drives results. In life sciences, the evidence is concrete. A systematic study of virtual drug screening found that deliberately improving the properties of the chemical data, including its representation, quality, quantity, and composition, was the productive lever, not by applying more sophisticated models.

A broader review of AI in drug discovery goes further. Performance gains attributed to fancier model architectures often vanished once the training data was cleaned, of known quality artefacts. This suggests that expert-driven data curation, not algorithmic sophistication, drove the reported wins.

The same story is now playing out at enterprise scale far beyond the lab. A 2025 MIT report found that 95% of enterprise generative-AI pilots delivered no measurable business impact. The cause was not model quality or regulation, but a learning gap: fragile workflows and tools that fail to integrate with, retain, or adapt to the organization's own data.

Disconnected data is not a life-sciences peculiarity. It is the dominant reason AI stalls everywhere. It just happens to be most expensive here. Pharmaceutical research has the same structure, at a vastly larger scale, and with one critical difference: the inputs are not sitting in one place waiting to be fed in. They are scattered.

Critical knowledge lives across lab informatics systems, omics repositories (large-scale datasets from genomics, proteomics, and related fields), clinical systems such as Electronic Data Capture (EDC) and Electronic Health Records (EHR), and an ever-growing body of external literature. Each silo speaks its own dialect, because different teams built each one for a single purpose in a different decade. Nobody chose this fragmentation. It is simply the by-product of an industry where discovery, toxicology, clinical operations, regulatory, and manufacturing each demand their own specialists and their own tooling. Every time a program crosses from one function to the next, some context is lost in translation. Those translation losses, accumulated across a decade-long program, are where the time and money quietly drain away.

Most organizations have the data. Very few have the data flow.

Where does disconnected data actually hurt in life sciences?

This is not an abstract data-governance complaint. The cost shows up at the exact moment that matters most for patients and at the moments that decide whether a drug program succeeds or fails.

When expert decision-making stalls before it starts

Some of the highest-stakes decisions in oncology happen in what the field calls a tumor board: a room (or video call) where oncologists, pathologists, geneticists, and radiologists sit down together to weigh all the evidence for a single patient and agree on a treatment path. It is one of the few moments in medicine where every relevant discipline looks at the same case at the same time.

ML6 built an AI-supported tumor board with the Princess Máxima Center for pediatric oncology, where an international panel of leukemia and lymphoma specialists meets weekly to review complex cases. What struck us was not the clinical complexity — it was what happens before the discussion even begins.

To prepare a single case, specialists manually gather pathology reports, genomic profiles, imaging, and the latest relevant literature, each locked in a different system, each in a different format. Hours of senior clinical time go into assembling the picture before anyone can start reasoning about it. The bottleneck is not the expertise around the table. It is the work required to get the right information in front of that expertise.

The same pattern repeats across the industry. Validating a disease target, designing a preclinical study, generating real-world evidence (RWE) for post-market surveillance: all of them stall in the same place, waiting on data that exists but cannot be reached, trusted, or combined. Integrating these heterogeneous modalities, each with its own scale, noise structure, and biological logic, is arguably the most demanding challenge in the field, and the one where generic tools fall hardest.

At the point where two-thirds of drugs fail: finding the right patient

The same fragmentation carries a far larger price tag upstream, in the clinic. Around 90% of drugs that enter clinical trials fail, and much of that failure has little to do with the chemistry. Marc Tessier-Lavigne, formerly head of research at Genentech and now chief executive of the AI-native biotech Xaira, makes the point bluntly: a large share of trials collapse even when the science is sound, because clinicians didn’t correctly identify the patients who would actually respond. The target is valid, the molecule does its job, and yet the trial still reads as a failure because of the “simple” reason that the right population could not be pinned down.

The published evidence points the same way. Among novel therapeutics that fall over in late-stage trials, the single biggest reason is inadequate efficacy rather than safety, and trials that use biomarkers to pre-select likely responders clear their endpoints at close to double the rate of trials that do not (10.3% versus 5.5%). Put simply, identifying who a therapy is for is a data challenge long before it becomes a biological one. It requires data sources to be integrated closely enough to reveal the responding subpopulation, something that remains difficult when systems are siloed.

ML6's empathic AI work in patient care is built on a conversational-AI reference architecture that lets healthcare and life-sciences organizations tailor how they engage each patient, rather than addressing everyone as a single undifferentiated cohort. That kind of personalization is only ever as good as the connected data sitting beneath it.

Faster design, same downstream bottleneck

Artificial intelligence is accelerating the early stages of drug development, shortening the path from target identification to a clinical candidate. As Dr. Jean Philippe Vert noted in an interview, this work was always possible, but timelines have shifted dramatically, from five to seven years in the past to less than two years today. But a designed molecule is a file. Turning it into a medicine means manufacturing it at regulator-grade purity in facilities that took years and billions to certify. That capacity is concentrated among a handful of contract manufacturers. When Novo Holdings paid $16.5 billion for Catalent and pulled that production off the open market, it underlined the point: as design gets cheap, value shifts to the parts of the chain that are still scarce, like trials, manufacturing, and the connected data that coordinates them.

That coordination is exactly what fragmented systems cannot provide. Validating a target, designing a trial, generating post-market evidence: each step depends on data that exists somewhere in the organization but cannot be reached, trusted, or combined fast enough to keep pace with what AI now produces upstream.

What does a connected data layer for life sciences look like?

When we say ML6 builds the connective tissue for AI in life sciences, we mean something specific and buildable, not a slide-deck abstraction. Here is what it is made of.

Self-healing data pipelines. Biological and clinical data is messy, and formats break constantly. Pipelines that need a human every time a schema shifts do not scale. We build pipelines that detect, flag, and recover from issues automatically, so the flow stays reliable without a person babysitting it. A silent data error is not a glitch; it is a risk to a decision. However, as also highlighted in a recent blog post, human oversight remains essential, especially in highly regulated environments.

Harmonization. Connecting two systems is easy. Making them mean the same thing is the hard part. Harmonization is the layer that reconciles vocabularies, units, identifiers, and formats so that a genomic record and a clinical note can sit in the same query and actually agree on what a patient, a gene, or a measurement is. The standard framework for achieving this is adopting guiding principles to make data Findable, Accessible, Interoperable, and Reusable (FAIR). The hardest version of this is not technical; it is conceptual. Engineering teams describe data in terms of platforms, schemas, and pipelines.

In contrast, clinical and life-science teams talk about the same data in terms of assays, conditions, and patient-level provenance. A "feature" in a machine-learning pipeline and a "biomarker" in a clinic can be the same measurement carrying completely different assumptions depending on how it was obtained, and what it licenses you to decide. This vocabulary gap is a recurring source of failure in cross-disciplinary AI projects: systems that are sound in one frame can be meaningless in another. Doing it well means reconciling mental models, not just file formats.

Knowledge graphs. Once sources are harmonized, a knowledge graph turns them into a connected structure where relationships between a mutation, a drug, a trial, and an outcome become first-class and queryable. This is what converts a pile of harmonized records into something a human or an agent can reason over.

Multi-agent synthesis. On top of that unified layer, specialized agents can each take on a slice of a complex question, analyze the relevant evidence, and synthesize their findings into a structured output. This is exactly the architecture behind the Tumor Board work mentioned earlier. A team of virtual specialists, each grounded in a different slice of medical knowledge, analyzes the patient case and scientific literature ahead of the meeting and produces a structured insights-and-options report, so the live discussion starts from synthesis rather than assembly. ML6 has since generalized this pattern into reusable multi-agent building blocks, because the same coordination problem shows up well beyond oncology.

The pattern underneath all four is the same. Fragmented sources become a unified intelligence layer. That layer is the substrate every other ambition in AI-driven research quietly depends on.


What connected data makes possible for AI in life sciences

Here is the part that gets people excited, and it is worth being clear about why it works. Once the underlying data is connected and trustworthy, you can add an agentic layer on top that turns the unified layer into something a scientist can actually talk to.

Start with what already exists. We built a Clinical Co-pilot, a secure application in which specialized agents read a patient's unstructured clinical notes, interpret a lengthy insurance form, and populate it for a clinician's final review. This removes the kind of assembly work that burns expert time and delays patient access to treatment: hours of manual extraction from messy records. It works because the agents sit on top of structured, interpreted clinical data, not raw document soup.

Now extend the same idea. Imagine a scientist queries in plain language: show me every internal study and external trial that touched this target, and summarize what we learned about toxicity. Behind that one question, agents retrieve across lab systems, omics repositories, clinical records, and the literature; reconcile what they find; cite their sources; and hand back a synthesis in minutes rather than the days it would take a person to chase it down. The same foundation supports a literature-monitoring agent that flags new publications relevant to a live program, or a trial-design assistant that benchmarks a proposed study against comparable ones using real-world data.

Notice the dependency in every example. None of it is reliable if the layer underneath is fragmented. An agent answering questions over inconsistent identifiers and unharmonized records does not save time. It manufactures confident, well-formatted errors, which in this domain are worse than no answer at all. The agents are the value the reader sees. The data layer is why that value can be trusted. That is the whole argument in miniature.

Why should you fix the data layer before investing in models?

There is a strategic reason to start at the data layer rather than the model layer, and it is not just that the foundation has to come first.

The data layer is durable. Models keep improving, and they keep getting swapped out as better ones arrive. The connected, harmonized, governed data layer underneath them does not get thrown away with each new model generation. It compounds. Every result that flows through it makes the next query richer. That is the virtuous circle again, and in plain commercial terms it is the part of the stack with the longest useful life and the strongest case for ongoing maintenance and extension. However, companies should be careful when using third-party tools with sensitive data. Companies should assess the buy-versus-build trade-off.

This is also where trust is won or lost. In life sciences, one flawed input, one mismatched identifier, or one silently dropped record can compromise a decision that affects a patient. The failure mode is worse than it sounds, because AI does not fail loudly here. Feed bad or disconnected data to a capable model, and it produces an answer that looks rigorous and sounds authoritative but is confidently wrong. The worst part is that it is indistinguishable from a correct answer to anyone without the domain knowledge to check it. Without a trustworthy substrate underneath, AI does not democratize science so much as democratize its appearance: plausible, well-formatted output that nobody is equipped to challenge. Getting the data layer right is not a precondition for impressive demos. It is a precondition for being allowed to operate at all. As stated in one of our latest blog posts: We do not need AI experiences that feel faster or smoother; we need experiences that help us make better decisions with confidence.

The honest version: what connected data can and cannot promise

We will not claim that connecting your data silos will produce a new drug by next quarter. On short timelines, the skeptics are correct: AI medicines do not fall from the sky. Take rentosertib, one of the first compounds with both an AI-discovered target (TNIK, a protein implicated in idiopathic pulmonary fibrosis) and an AI-designed molecule. It shows how fast AI can move through discovery and the preclinical phase. But it also shows the limit: it only entered a Phase III trial in 2026, years after that early work, and only time will tell whether it clears the clinical-development stages that decide whether a drug actually reaches patients.

That is the honest shape of the opportunity. AI compresses the part of the pipeline that was already one of the cheapest and shortest. The long, expensive, decisive part still runs on trials, manufacturing, and, underneath it all, connected data. The organizations that win the next decade of AI in life sciences are the ones that fix their foundational data flow early, while others stay captivated by the models themselves. The substrate may be unglamorous, but it is precisely where the real advantage sits.