The biomarker validation gap in precision oncology
Real-world evidence is helping accelerate biomarker validation and clinical adoption in precision oncology.
Tobi: Fit for purpose starts with the decision. What are we trying to answer, who needs the answer, and what level of confidence do they need to act on it? One of the big misconceptions is that more data automatically means better evidence. In practice, more data can also mean more noise, more complexity, and more work if it is not connected to a clear question.
The simplest way to put it is that organizations need to be evidence-ready before they can really be AI-ready.
Tobi Oremulé, Head of Solutioning, BC Platforms
A dataset can be useful in one setting and still not be right for the question in front of you. For example, claims data might tell you when a therapy was administered and where, but it may not tell you why that therapy was chosen, what was tested before it, or what happened next in the patient journey. The point is not simply to collect more data. — it’s to reduce uncertainty for a specific decision-maker.
Kumar: I would add that fit for purpose is not only about whether the data exists. It is whether the right data can be identified, linked, harmonized, governed, and used in a way that can stand up to scrutiny. You need provenance, privacy-preserving access, compliance, and a structure that supports the decision being made. If those are not in place, the data may be available, but it still may not be usable evidence.
Tobi: A lot of teams start with the data they already have and then try to work out what they can do with it. The stronger approach is to start with the decision and work backward. Who is the evidence for? What are they trying to decide? What would change if they had a better answer?
That changes the whole evidence plan. It affects the study design, what data sources you need, what governance is required, and what kind of analytics will be appropriate. Starting with the decision helps teams avoid building evidence around what is convenient rather than what is actually needed.
Kumar: Evidence questions now come from many parts of the product lifecycle. They can come from clinical development, regulatory, HEOR, market access, medical affairs, or post-launch teams. Often, these stakeholders are asking different questions, but the underlying foundation can support more than one of them if it is designed that way from the start.
That is why we see value in moving away from one-off studies and toward evidence generation as a reusable capability. The question should not only be, “How do we answer this question?” It should also be, “How do we answer it in a way that helps with the next question?”
Kumar: The dataset has to answer the question, but it also has to give people confidence in the answer. That means looking at things like linkage, provenance, compliance, longitudinal depth, breadth across the relevant populations, and whether the data can reconstruct the patient journey in a meaningful way.
For many evidence questions, one dataset is not enough. You may need EHR data, claims, imaging, biomarkers, diagnostics, laboratory data, clinical notes, or other sources. The value comes from bringing the right sources together and making them coherent, traceable, and usable.
Tobi: A simple way to test it is to ask: good for what, for whom, and against what evidence bar? If a team can answer those questions clearly, they will have a much better sense of whether the dataset is actually fit for purpose.
Kumar: Organizations need to think in terms of a network, not a single source. First, you need to understand where the right patients are, which sites have the data you need, what data types are required, and what access controls apply.
In many cases, the evidence question requires multi-source, multi-modal data. That could include EHRs, claims, imaging, biomarker results, diagnostics, and other data types. It also means doing the feasibility work upfront and understanding consent, privacy, governance, and site requirements before you get too far into the analysis.
The important point is that this should not be treated as a one-time exercise. If you are going to keep asking questions over time, the model needs to support refresh, reuse, and repeatable analysis.
Kumar: Governance cannot be an afterthought. In a federated model, the technology can be deployed at the site level, within a trusted research environment. The data can be harmonized, anonymized, linked, and made analytics-ready locally, while the analysis can happen without unnecessary movement of the raw data.
Governance cannot be an afterthought. The data can be harmonized, anonymized, linked, and made analytics-ready locally, while the analysis happens without unnecessary movement of the raw data.
Narasimha Kumar, Global Head of Technology and Data Services, BC Platforms
There are different ways this can work. In some cases, aggregate outputs can be brought together centrally. In others, approved analytics can be sent to the local environment and only the results come back. Federated learning follows the same principle for AI models: the model can learn from distributed data without all of the underlying data having to move.
Tobi: That matters especially in regions where institutions need to keep control of their own data. The goal is to let organizations get the insights they need across different sites or countries, while still respecting local governance requirements. It creates a balance: the data holders keep control, and the evidence teams can still generate useful insights.
Tobi: This is one of the areas where I think the industry is leaving some of its greatest value on the table. Historically, evidence generation has been organized around individual projects: a regulatory study, an HEOR study, a market access study, and so on. Each one can require new cohorts, new governance processes, new mappings, and new workflows. That becomes time-consuming and costly when you have to repeat it again and again.
The organizations at the forefront are taking a different approach. They are building evidence generation as a capability. They create an evidence layer once and then leverage it again and again across teams and decisions. The question changes from, “How do I answer this study?” to, “How do I answer this study in a way that strengthens the next one?”
That’s where the economic value starts to become more compelling. You save time because you are not starting over each time with sites, governance, mappings, and workflows. We often tell customers not to measure the ROI only on the first study. Measure it by how much cheaper and easier the next one becomes.
Kumar: I agree. Most traditional RWE platforms have been designed to answer one study or one question at a time. But many real-world evidence studies are not static. They may include retrospective data, prospective refreshes, changing protocols, consent management, and new questions over time.
So reuse requires agility with compliance. If the network, governance, data assets, and patient journey are already in place, the next questions become incremental rather than entirely new builds. That is very important for pharmaceutical organizations that need to answer evidence questions across a product, a therapeutic area, or a portfolio.
Kumar: AI can add value in a number of places: classification, metadata management, computer vision, clinical NLP, analytics, and model training. But AI is only as good as the data, governance, and context underneath it.
You need a strong data foundation, common data models, traceability, and the flexibility to apply different models, tools, and analytic approaches as needs evolve. Trusted research environments should support that flexibility while maintaining security, compliance, and scientific rigor.
Tobi: The simplest way to put it is that organizations need to be evidence-ready before they can really be AI-ready. If the data is not governed, traceable, harmonized, and fit for purpose, AI will only multiply the problems that already exist. AI may help compress the time to an answer, but it does not automatically create a trusted answer.
Kumar: Readiness is about alignment. Evidence generation cannot sit in one department only. It needs alignment across medical affairs, commercial, HEOR, regulatory, technology, data governance, and other teams.
There are three major areas to assess: the data foundation, the technology environment, and the governance model. That includes data quality, data integrity, compliance, privacy-preserving access, consent management, AI governance, and the ability to support multiple stakeholder needs over time.
Tobi: Start with consequence. Where would better evidence change an important decision? That is not always the same as the biggest data gap. Some large data gaps sit under decisions that no one is likely to act on.
It’s also important to separate local gaps from foundational gaps. A missing dataset for a single study may be a local issue. But if you do not have a governance pathway, a harmonization layer, or a secure analytics environment, that will block many future questions. Those foundational gaps should be addressed early, but still in the context of specific decisions, sponsors, and timelines.
Kumar: Prioritization should also be guided by evidence quality. Teams need clear design principles for what makes evidence acceptable, defensible, and compliant. The goal is not just to move faster. It’s to move faster in a secure, compliant, and reusable way. That should be planned as a multi-year strategy, not a short-term project.
Tobi: One recent oncology initiative in Japan is a good example. Our client is a large global pharma. The original request looked like a set of separate evidence studies across gastric cancer and prostate cancer, with questions around biomarker testing, treatment sequencing, and patient outcomes.
As we worked through the discussions, it became clear that the larger need was not just to answer a few separate questions. The organization needed a more sustainable oncology evidence platform that could work across fragmented healthcare datasets, institutional requirements, and disease-specific evidence needs.
That changed the conversation. Instead of treating gastric cancer and prostate cancer as isolated studies, the focus shifted to building a reusable foundation that could support multiple oncology programs over time. Future questions could then build on the same governance pathway, technology environment, and evidence infrastructure.
Tobi: Pick one decision, build the foundation under it, and make sure the next question is cheaper than the last. The future belongs to organizations that treat evidence as part of the infrastructure, not as the output of one vendor engagement or one isolated project.
Pick one decision, build the foundation under it, and make sure the next question is cheaper than the last.
Tobi Oremulé, Head of Solutioning, BC Platforms
Kumar: From my perspective, build for reuse, long-term ROI, and agility with compliance. That is the way to build reusable evidence networks in a compliant and secure manner.
The path from fragmented real-world data to trusted, decision-ready real-world evidence starts with a shift in mindset. Organizations do not need more data for its own sake. They need reusable evidence foundations that connect the right data, governance, access model, technology, and expertise around the decisions that matter.
As evidence needs become more complex — and as AI becomes more embedded in evidence workflows — the organizations best positioned to lead will be those that can generate trusted evidence repeatedly, securely, and at scale. By designing for reuse from the start, teams can reduce reinvention, improve consistency, and make each future evidence question faster and easier to answer.
Discover how fit-for-purpose data, strong governance, and privacy-preserving analytics can help transform fragmented healthcare data into reusable evidence foundations.
Fit-for-purpose real-world data is data that is relevant, reliable, and appropriate for answering a specific evidence question and supporting a specific decision.
Reusable evidence foundations combine data, governance, analytics, and technology into a framework that can support multiple evidence-generation initiatives over time.
Federated analytics allow researchers to analyze data across multiple sites while keeping sensitive data under local control, helping organizations generate evidence while maintaining privacy and compliance.
Trusted research environments provide secure, governed settings for accessing and analyzing healthcare data while supporting compliance, privacy, and scientific rigor.
Organizations become AI-ready by establishing strong data foundations, governance processes, data harmonization practices, and traceable evidence-generation workflows before introducing AI capabilities.