From Source Data to Regulatory-Grade Evidence: Five Insights from the 2026 ASA Biopharmaceutical Section Regulatory-Industry Statistics Workshop (RISW)
What the workshop reinforced about fit-for-purpose Real-Word Data, emerging methods, and the expertise needed to turn source data into submission-ready evidence.
Five insights I took away from RISW
1. Fit for purpose and data quality come before methodology or AI.
2. Data quality and fitness for purpose are related, but they are not the same thing. High-quality data can still be unsuitable for a particular regulatory question.
3. Missingness is not simply an empty field. Why information is absent can determine whether RWD can support the intended analysis and can affect the statistical approach.
4. New methods keep circling back to the data. Bayesian methods, digital twins, and AI create new possibilities, but they do not eliminate the need for fit-for-purpose, high-quality data.
5. “Human in the loop” is not enough. Regulatory evidence requires the right combination of statistical, epidemiological, clinical, data, standards, and regulatory expertise.
Last week I attended the 2026 American Statistical Association (ASA) Biopharmaceutical Section Regulatory-Industry Statistics Workshop (RISW) for the first time. The workshop brought together FDA statisticians, industry researchers, academics, data providers, and other experts to discuss topics ranging from real-world evidence (RWE) and estimands to Bayesian methods, digital twins, and, of course, artificial intelligence. It was encouraging to see how the field continues to evolve nearly a decade after passage of the 21st Century Cures Act, with its focus on accelerating medical product development and bringing innovations to patients more efficiently.
My background is in data standards and regulatory review, and more recently in RWE standards and RWD fit-for-purpose assessments, so I consider myself statistically adjacent. To do this work well, I need to understand what statistical reviewers and sponsor biostatisticians require to evaluate and analyze data supporting the safety and effectiveness of drugs and biologics. With RWE, it is not enough to know how to standardize data for submission. We also need to know whether the data being standardized can actually support the scientific and regulatory question the analysis is intended to answer. If they cannot, transforming them into a submission standard will not make the data reviewable.
Across very different sessions, I heard a remarkably consistent message: increasingly sophisticated methods cannot compensate for data that are not fit for purpose, are poor quality, or both. At the same time, there is tremendous promise in applying new methods to high-quality, fit-for-purpose data. The path from source data to regulatory-grade evidence depends on the scientific question, the data source, availability of key variables, missingness, study design, and the ability to provide traceability. Many pieces, along with expert judgment, must come together to meet the evidentiary standard required.
1. Fit-for-purpose, high-quality data come first
In a session on harnessing RWE in regulatory submissions, an FDA presenter discussed challenges in using RWD, whether in an RWD study or to supplement another clinical study. A recurring point was that data must be suitable for the regulatory question and sufficiently reliable and relevant for the intended use. That means understanding completeness, accuracy, provenance, traceability, and whether key data elements are captured related to efficacy and safety. Because treatment selection is usually not randomized in RWD studies, sponsors must also assess potential bias in both the data and the study design, including missing data, misclassification, measurement issues, immortal time bias, and confounding.
The session also underscored that the evidentiary burden for an externally controlled trial (ECT) should not be underestimated. In an ECT, outcomes in a treated group are compared with outcomes in a control group drawn from outside the trial who have not received the same treatment. Demonstrating that those groups are sufficiently comparable for a credible comparative analysis requires careful attention to data, design, and analysis.
In a later session, “Target Trial Emulation for Regulatory Grade RWE Generation,” FDA’s Dr. Hana Lee returned to the organizing logic of FDA’s 2018 Framework for its Real-World Evidence Program. The framework asks three questions when evaluating RWE:
1. Whether the RWD are fit for use.
2. Whether the trial or study design used to generate RWE can provide adequate scientific evidence to answer, or help answer, the regulatory question.
3. Whether the study conduct meets FDA regulatory requirements.
That sequence matters. Before asking which method to use, sponsors need to define the regulatory and scientific decision they are trying to support. Only then can they determine what evidence is required and whether a candidate data source contains the information needed to generate it.
RWD can be high quality in a general sense and still not be fit for a particular purpose. Conversely, RWD does not have to be perfect to be useful. It does, however, need to be adequate for the intended regulatory use and that means evaluating it against the specific question the evidence is intended to answer.
Dr. Lee pointed to FDA’s guidance on externally controlled trials and to ICH E9(R1), the addendum on estimands and sensitivity analysis, as useful references when thinking about fitness for purpose and comparability. The broader lesson was practical: define what must be comparable and measurable before deciding that a data source can support the study.
She also emphasized an order of operations that appeared repeatedly throughout the workshop: data first, then design, then analysis. Start with the best data available for the question. If the necessary information is not available or cannot support a credible comparison, a different study approach may be more appropriate.
In the same session, Dr. Shirley Wang presented an ECT example in which EHR-based RWE was used to replicate a randomized clinical trial (RCT). The conclusion: EHR-based RWE studies can closely replicate RCT results when data-fitness checks indicate that the data are fit for purpose. She highlighted three requirements:
· A question-specific assessment of data suitability.
· Alignment of the question with the study design.
· Evaluation of data fitness together with the study design.
Other emulations encountered different challenges, including rapidly evolving standard of care treatments that limited power and comparability with randomized-trial intent-to-treat analyses, as well as insufficient curation of needed variables. A Venn diagram in the presentation captured the problem neatly: question, design, and data each matter, but all three need to overlap to meet the FDA’s evidentiary bar.
Take-home message: EHR-based RWE studies are reliable when the data are demonstrably fit for purpose and data quality is sufficient for the intended analysis.
2. Data quality is necessary, but “quality” depends on the question
The session “Statistical Review and Inspection of Real-world Evidence in Regulatory Submissions: Challenges, Risks, and Emerging Solutions” brought FDA, sponsors, and a data provider together to examine practical and statistical challenges in reviewing regulatory submissions derived from RWD sources. Moderator Mayur Saxena, CEO of Droice Labs, asked why RWE appears relatively infrequently in regulatory approvals despite its use across many development programs. Dr. Lee’s succinct response was: “Usually it is because of the data.” She noted that data quality is getting better but recurring problems such as bias, missing data, nonconcurrent treatment groups, and uncertainty in outcome measurement persist.
Dr. Khaled Sarsour of Astellas commented that the many sources of bias in nonrandomized studies could be grouped into three categories:
1. Selection bias: Does the sample represent the population of interest?
2. Confounding: Are important prognostic factors adequately measured and addressed?
3. Measurement: Are relevant factors measured accurately and consistently across groups?
FDA guidance and other methodological frameworks can help sponsors design studies and develop transparent plans for addressing these sources of bias.
Another discussion focused on how to define and measure data quality. Louis Brooks of Optum noted that this is one of the hardest problems in working with data assets and observed that “quality boils down to a standard you put in place.” Dr. Lee, by contrast, returned to FDA’s concepts of reliability and relevance. The exchange reinforced an important point: a generic quality label is less useful than a clear standard tied to the intended analysis.
The discussion of completeness illustrated why definitions matter. From one data-provider perspective, a record may be considered complete if it accurately reflects what the provider did, even when a test was never ordered and therefore no result exists. For a regulatory analysis, however, that absent measurement may be consequential. If a lab value is needed to assess efficacy, establish comparability, or evaluate toxicity, its absence can limit or even preclude the intended analysis. Sponsors evaluating RWD need be aware of these distinctions and understand the potential impact on analysis.
3. Missingness is information, not just an empty field
The short course “Advanced Estimands and Missing Data Imputation in Clinical Trials” was particularly useful because it connected statistical thinking with data collection and data standards. The course addressed practical implementation following the ICH E9(R1) addendum on estimands. Under that framework, the treatment effect of interest is defined through five attributes: the treatment condition, target population, outcome variable, population-level summary, and strategy for handling intercurrent events. Importantly, an intercurrent event is not automatically considered the same as missing data. Treatment discontinuation, rescue therapy, treatment switching, withdrawal, or death can change how an outcome should be interpreted or whether a measurement remains relevant to the clinical question. Data are missing when information that would be meaningful for the specified estimand was not collected or is unavailable.
That distinction has direct implications for RWD. A blank value is not a sufficient description of missingness. Why is the value absent? Was the test not ordered or not performed? Was care delivered outside the captured network? Did an intercurrent event change the meaning of the measurement? Is the value missing at random, or is the variable systematically unavailable for a clinically important subgroup? These are different problems with different consequences for bias and analysis.
The course also highlighted the role of data standards in preserving information. Existing SDTM domains can capture events such as disposition changes and concomitant or rescue medications, while proposed ADaM approaches such as the OCCDS-ICE structure can consolidate intercurrent events and retain links back to source records. The larger lesson was that a data standards need to preserve the information required to understand and interpret the evidence, not merely produce a technically conformant dataset.
4. New methods keep circling back to the same foundation: the data
Digital twins, Bayesian methods, and AI/ML were among the workshop’s most forward-looking topics. Yet the discussion repeatedly returned to familiar questions about data fitness and data quality. The takeaway was simple: no statistical method or AI system can rescue data that are fundamentally unsuitable for the question being asked.
Digital twins: data harmonization can be harder than modeling
Presentations described digital twins as predictive model or virtual representations of a patient that simulates how they would respond to different treatments using baseline characteristics. Each patient gets a “virtual twin” under alternative treatment scenarios and machine learning predicts the outcome. Where a traditional clinical trial seeks to find the average treatment effect, the Digital Twin Approach asks what is likely to happen to this individual patient under each treatment option? Potential applications range from control-arm augmentation, patient enrichment, treatment optimization, to personal medicine at scale.
The practical constraints were familiar. One industry presentation emphasized that data harmonization can become the rate-limiting step, with endpoint alignment and integration of historical trial data consuming substantially more effort than the modeling. An FDA research example involving a synthetic control arm emphasized the need for sufficient high-quality training data and ongoing validation. An MD Anderson presentation made the point very clearly: different digital-twin uses require evidence that is fit for that specific use, together with an explicit assessment of uncertainty in the data.
Bayesian methods: external information has to earn its role
In an opening plenary session, James Travis of CDER discussed FDA’s January 2026 draft guidance, “Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products.” The guidance focuses on the use of Bayesian methods to support primary inference in clinical trials intended to support the effectiveness and safety of drugs. Bayesian approaches have been used in post market analysis for over a decade and the ICH E11A guideline, released in draft in 2022, provides a standardized, global framework for using pediatric extrapolation mentions Bayesian methods (along with frequentist methods). It discusses the use of external information and notes the process of determining a prior (a key element in Bayesian analysis) should begin with the identification and review of all the available relevant external information. Data quality and reliability need to be adequate for the type of regulatory decision to be informed by the analysis. Thus, before external RWD can meaningfully inform a prior or the degree of borrowing, it is necessary to determine how it was generated, what the data represent, whether key variables are comparable, what is missing, what transformations occurred, and what uncertainty and limitations these factors may place on the use of the data.
AI: automation requires validation
The AI discussions raised a related concern: transformation itself can change the evidence. In a session on AI/ML and patient-reported outcomes (PROs), the question was not simply whether AI could reduce workload. It was whether the patient’s original response, the construct being measured, and the resulting score remained semantically and analytically valid after AI-mediated processing.
The implication extends beyond PROs. If AI is used to extract variables from clinical notes, map terminology, classify outcomes, derive phenotypes, impute values, or automate transformations, validation should focus on the specific task, population, model version, and context of use. The transformation also needs to remain traceable from source to submitted evidence. There is no generic “validated AI” for regulatory evidence. Credibility is task-, population-, version-, and context-specific.
5. “Human in the loop” is not enough: the right expertise must be in the loop
A final theme cut across nearly every topic: regulatory evidence, especially RWE, is interdisciplinary. Statisticians help define questions, design studies, quantify uncertainty, and evaluate bias. Clinicians provide medical context. Epidemiologists connect the study question to populations, exposures, outcomes, and confounding. Data scientists and programmers operationalize those concepts in data. Standards experts help make the evidence traceable and exchangeable. Regulatory experts connect the work to the intended decision.
Several speakers raised the risk that AI can create the illusion that one discipline can substitute for another. A clinician can ask an AI system a statistical question; a statistician can ask it a clinical question; a programmer can ask it to infer meaning from a variable. The output may sound plausible while still missing the disciplinary context needed to recognize a faulty assumption.
For RWD programs, “human in the loop” is therefore not enough. The right subject-matter experts need to be involved at the right points in the evidence lifecycle. Early collaboration also matters between sponsors and FDA, particularly when a novel data source, external control, Bayesian approach, digital twin, or AI-enabled transformation could materially affect interpretation. The critical regulatory questions are less costly to answer before the data are purchased, the study design finalized, and the analysis run.
From source data to submission-ready evidence
I went to RISW to better understand the statistical thinking surrounding emerging methods. I came away with an even stronger appreciation for the work that needs to happen before the final analysis is run.
For organizations planning to use RWD in regulatory submissions, I see three connected challenges. First, determine whether the data are fit for the intended purpose. Second, determine whether data quality is sufficient for the intended regulatory use. Third, transform those data into analysis- and submission-ready datasets without losing the meaning, provenance, assumptions, and traceability that make the evidence interpretable. That transformation should also make missingness visible, preserve what is known about why information is absent, and support assessment of its potential impact on the analysis.
These challenges form an important bridge between data strategy and regulatory science. A fit-for-purpose assessment is not a generic data-quality checklist. It is driven by the scientific question, the study design, and the regulatory role the evidence is intended to play.
The methods discussed at RISW will continue to evolve. The more important question is simpler: can we show, from source through analysis and submission, why this data, this design, and these analyses provide adequate and well-controlled evidence for the decision we are asking FDA to make?
About the author
Ingeborg Holt is the founder of Orizaba Solutions, where she works at the intersection of real-world data, data standards, interoperability, and regulatory submission readiness. Her work includes fit-for-purpose assessment of RWD and transformation of complex healthcare data into traceable, analysis- and submission-ready evidence.