From Digital Health Technologies to Regulatory-Grade Evidence: Key Insights for Drug Development

Insights from the FDA and Duke-Margolis workshop on statistical considerations for digitally-derived endpoints in clinical trials

‍ ‍

Introduction

‍ ‍

Digital health technologies (DHTs) are defined by FDA as "a system that uses computing platforms, connectivity, software, and/or sensors for healthcare and related uses." They span a wide range of tools, from general wellness applications to complex medical devices, and include wearable sensors, continuous glucose monitors (CGMs), and accelerometers.

‍ ‍

Traditional clinical endpoints are generally collected at scheduled visits under conditions that are highly controlled and standardized through a protocol, staff training, and monitoring. Digitally-derived endpoints instead originate from near-continuous measurements collected while participants go about their daily lives.

‍ ‍

This capability has attracted interest from sponsors, patient advocates, and regulators because it has the potential to revolutionize how data are collected in clinical trials. Rather than limiting data collection to clinical visits, DHTs may reduce patient burden, including the need for clinic visits, which can lessen geographic and logistical barriers and enable participation by a wider range of patients. DHTs can produce higher-quality, granular data spanning weeks, allowing changes to be measured that may be missed by intermittent, scheduled clinical assessments. They can also reduce recall bias in patient-reported outcomes. This type of granular, longitudinal data collection may also fill knowledge gaps in disease areas, support earlier disease detection, and facilitate the development of new digital endpoints.

‍ ‍

However, the same characteristics that make digital data attractive also make them statistically challenging.

‍ ‍

DHT data can contain greater biological variability and measurement noise than clinical assessments, complicating analysis. The data are collected at high frequency and volume, yielding a rich, longitudinal profile for each participant, but they often require specialized analytical pipelines for processing. Endpoints are derived through complex functions, and existing validation benchmarks are often unavailable. Missing data can be difficult to identify, and, even when identified, the reason for missingness is not always explainable, complicating the assumptions underlying statistical analysis.

‍ ‍

The release of FDA’s Digital Health Technologies for Remote Data Acquisition in Clinical Investigations Guidance in 2023 was a milestone for the use of DHT data to support regulatory decision-making. It showed how DHT data collected from patients during a clinical investigation could be included in a submission, but there is more work to be done. The industry's challenge is to build the scientific and regulatory infrastructure needed to evaluate DHT data using methods that are thorough, appropriate, transparent, and reliable enough to support new drug and biologic approvals.

‍ ‍

The workshop's regulatory focus

‍ ‍

The workshop, Statistical Considerations for Digitally-Derived Endpoints in Clinical Trials, was convened by the Duke-Margolis Institute for Health Policy under a cooperative agreement with FDA to examine progress in using DHTs to support regulatory decision-making for new drug approvals and identify the remaining challenges to their more effective, efficient, and predictable use.

‍ ‍

It brought together FDA statisticians and clinicians, academic researchers, industry representatives, and patient advocates to examine the development, validation, analysis, and interpretation of digitally-derived endpoints for regulatory decision-making. The workshop examined progress in the use of DHTs and identified remaining challenges, with the aim of helping FDA determine its next steps toward unlocking the full potential of DHTs to support medical product development and innovation.

‍ ‍

Mark Levenson, Director of CDER's Office of Biostatistics, described the central challenge of deriving endpoints from DHTs in his opening remarks: How do we translate continuous, multi-dimensional data from a DHT into a clinically meaningful endpoint?

‍ ‍

Key insights

‍ ‍

Start with what matters to patients, not what the technology can measure

‍ ‍

Perhaps the most important discussion concerned endpoint selection itself.

‍ ‍

The workshop repeatedly emphasized that DHT capabilities should not drive endpoint selection. A process similar to the one used to develop patient-focused outcome measures for clinical trials, as described in FDA's patient-focused drug development (PFDD) framework, should be used with DHTs. Investigators should begin by understanding the disease or condition, identifying meaningful aspects of health, identifying the concept of interest, and defining the context of use. Only then should they determine whether a DHT and digitally-derived measure can appropriately assess that concept.

‍ ‍

Focusing on what is important to patients is especially important because a single device may generate many candidate metrics. For an itching disorder, for example, a sensor might characterize frequency, duration, or intensity. Analytical convenience should not be used to determine which of these represents the outcome most meaningful to patients.

‍ ‍

A well-justified endpoint rationale should address the concept of interest, trial objective, endpoint's role in the study, intended indication, fitness of the DHT and measure for the planned trial, importance to patients, and the endpoint's strengths and limitations.

‍ ‍

A digitally-derived endpoint is the output of a process, not simply a sensor

‍ ‍

One of the most important concepts stressed by subject-matter experts at the workshop was that a digitally-derived endpoint should not be viewed simply as the output of a sensor, but as the output of a process.

‍ ‍

Using an accelerometer as an example, Vadim Zipunnikov from Johns Hopkins Bloomberg School of Public Health illustrated the many decisions that occur between data collection and a final endpoint: device selection, body placement, sampling rate, raw signal processing, quality checks, epoch segmentation, non-wear identification, definitions of valid days, sleep/wake classification, activity and posture identification, missing-data handling, aggregation, and, finally, endpoint derivation. In a separate presentation, Taehyun Jung, a Senior Statistical Reviewer in CDER, focused on just part of that process, going from the sensor data to the final endpoint (Figure 1), and noted that error could be introduced at any step along the way.

Figure 1: DHT Data Processing Pipeline

‍ ‍

Two teams could begin with similar raw data and make different decisions at these stages, potentially producing different results. The implication is important: the provenance of the final endpoint includes not only the endpoint measured, but the entire processing pathway.

‍ ‍

Validation is context-specific

‍ ‍

There is currently no gold standard against which to measure a DHT; the appropriate standard depends on the context of use. A measurement that performs well in a clinical setting may not perform well in a real-world setting. Validation does not automatically transfer from one setting to another. Analytical validation asks whether a digitally-derived measure accurately and reliably represents its intended signal under specified conditions. Favorable performance under one set of conditions does not, by itself, establish generalizability to another population, setting, or context of use. The overall analytical validation conclusion is built from determinations about each source for a defined context of use.

‍ ‍

Workshop participants described several potential sources of measurement error, including device performance, algorithm instability, adherence and data loss, differences between clinic and home environments, and population characteristics such as age, disease severity, and comorbidities.

‍ ‍

Clinical validation provides evidence that an analytically valid, digitally-derived measure correctly assesses the concept of interest in the target population and corresponds to an outcome meaningful to how a patient feels, functions, or survives.

‍ ‍

Missing data needs to be understood

‍ ‍

Missing data received substantial attention throughout the workshop. With DHTs, missingness can occur at multiple levels: individual epochs, assessments, days, monitoring periods, or entire participants. It may result from a participant not wearing the device because the device is burdensome or the participant misunderstands instructions, technical failure, loss of connectivity, or other behavioral, technical, or clinical factors.

‍ ‍

Critically, missingness may be informative, and understanding the reason or reasons for it can help determine whether that is the case. Simply treating missing observations as randomly missing could bias an endpoint. Speakers emphasized the need to understand when the amount of missing data affects the measurement of a given metric to the point that it is no longer reliable.

‍ ‍

Standardization facilitates communication and efficient review

‍ ‍

The FDA’s CGM technical specification, released earlier this year, was presented as a potential model for other digital technologies because it combines data standards, metadata, traceability, device information, missing-data documentation, and analytical specifications. Workshop participants also confirmed its usefulness.

‍ ‍

In the workshop's closing remarks, Mathilde Kam, Division Director for the Division of Analytics and Informatics in CDER's Office of Biostatistics, characterized standardization as fundamental. She noted that inconsistent terminology and data structures create barriers to cross-trial comparability and efficient regulatory review, while common standards can make analytical decisions more transparent and reproducible.

‍ ‍

Standardization is valuable in this context because it can expose assumptions, improve traceability, facilitate reproducibility, and make evidence easier to evaluate.

‍ ‍

A parallel with real-world data

‍ ‍

Although real-world data (RWD) are generally collected during routine care rather than prospectively within a clinical trial, many of the same fit-for-purpose challenges apply. In both settings, investigators need to understand how and why the data were generated, what data are missing and why, which transformations and assumptions shaped the data being analyzed, how well the data represent the intended population and clinical concept, and how much confidence the resulting evidence can support.

‍ ‍

10 Practical implications for sponsors

‍ ‍

Drawing from presentations throughout the workshop, I identified 10 practical lessons for sponsors submitting DHT-based data, several of which reflect insights shared by Cynthia Fisher of CDER’s Division of Analytics and Informatics based on patterns observed in DHT-based submissions.

‍ ‍

‍1.       DHT capabilities should not drive endpoint selection. The patient perspective must remain central, and digitally-derived endpoints should be selected only when they are meaningful to patients.

2.       Understanding measurement accuracy, quantifying bias and precision, and characterizing measurement error are all needed to support fit-for-purpose evidence.

3.       The entire process used to derive the digital endpoint should be validated, not only the final output.

4.       Validate the technology and algorithm in the intended population and setting.

5.       The goal of validation is not perfection. It is to understand sources of error well enough to determine whether the measure can support the question being asked.

6.       Using an FDA-cleared device does not resolve questions about the fitness for purpose of the device or its output for a particular endpoint. FDA clearance establishes that a device can measure something accurately; this is a separate question from fitness for purpose, which is context-dependent.

7.       Clearly define what counts as missing data and when missing data may compromise an endpoint measurement, and prespecify missing-data strategies.

8.       Evaluate the reasons for missingness. Sponsors should not assume it is random and should ensure that missingness is not systematic.

9.       Standardization provides a framework and consistent terminology for representing data. It facilitates communication, traceability, cross-trial comparability, and review efficiency.\

10.   Early engagement with FDA is highly encouraged. Validation, missing-data strategies, and analytical decision rules are easier to address before data have been collected and the analysis has been completed.

Challenges remain

‍ ‍

Statistical methodology must evolve alongside the technology. DHTs generate data that are fundamentally different from traditionally collected clinical trial data in structure, volume, and complexity. However, the standards applied to digitally-derived endpoints must be as high as those applied to any clinical endpoint intended to support a labeling claim. Promising approaches to handling missing data and validating novel endpoints were presented at the workshop, but more work is needed to advance understanding and methodology in this area.

‍ ‍

Greater standardization is needed to align terminology and analytical conventions. The Technical Specifications Document for submitting CGM data can serve as a model for developing other DHT specifications, but a level of maturity and consensus must be reached before these specifications have real value for sponsors and FDA.

‍ ‍

Conclusion

‍ ‍

As clinical research increasingly incorporates data generated outside traditional trial visits, the underlying scientific questions needed to establish rigorous, credible evidence for regulatory review may become more important than the labels applied to different data sources.

‍ ‍

Where did the data come from? What happened to the data along the way? What do the data actually measure? What is missing? What assumptions were made? How well do the data represent the intended population and clinical concept? And, ultimately, how much confidence can we place in the conclusion?

‍ ‍

These are not exclusively digital endpoint questions. They are fit-for-purpose evidence questions.

‍ ‍

Next
Next

From RWE Policy to Practice: Building Evidence Decision-Makers Can Trust