Biostatistics • July 30, 2026

Handling Missing Data in Clinical Trials: Advanced Imputation and Sensitivity Analysis

Glass clinical trial data matrix branching into imputation and sensitivity analysis pathways

Missing-data analysis should begin by defining the estimand and the missingness mechanism, not by selecting an imputation algorithm. Multiple imputation can address uncertainty under a missing-at-random assumption, while pattern-mixture, selection, and delta-adjustment analyses test how conclusions change when missing-not-at-random behavior is plausible.

Missing observations are common in clinical trials. Participants discontinue because of adverse events, lose access to follow-up, receive rescue medication, move to another treatment, or die before the scheduled assessment. The missingness is often related to the outcome itself or to an evolving clinical state. Treating these observations as a minor data-cleaning issue can therefore change the treatment effect that the trial reports.

A complete analysis connects four elements: the clinical question, the estimand, the missingness process, and the sensitivity analysis. The same dataset can support different treatment-effect questions depending on whether the analysis follows participants after discontinuation, estimates outcomes under continued treatment, or treats a post-randomization event as part of a composite endpoint.

1. MCAR, MAR, and MNAR are assumptions about the data process

Missing completely at random (MCAR) means that the probability of missingness is unrelated to both observed and unobserved data. This is rarely credible in clinical trials, although accidental specimen loss or a random technical failure may approximate it. Complete-case analysis is unbiased under MCAR when the analysis model is otherwise correct, but it can be inefficient because it discards observed information.

Missing at random (MAR) means that missingness may depend on information already observed, but not on the unobserved value after conditioning on the available data. A participant's probability of missing a visit may depend on prior symptoms, treatment assignment, and earlier biomarker values. MAR cannot be verified from the observed dataset; it is a modeling assumption supported by design knowledge and rich auxiliary variables.

Missing not at random (MNAR) describes situations in which missingness still depends on the unobserved value after accounting for observed information. A participant with worsening symptoms may be more likely to discontinue precisely because the unobserved outcome is unfavorable. MNAR is not a single model. It is a family of plausible departures from MAR that should be explored through sensitivity analysis.

2. Multiple imputation under MAR

Multiple imputation replaces each missing value with a set of plausible values generated from an imputation model. Each completed dataset is analyzed using the planned outcome model, and estimates are combined using rules that incorporate both within-imputation and between-imputation uncertainty. Unlike single imputation, multiple imputation does not pretend that the imputed value is known without error.

The imputation model should include treatment assignment, baseline prognostic factors, observed outcome history, reasons for discontinuation, auxiliary variables, and variables related to missingness. It should be at least as rich as the analysis model in relevant predictors. The number of imputations should be sufficient for the fraction of missing information, especially when the missingness rate is high or the analysis targets a tail probability.

3. Evidence summary table

Methodology / guidanceKey sourceLevel of evidence
Missing data prevention and analysisNational Research Council, The Prevention and Treatment of Missing Data in Clinical TrialsHigh: methodological guidance
Confirmatory trial missing dataEMA Guideline on Missing Data in Confirmatory Clinical TrialsHigh: regulatory guideline
Estimands and missingnessICH E9(R1) Addendum on Estimands and Sensitivity AnalysisHigh: international regulatory framework
Multiple imputationRubin, Multiple Imputation for Nonresponse in SurveysHigh: foundational statistical method

4. Why MAR-based results need sensitivity analysis

A primary analysis under MAR may be appropriate, but it cannot demonstrate that MNAR behavior would not alter the conclusion. Sensitivity analysis evaluates how much unobserved outcomes would need to differ from MAR-implied values to change the clinical interpretation. This is more informative than declaring an analysis robust because two similar models produced similar estimates.

Controlled multiple imputation applies a structured shift to imputed values after MAR imputation. A delta adjustment can make unobserved outcomes systematically worse or better than observed outcomes by a clinically interpretable amount. The analysis can identify the tipping point at which the treatment effect crosses a prespecified threshold, such as no benefit or a minimum clinically important difference.

5. Pattern-mixture and selection models

Pattern-mixture models stratify data according to missingness patterns and model the outcome distribution within each pattern. The unobserved part must be identified through assumptions, often expressed as a sensitivity parameter. Selection models factor the joint distribution into an outcome model and a missingness model, allowing missingness probability to depend on the unobserved outcome through a sensitivity parameter.

Neither framework is automatically superior. Pattern-mixture models often make clinical departures from MAR easier to explain, while selection models can align naturally with a proposed missingness mechanism. The important requirement is that the sensitivity parameter has a transparent interpretation and that the explored range is clinically defensible rather than selected because it preserves statistical significance.

6. Actionable steps for a defensible missing-data analysis

StepAnalysis phaseKey deliverable
Step 1Define the estimand and classify each intercurrent event affecting outcome observation.Estimand and event matrix
Step 2Describe missingness by treatment, time, reason, baseline risk, and observed outcome history.Missingness diagnostic
Step 3Pre-specify the primary MAR-compatible analysis and include auxiliary predictors.Primary analysis plan
Step 4Run controlled MI, delta-adjustment, or pattern-mixture sensitivity scenarios.MNAR sensitivity grid
Step 5Report the tipping point and clinical interpretation, not only p-values.Robustness conclusion

7. Prevention is part of missing-data methodology

Statistical handling cannot replace prevention. Protocols should minimize participant burden, collect reasons for missed visits, allow flexible visit windows where scientifically acceptable, and continue outcome collection after treatment discontinuation when the estimand requires it. Trial teams should document rescue treatment, treatment switching, and withdrawal reasons in a way that supports both the primary analysis and sensitivity analyses.

Operational planning also matters. A high rate of missing data may indicate that the intervention is poorly tolerated, that the visit schedule is unrealistic, or that the endpoint is difficult to measure. These are clinical and design signals, not just statistical nuisances. Monitoring should therefore review missingness patterns during the trial without using unblinded outcome information to make inappropriate adaptations.

8. Reporting standards for peer review

A strong manuscript reports the amount and timing of missing data, reasons for missingness, treatment-group differences, the estimand, primary missing-data assumption, imputation model, diagnostics, and sensitivity analyses. It should explain whether the conclusion is stable across clinically plausible MNAR scenarios. Reporting only the number of complete cases and a final model coefficient does not allow readers to evaluate risk of bias.

Elevate your clinical trial analysis with Lingcore SCI tools

Missing-data work requires alignment between estimands, imputation models, and sensitivity assumptions. Lingcore SCI provides specialized tools for medical researchers:

Conclusion

Missing data are part of clinical trial design and the estimand, not a final-stage spreadsheet problem. A credible analysis describes why values are missing, uses a transparent primary assumption, and tests clinically plausible departures from that assumption. Multiple imputation can provide efficient inference under MAR, while controlled imputation, pattern-mixture models, selection models, and delta-adjustment show whether the conclusion survives MNAR uncertainty. This combination gives clinicians, regulators, and peer reviewers a more honest assessment of what the trial can support.