Missing Data in Medical Tourism Research

Unreviewed Written 9 October 2026| 5 sources| The famous imputation warning traced to Dempster and Rubin, not to the handbook
Missing Data in Medical Tourism Research
Paper medical records being filed in a records office
Paper medical records being filed. Whether a field was never collected, collected and lost, or collected and withheld cannot be told apart from the finished dataset. Photograph by Tech. Sgt. Eric Burks, U.S. Air Force, Public domain, via Wikimedia Commons.
Verified against primary record
Three mechanismsMissing completely at random, missing at random, and not missing at random[1]
On filling gapsThe idea of imputation is both seductive and dangerous[1]
Single imputationKnown to underestimate the variance[1]
Records readA methodological handbook, a systematic review handbook, a statistical metadata record and two studies, 9 October 2026
Independently reported
An honest counterweightConcern about bias from incomplete outcome data is driven mainly by theoretical considerations[2]
Bands apply only to the rows beneath them. The imputation sentence is the handbook quoting Dempster and Rubin, 1983, not the handbook’s own wording.

Missing data is the ordinary condition of medical travel research rather than an exception to it. Records in this field are incomplete at source, and the methodological question is not how to avoid gaps but what may legitimately be done about them.

Three mechanisms, with different consequences

Why a value is absent matters more than how many are absent. The joint OECD and European Commission handbook on composite indicators sets out the standard three-way distinction. Under missing completely at random, missing values do not depend on the variable of interest or on any other observed variable in the data set. Under missing at random, missing values do not depend on the variable of interest, but are conditional on other variables in the data set. Under not missing at random, missing values depend on the values themselves.[1] The handbook abbreviates the third case as NMAR; the more common modern abbreviation is MNAR.

Medical travel data is overwhelmingly of the third kind, which is the worst. A patient whose result was poor is more likely to stop responding to a destination clinic. A complication treated at home is absent from the destination’s record precisely because it was a complication. The gaps are caused by the thing being measured, so no amount of data about the remaining patients recovers what is missing.

The warning about filling gaps, correctly attributed

The most quoted sentence in this area appears in that handbook: the idea of imputation is both seductive and dangerous. It is seductive because it can lull the user into believing that the data are complete after all, and dangerous because it lumps together situations where the problem is minor enough to be handled that way and situations where standard estimators applied to real and imputed data have substantial bias.[1]

The sentence is not the handbook’s own. The text immediately preceding it attributes it to Dempster and Rubin, writing in 1983, and the handbook’s section heading softens the modality to a conditional.[1] It is widely reproduced as an OECD position, and it should be cited as Dempster and Rubin as quoted in the handbook.

The handbook’s own prescriptions are narrower and more practical. The uncertainty in imputed data should be reflected by variance estimates. Single imputation is known to underestimate the variance. And no imputation model is free of assumptions, so results should be thoroughly checked for their statistical properties.[1] The objection is therefore not to imputation but to imputation that conceals how much was invented.

A counterweight worth carrying

The systematic review literature is more measured than the composite indicator literature. Guidance on assessing trials records that missing outcome data, due to attrition during the study or exclusions from the analysis, raise the possibility that the observed effect estimate is biased, and then adds that concerns over bias resulting from incomplete outcome data are driven mainly by theoretical considerations, with empirical studies mostly finding no clear evidence of bias.[2]

That is a useful corrective to the reflex that any missing data invalidates a result. It is also from an archived edition of that guidance, superseded by later versions, and is cited here as such.

Where the gaps actually are

Three documented examples show how this works in practice, and none of them involves a researcher choosing to impute.

Official statistics carry an explicit unknown category. European hospital discharge metadata records that some countries cannot report country of residence and report those cases as having an unknown country of residence, and that some countries report the hospital’s location instead of the patient’s residence.[3] The residence variable that would identify a foreign patient is the one that goes missing, and it goes missing non-randomly, by country.

Clinical records are incomplete at source. A commissioned review of complications presenting to the United Kingdom’s health service records that its data came from medical notes, which can be incomplete or wrongly coded, and that in consequence case numbers are under-reported and costs are under-estimated.[4] The direction of the error is known even though its size is not.

Published papers can be internally inconsistent. The best population survey study in this field reports its screening sample as 93,965 respondents with 535 positive answers in one section and 93,492 with 517 in its abstract and main table, and describes its jurisdictions as ten states and a territory in one place and eleven in another.[5] These are small discrepancies in a careful paper, noted here because the correct response is to disclose both figures rather than to pick one silently.

What to ask

How many values are missing, and from which variables. Why they are missing, in the three-way sense above. Whether anything was imputed, by what method, and whether the resulting uncertainty is reflected in the published intervals. And whether the paper reports a sensitivity analysis showing what the conclusion would have been under different assumptions about the missing cases. Where none of those is answered, a complete-looking table is not evidence that the data were complete.

See also

References

  1. Nardo M, Saisana M, Saltelli A, Tarantola S, Hoffmann A, Giovannini E. Handbook on constructing composite indicators, methodology and user guide, section 1.3, imputation of missing data. OECD and European Commission Joint Research Centre, 2008, ISBN 978-92-64-04345-9, JRC47008. Verified against primary record: section read in the handbook, including the attribution of the quoted sentence to Dempster AP and Rubin DB, in Incomplete Data in Sample Surveys volume 2, 1983. Retrieved 9 October 2026.
  2. Cochrane. Cochrane Handbook for Systematic Reviews of Interventions, archived version 5.1, section 8.13.1, rationale for concern about bias. Independently reported: archived methodological guidance opened and read; superseded by later editions. Retrieved 9 October 2026.
  3. Eurostat. Hospital discharges and length of stay, ESMS metadata, hlth_hosd. Metadata last updated 15 December 2025. Verified against primary record: residence reporting and country notes read in the metadata. Retrieved 9 October 2026.
  4. England C, Bromham N, Needham-Taylor A, Hounsome J, Gillen E, Ingram BJ, Davies J, Edwards A, Lewis R. Complications and costs to the UK National Health Service due to outward medical tourism for elective surgery, a rapid review. BMJ Open, 2026;16(1):e109050. Independently reported: publicly commissioned review, preprint full text read. Retrieved 9 October 2026.
  5. Stoney RJ, Kozarsky P, Walker AT, Gaines J. Population-based surveillance of medical tourism among U.S. residents from 11 states and territories. Infection Control and Hospital Epidemiology, 2022;43(7):870-875. Independently reported: author manuscript read on the agency’s document repository. Retrieved 9 October 2026.

Sourcing note: the handbook section, the archived review guidance, the statistical metadata and the two studies were opened and read on 9 October 2026. The sentence about imputation being seductive and dangerous is reproduced here with the attribution the handbook itself gives it, to Dempster and Rubin writing in 1983; it is commonly presented as an OECD statement and that is inaccurate. The review guidance is an archived edition and is labelled so. The discrepancies noted in the survey study are reported as published rather than resolved.