Sampling and Sampling Frames in Big Data Epidemiology

AbstractPurpose of ReviewThe ‘big data’ revolution affords the opportunity to reuse administrative datasets for public health research. While such datasets offer dramatically increased statistical power compared with conventional primary data collection, typically at much lower cost, their use also raises substantial infere ntial challenges. In particular, it can be difficult to make population inferences because the sampling frames for many administrative datasets are undefined. We reviewed options for accounting for sampling in big data epidemiology.Recent FindingsWe identified three common strategies for accounting for sampling when the data available were not collected from a deliberately constructed sample: (1) explicitly reconstruct the sampling frame, (2) test the potential impacts of sampling using sensitivity analyses, and (3) limit inference to sample.SummaryInference from big data can be challenging because the impacts of sampling are unclear. Attention to sampling frames can minimize risks of bias.
Source: Current Epidemiology Reports - Category: Epidemiology Source Type: research