Enhancing accuracy in surveys obtained by respondent-driven sampling with bias adjustment techniques
摘要
Respondent-driven sampling (RDS) is a network-based sampling technique widely used to survey hidden or hard-to-reach populations. Existing estimators rely on strong assumptions intended to ensure the unbiasedness of RDS estimators. Even though survey practitioners employ various practical tools to account for them, these assumptions are often unrealistic and challenging to satisfy in practice. In particular, when working with real-world survey data obtained through this methodology, RDS estimators may demonstrate poor performance, exhibiting notable bias. In this context, several strategies to adjust for selection biases inherent in non-probability samples, such as those that use propensity score adjustment, statistical matching and double robust estimators, are available and can outperform conventional RDS estimators. This study investigates the performance of such estimators with a simulated population designed to deviate from certain assumptions of respondent-driven sampling, and with an empirically collected network dataset. We examine how the number of seeds affects the performance of the RDS methodology by using different number of seeds to the first simulated population. We also examine network structures featuring individuals with exceptionally high centrality, along with structures where connections are either concentrated within small groups or more evenly distributed across groups. Our analysis reveals that these alternative estimators can outperform conventional RDS estimators, particularly in scenarios where the underlying assumptions of RDS are violated. This study represents the first application of these methods and techniques within an RDS framework.