Accurately Surveying Women
Using List Experiments to Measure Intimate Partner Violence (IPV): Evidence from rural Burkina Faso
Sarah Deschenes
I implement a measurement exercise and compare the estimates of the prevalence of intimate partner violence (IPV) obtained with direct questions and with a list experiment (LE), an indirect non-conventional measure for sensitive opinion or behavior. I measure the prevalence of several types of IPV (less severe physical violence, severe physical violence and marital rape) among women living in rural areas of Burkina Faso and find that the direct measures of IPV underestimate by 7 to 9 percentage points the most intense forms of violence. I also find that IPV correlates differently with a series of respondents' characteristics according to the measure of IPV used, including whether the first child of a woman is a girl. To the best of my knowledge, this paper is among the first to study the bias of IPV in a Western African country and whether this bias differs according to the intensity of IPV. This work fuels the debate around the relevance of the current measures of IPV in household surveys that seem to significantly underestimate the prevalence of IPV.
Who is Asking and How? Effects of Enumerator Gender and Survey Methods on Measuring Women’s Life Experience (WLE)
Aditi Kadam, Ellen McCullough, Tamara McGavock, Nicholas Magnan, & Thomas Assefa
Under-reporting and misreporting bias researchers’ perceptions about the respondent’s true beliefs about social norms and make them a concern for studying sensitive topics such as social norms and beliefs. Given limited causal evidence on effective ways of capturing sensitive norms, we use a sample of 637 women from the northern Amhara region in Ethiopia to test whether survey methods that may increase respondent privacy affect the reporting of social norms (such as women’s perceived independence, empowerment, and safety). We randomize the type of interview (phone vs in person), and gender of an interviewer to measure the direction of expected bias. Additionally, we explore the effect of rapport conditional on interview type by comparing responses between women who have had prior frequent interactions with our study team and those who have not. Preliminary results show that the reporting of sensitive topics such as physical safety, emotional well-being, and freedom of movement is not significantly different for respondents interviewed via mobile phone and respondents interviewed in person. Gender of enumerator and built-up rapport only affect the reporting of independence of movement. Respondents report lower independence to female enumerators and higher independence if they have a rapport with the survey team. Respondents may be reluctant to reveal honest answers because of social pressure, lack of privacy, shame, taboo, fear of safety, or have normalized patriarchal norms and violence in their minds. Given the importance and increasingly common usage of these questions in surveys for measuring women’s empowerment, traditional methods may compromise accurate measurement by introducing measurement bias. We develop a mechanism to document the direction of this expected bias in traditional surveys as well as in alternate survey methods.
No Clear Evidence Favoring the List Experiment to Measure Intimate Partner Violence in Nairobi City County
Winnie Mughogho, Dhwani Yagnaraman, Nicholas Owsley, & Patrick Forscher
Much of social science research relies on self-reported data. In recognising that survey respondents may not give truthful and accurate reports on sensitive questions when asked directly, indirect methods of questioning are now being adopted in research. In this study, we apply one such indirect technique, the List Experiment, to measure and characterize the experience of physical and emotional intimate partner violence (IPV) in a low income sample. We investigate whether there are any inconsistencies in between direct and indirect reports, and across what subgroups and mode of data collection. In addition, we tested the validity of the technique to ensure that we are measuring what it claims to measure. From our study we find that on the most part there are no statistical differences between direct and indirect reports of IPV. The direct reports tend to yield higher estimates with greater precision than indirect reports. We find higher direct reports of IPV from women, those without a college and those who are not married. Similar estimates are found for direct reports for the phone and the self administered survey. The subgroup analysis yields mixed findings for whether there is consistent reporting of IPV. We test compliance to the List Experiment with a screener list and find that those with a college education and those who took part in the phone survey show greater compliance. This could be attributed to the fact that the presence of an enumerator and some education improves comprehension of the method. The more compliant groups exhibited the most similar estimates to direct reports. From our findings there is no clear evidence that the List Experiment provides better estimates than direct reports.
Measurement Error & Statistical Power
How Accurate are Ex Ante Power Analyses? An Analysis from 3ie’s Replication Programme
Bob Reed, Alex Tian, Tom Coupe, Sayak Khatua, & Ben Wood
Almost all funding organizations require that an ex ante power analysis accompany any request for funding. Calculation of ex ante power is based on a number of factors, such as sample size, the hypothesized effect size, the treatment allocation, the unconditional variance of the outcome variable, the proportion of total variation in the dependent variable explained by the right-hand side covariates, and the intra-cluster correlation coefficient, among other things. Actual (ex post) power can deviate from ex ante power for many reasons. Yet to date, there is no generally acceptable method for calculating ex post power. This paper modifies a procedure suggested by McKenzie and Ozier (2019) for calculating ex post MDE. We call this method the SE-ES method. We carry out Monte Carlo experiments to determine its performance in a variety of data environments designed to match real life data. We then apply this method to a set of 49 estimated treatment effects from 23 studies that were funded by the International Initiative for Impact Evaluation (3ie). Each of these studies were required to provide ex ante power calculations as part of their funding application. Upon completion, all of the studies supplied their final data and code, enabling us to exactly reproduce their results. We applied the SE-ES method to estimate ex post power for each of the 49 estimated effects and compared them with their ex ante analogs. Overall, we find that the average of the ex post powers was approximately equal to the average of the ex ante powers, though there were some egregious outliers. We then estimated the extent to which differences between ex post and ex ante power could be attributed to differences in actual versus ex ante values for number of clusters and ICC.
Measuring Agricultural Survey Bias Across Couples with GIS and Lab-in-the-Field
Rachel Sayers, Ariel BenYishay, Jessica Wells, & Katherine Nolan
Household and farm surveys frequently interview an individual household member about the agricultural plot characteristics, inputs, and production across all plots farmed by the household. However, the gender of the respondent may yield differing reports on these measures in surveys, especially in settings where men and women farm different plots. Such disagreement in responses about a household’s plots may be due to information asymmetry from specialization, likely leading to random measurement error, or deliberate concealment, likely leading to systematic measurement error. The structure of this measurement error and how it relates to the ownership status of the plot (i.e. owned by the respondent, owned by the respondent’s spouse, or jointly owned) and household bargaining power is critical to produce unbiased estimates based on most survey data. We investigate the structure of this measurement error using survey data from couples that make decisions about agricultural plots in northern Ghana. We pair this survey data with outcome measures derived from high resolution satellite imagery matched to GPS plot walks conducted during survey collection, in order to validate self-reports on agricultural plot characteristics. Using this method, we are able to characterize the structure of measurement error inherent in the self-reports. Finally, we utilize lab-in-the-field games on bargaining power and willingness to pay to capture income, along with survey self-reports about decision-making to determine how the magnitude of measurement error varies based on household bargaining power and income hiding.
Using Validation Data to Correct for Treatment-Induced Measurement Error
Joshua Deutchman & Emilia Tjernstrom
In practice, measuring key variables in applied economics involves trade-offs between cost and quality (Carletto et al., 2021). Researchers must often choose between cheap but noisy and possibly biased survey measures and expensive but more accurate measures of outcomes. In agricultural settings in particular, cultivated acreage is a common outcome of interest, both as it relates to intensification and as a way to turn harvest data into yields and thereby estimate productivity. Recent work finds that measurement error in cultivated acreage is widespread in survey datasets and can bias estimates of the relationship between land and productivity (Carletto et al., 2013, 2015; Abay et al., 2019a; Dillon et al., 2019; Abay et al., 2021).
In this paper, we study the implications of a particular form of measurement error: differential mis-reporting by treated and control participants in a field experiment. We document this form of measurement error in self-reported cultivated acreage by farmers in a recent experiment in western Kenya (Deutschmann et al., 2021) and discuss some potential mechanisms for its presence. The main study uses GPS measures of cultivated acreage. We demonstrate that our interpretation of the results of the experiment would change if we relied on self-reported acreage.
We adapt a method from Carroll et al. (2006) and Buonaccorsi and Tosteson (1993) to correct for this error and demonstrate its effectiveness in our setting. In this method, researchers use a validation subsample in which one observes both biased and non-biased measures to correct for bias in a larger sample. We validate the procedure using data from Deutschmann et al. (2021), and show that the procedure performs well with relatively small validation samples. We discuss how researchers can use our results when planning data collection efforts, in the spirit of the data quality production function of Dillon et al. (2020).
Integrating Satellite Data with Surveys
Using Satellite Imagery to Inform Survey Design and Poverty Measurement
Manuel Cardona & Elliot Collins
Economic development and aid programs frequently call for targeting impoverished households or communities, raising the question of how such households can be easily identified. Carefully measuring poverty calls for detailed household consumption or income surveys that are usually too time-consuming and expensive, which has motivated several shorter, simpler proxy methods. In this study, we show how recent wealth mapping methods based on satellite imagery can be used to inform survey design and improve on widely adopted survey-based poverty measurement, yielding greater accuracy than either method alone without increasing costs. We begin with analysis of the Poverty Probability Index (PPI), a widely adopted set of country-specific survey modules used by hundreds of organizations across more than 25 countries. The PPI uses data from nationally representative household surveys to construct low-dimensional predictive models of household poverty. This model can then be translated into a simple scorecard using as few as ten survey questions. In parallel efforts, researchers have constructed high-resolution poverty maps based on satellite imagery and similar nationally representative household surveys. Compared to the PPI, these models provide much more spatial granularity, identifying variation in poverty at the 5 km level rather than at the level of large administrative regions. However, poverty maps naturally struggle to measure welfare variation within a particular area, which can be understood using survey-based proxy methods like the PPI. Using a variety of machine learning techniques that incorporate satellite imagery data into the PPI, we are able to construct a predictive model that outperforms either approach alone, improving upon the shortcomings of each method and providing insights to improve household survey design.
Impacts of Presurvey Messaging Content on Response Rates in Low-Income Countries
Steven Glazerman, Andrew Dillon, Dean Karlan, & Michael Rosenbaum
Survey methodologists have long experimented with different ways to improve response rates by pre-contacting sample members with messaging designed to predispose them to participation and cooperation. The methodological literature includes many studies of postcards, letters, endorsements, and prepaid incentives or gifts. For researchers working in low-income countries, however, the insights from this research have little relevance. Mailing gifts or letters is rarely feasible and the social and economic context for survey participation behavior is vastly different from that of high-income countries.
Building on a randomized experiment that we implemented to test the use of SMS messages to improve response rates to RDD surveys (Dillon et al. 2021), we designed an experiment to test more variations in pre-survey message content for followup surveys in five countries: Burkina Faso, Colombia, Cote d’Ivoire, Rwanda, and Zambia. We designed messaging to appeal alternatively to salience, self-interest, or a combination of the two. To test how salience works, we randomized whether the pre-survey SMS included information learned from the first round of the survey and what form that took: general or specific. We also randomly varied whether the information was about food insecurity or unemployment, two of the survey’s major areas of inquiry. To test how self-interest works, we cross-randomized a reminder about the monetary incentive. There were ten treatment arms in total.
We examine the impact of these survey messages on several outcomes. The first set is successful contact, cooperation, and completion: Which type of message content most improves response rates? The second set of outcomes is sample composition: How does message content influence the representativeness of the sample? The third set of outcomes is the survey responses themselves. Does message content translate into different inferences from the study itself?
Survey ParaData & Metadata as Research Infrastructure
Using Paradata to Assess Respondent Burden and Interviewer Effects in Household Surveys: Evidence from Low- and Middle-Income Countries
Ardina Hasanbasri, Talip Kilic, Gayatri Koolwal, & Heather Moylan
Household surveys constitute a vital component of national statistical systems; inform official statistics on an extensive range of social and economic phenomena; and are required for tracking progress towards national and international development goals. In low- and middle-income countries, household surveys have continued to grow in terms of topical coverage and complexity - including in terms of intra-household, individual-disaggregated data collection, with direct, yet understudied, implications for respondent burden, survey costs, and data quality. As national statistical offices (NSOs) in these contexts have accelerated the transition to computer-assisted personal interviewing (CAPI) and computer-assisted telephone interviewing (CATI), there is increasing scope for leveraging survey paradata to derive precise (i.e. questionnaire module- or question-specific) insights regarding interview durations, associated monetary costs, and potential interviewer effects embedded therein. Our study does precise that, by leveraging the time-stamped paradata generated by the Survey Solutions CAPI platform used by the large-scale national multi-topic household surveys that were implemented in Cambodia, Ethiopia and Tanzania with support from the Living Standard Measurement Study Plus (LSMS+) program. In addition to household interviews that have elicited core socioeconomic data, these surveys conducted private interviews with adult household members using cross-country comparable questionnaire modules to elicit self-reported information regarding work and employment and ownership of and rights to physical and financial assets. The use of paradata to study respondent burden and interviewer effects in household surveys is nascent in developing countries. Our analysis yield a range of operationally-relevant reference points regarding duration of modules included in multi-topic household and individual questionnaires and the underlying interviewer effects. These insights would be of interest to survey practitioners planning surveys based on comparable questionnaires and fieldwork protocols.
A Metadata Schema for Data from Experiments in the Social Sciences
Sarah Kopper, Jasmin Claire Fliegner, Jack Cavanagh, Sarah Kopper, & Anja Sautmann
The use of randomized controlled trials (RCTs) in the social sciences has greatly expanded, resulting in newly abundant data that is high-quality but often not easily accessible due to barriers to discovery and reuse. Increasing this accessibility is important: RCT data can be reused to perform methods research in program evaluation, to systematize evidence for policy makers, and for replication and training purposes that would also increase equity in data usage. In this paper, we propose a metadata schema that can serve as the basis for one (or many, interoperable) catalogs of RCT data to make high-quality social science RCT data easily findable, searchable and reusable for secondary research. The metadata schema takes into account the unique properties of such data and defines a set of fields and associated encoding schemes (acceptable formats and values) that can be used to describe any dataset associated with a social science RCT. We also create a set of recommendations for implementing a catalog or database based on this metadata schema.
Survey Data Quality: Survey Inattention & Reliable Measurement of Vaccines
Measurement of Survey Inattention in a Low Income Sample
Winnie Mughogho & Busara Center for Behavioral Economics
Inattention is common in self-administered surveys, adds noise and bias to survey data, and limits the data’s capacity to yield accurate statistical estimates. The literature on attention measurement centres on identifying inattentive respondents in surveys and recommends filtering these out of the sample. If however these measures of attention are related to sample characteristics, this filtering can lead to a biased sample. Given the differences in attention measures and their complexity, it is likely that measured attention is related to characteristics such as education. Rather than filter these out of the sample, it is better to try and increase respondent attention. In this study we therefore i) compare how different attention measures adapted from the literature fare in a sample consisting of students from the University of Nairobi and respondents from the urban slum of Kibera in Nairobi, Kenya. ii) assess whether attention is related to respondent characteristics, iii) explore whether a bonus financial incentive can reduce inattention in these surveys iv) test the validity of the attention measures to ensure that they measure what they claim to. From our results we find that measured attention is higher for students than the Kibera population. Measured attention is correlated with education level, age and employment status. A regression analysis reveals that the incentive does not increase attention overall for the sample, but increases attention in the low income Kibera sample. The attention measures are positively correlated with each other at 5% level of significance providing evidence of convergent validity. We also find evidence of construct validity in our attention measures using Confirmatory Factor Analysis.
Reliability of Survey and Administrative Data on Vaccinations
Philip Wollburg, Yannick Markhof, & Alberto Zezza
Have COVID-19 vaccination campaigns been misinformed? This study investigates the alignment of administrative vaccination data with survey data from national high-frequency phone surveys and face-to-face data collection. In the context of COVID-19, administrative statistics are the primary resource informing the progress of vaccination campaigns, but survey data is being used for information on vaccine hesitancy, barriers to access, and other ways to expedite vaccination efforts. Past research from before the pandemic and anecdotal evidence from COVID-19 have indicated that both data sources are subject to a number of potential biases that threaten their ability to provide accurate insights to vaccination campaigns. We study this issue in the context of Sub-Saharan Africa, a region that is trailing the rest of the world in reported vaccination rates. We find that vaccination rates estimated from survey data consistently exceed administrative figures across our study countries, but generally remain within the boundaries of vaccine procurement statistics. We investigate sampling and non-sampling related sources of this misalignment, using survey experiments, recalibration of sampling weights, and analyzing the administrative data sources. Based on our findings, we develop recommendations for survey design. As such, our contribution is relevant beyond the context of COVID-19 and matters for a large body of applied research on vaccine uptake and vaccination campaigns.
Survey Design Issues for Reliability: Dietary Diversity & Retail Productivity
Measuring Dietary Diversity with High Frequency Mobile Phone Interviews in Ethiopia
Thomas Assefa, Ellen McCullough, & Tamara McGavock
With a long reference period, researchers can measure diet diversity more accurately, especially for households that consume some items infrequently. Longer recall periods, however, are associated with a higher cognitive burden on respondents and higher recall errors. To address the fundamental trade-off between cognitive burden and recall period, we conducted an experiment to validate a novel method of constructing dietary diversity measures using high-frequency phone surveys in Ethiopia. We randomly assigned households to report their diets through either a traditional survey (a single interview recalling the woman’s 24-hour diet and the household’s 7-day diet) or through the novel method (via a series of phone calls received twice a day over a 7- day window, with each call corresponding to a short, bounded recall period). Our results show that the 7-day Household Dietary Diversity Score (HDDS) measured using high-frequency phone calls is significantly lower than the HDDS measured via in-person survey. This result likely occurs due to telescoping, in which respondents include additional items that might have been consumed outside of the 7-day recall period. Curiously, we find that Women’s 24-hour Dietary Diversity Score (WDDS) measured using high-frequency phone calls is significantly higher than the WDDS score measured using a single interview recalling the 24-hour period. This result likely occurs because respondents occasionally forget to report items over the 24-hour recall that they remember during a shorter, more clearly bounded recall period. Our results shed light on recall errors in reporting dietary diversity, suggesting the direction of bias can change depending on the length of the reference period. Our method offers a promising approach to extend respondents’ reference periods without exacerbating recall biases, which has important implications for those seeking to measure dietary diversity as a programmatic outcome.
Estimating Retail Productivity
Ajay Shenoy & Brenda Samaniego de la Parra
We propose to refine and apply a new method to estimate the productivity of small and medium-sized retailers in developing countries. A successful retail shop must attract customers, manage a storefront, and maintain ample inventory across many products. The project will 1) develop a model that will decompose retail productivity into these three dimensions; 2) estimate productivity for a large sample of single-establishment shops; 3) identify the managerial practices that best predict each dimension of productivity; 4) and estimate the medium-run persistence of productivity, and the practices that predict productivity growth. We will refine and estimate the model by creating a unique panel dataset of thousands of shops in Lusaka, Zambia, a city whose retail sector is broadly comparable to that of most cities in the developing world. The panel will combine a detailed in-person survey of the entrepreneur's characteristics and managerial practices with a high-frequency panel dataset of outcomes and input choices. This first round of data collection will inform our first three objectives. We will then collect a similar dataset on the same firms two years later, which will inform the final objective.
The project has not yet begun collecting the panel of firms in Zambia. We offer a limited demonstration using existing data from a survey of entrepreneurs in Malawi. The results confirm that a more limited model would have missed key aspects of the firm's productivity. We provide analytical arguments to show that existing methods would produce either biased or incomplete measures of productivity. This work will also be the foundation of a subsequent project that will use the measures of productivity to guide a randomized mentorship intervention.
Survey Mode & Representative Samples
Representativeness of Remote Methods in LMICs: A Cross-National Study of Pandemic-Era Surveys
Elliot Collins, Shana Warren, Savanna Henderson, & Michael Rosenbaum
Remote survey methods have been frequently used in LMIC’s in recent years, raising the important question of whether such methods yield nationally representative samples. In this article we assess this sampling bias question for a range of remote survey methods conducted during the first two years of the COVID-19 pandemic. We first consider both how estimates differ from those in nationally-representative face-to-face household surveys conducted prior to March 2020, and in the early months of the return to face-to-face surveys. We bring together random-digit-dial generated samples, phone follow-ups to face-to-face baseline surveys, and Facebook-sourced studies from a range of high-quality surveys conducted by leading research organizations. We then directly compare survey modes to one another in cases where multiple survey projects were conducted in the same country in roughly the same period, providing evidence on how sampling methods and survey mode affect the demographic profiles in different survey projects, while controlling for temporal effects of the pandemic.
We find that across a range of remote survey methods, most samples are biased to over-represent men, heads of households, younger, more educated, and more often employed respondents, though there are important differences across contexts and survey modes. The survey weighting strategies used in the studies we consider are not able to fully correct for these imbalances. Instead of reliance on weighting, we recommend that the thorough characterization of remotely recruited samples should be integrated into analysis and interpretation of remote methods.
This research note is part of a series assessing remote measurement in LMICs, bringing together studies conducted by a wide range of organizations to quantify key sources of bias in remote surveys and characterize the trade-offs between different survey modes, with the goal of informing future research and policy projects."
Does Survey Mode Matter? Evidence from Phone and In-Person Agricultural Surveys in India
Ellen Anderson, Travis Lybbert, Ashish Shenoy, Rupika Singh, & Daniel Stein
Phone surveys have become more common in developing countries and are increasingly used for data collection in randomized control trials. However, the consequences of switching from in-person to phone surveys in the context of RCTs are not fully understood. Of particular concern is whether measurement error in outcomes collected from phone surveys could bias treatment effects. We compare responses from phone and in-person surveys conducted for an overlapping set of questions and households that were part of an RCT studying the effects of a pulses farming promotion program on crop production in Bihar, India. We find differences in properties of the outcomes distributions between the phone and in-person surveys. The differences are driven by both selection and mode effects. However, we find similar treatment effects of the program by survey mode for both intent to treat and local average treatment effects.
Methods for Measuring Violence & Anthropometrics
Assessing the Feasibility of Caregiver-Administered Anthropometric Measurements
Doug Parkerson, Mpela Chembe, Günther Fink, Peter Rockers, & Dorothy Sikazwe
Accurate measurement of children’s height and weight is of central importance for the assessment of children’s nutritional status as well as for the evaluation of nutrition-specific interventions. During the COVID-19 pandemic, governments across the globe imposed social distancing requirements to restrict the spread of the virus. Social distancing made it impossible to follow the standard protocols for collecting basic anthropometric data for children, which require contact by trained enumerators with children to measure their height, weight, and mid-upper arm circumference (MUAC). From April through June 18th of 2021, the government of Zambia lifted its social distancing requirements, allowing us to apply the standard anthropometry protocol in a household survey in three districts. We compare the standard anthropometry protocol to a “no-contact” protocol that instructs caregivers to measure the height, weight, and MUAC of children in their household between the ages of 3 and 10 months.
Field teams took anthropometric measurements for 76 children using both the no-contact protocol and the standard protocol. On average for the 76 children measured by both methods, we find virtually no difference between the measurements in height (0.03 kg), weight (0.01 cm), and MUAC (-0.55 mm). We conclude that caregiver-administered anthropometric measurements are a promising alternative to the standard protocol in situations where outside assessors are unable to come in close contact with study participants.
Evidence-Based Decision-Making on Research Ethics in the Social Sciences
Graeme Blair, Rebecca Wolfe, & Rebecca Littman
Under the ethical principle of beneficence, researchers are required to show that the potential benefits of a proposed study to participants or society at large outweigh the risks of potential harm. Yet there is a challenge with this status quo: determining risks and benefits is not straightforward. Researchers and institutional review boards (IRBs) often rely on anecdotal evidence and assumptions about the probability and magnitude of risks and benefits. Unfortunately, there is mounting evidence that scientists’ accuracy at estimating how participants will respond to research is limited. We propose that to improve ethical standards in the social sciences, researchers should use the same tools they use to develop theory and test policy interventions: randomized experiments to estimate the causal effects of research on participants. To illustrate our proposal, we present data from three studies in conflict-affected areas of Nigeria where we randomized whether participants were asked about experiences with violence before or after questions on psychological distress. Contrary to the assumptions many researchers and IRBs hold, we do not find adverse effects of asking about violence on psychological wellbeing. We conclude by outlining a series of steps that researchers can take to accumulate evidence on beneficence and improve ethical decision-making.