CREAT - Census Bureau

Optimal Stratified Sampling for Probability-Based Online Panels

September 2025

Written by: Jonathan Eggleston

Working Paper Number:

CES-25-69

Abstract

Online probability-based panels have emerged as a cost-efficient means of conducting surveys in the 21st century. While there have been various recent advancements in sampling techniques for online panels, several critical aspects of sampling theory for online panels are lacking. Much of current sampling theory from the middle of the 20th century, when response rates were high, and online panels did not exist. This paper presents a mathematical model of stratified sampling for online panels that takes into account historical response rates and survey costs. Through some simplifying assumptions, the model shows that the optimal sample allocation for online panels can largely resemble the solution for a cross-sectional survey. To apply the model, I use the Census Household Panel to show how this method could improve the average precision of key estimates. Holding fielding costs constant, the new sample rates improve the average precision of estimates between 1.47 and 17.25 percent, depending on the importance weight given to an overall population mean compared to mean estimates for racial and ethnic subgroups.

Document Tags and Keywords

Keywords:

data census, census data, survey, respondent, average, hispanic, trend, budget, population, rate, census bureau, sampling, sample, use census, assessing

Tags:

Computer Assisted Telephone Interviews and Computer Assisted Personal Interviews, American Community Survey, Health and Retirement Study, National Opinion Research Center, Census Bureau Disclosure Review Board

Similar Working Papers

The 10 most similar working papers to the working paper 'Optimal Stratified Sampling for Probability-Based Online Panels' are listed below in order of similarity.

Working Paper

CTC and ACTC Participation Results and IRS-Census Match Methodology, Tax Year 2020

December 2024

Authors: Charles Hokayem, Ethan Krohn, Ciyata Coleman, Sanghun (Eric) Kim, Krishnan Patel, Dean Plueger

Working Paper Number:

CES-24-76

The Child Tax Credit (CTC) and Additional Child Tax Credit (ACTC) offer assistance to help ease the financial burden of families with children. This paper provides taxpayer and dollar participation estimates for the CTC and ACTC covering tax year 2020. The estimates derive from an approach that relies on linking the 2021 Current Population Survey Annual Social and Economic Supplement (CPS ASEC) to IRS administrative data. This approach, called the Exact Match, uses survey data to identify CTC/ACTC eligible taxpayers and IRS administrative data to indicate which eligible taxpayers claimed and received the credit. Overall in tax year 2020, eligible taxpayers participated in the CTC and ACTC program at a rate of 93 percent while dollar participation was 91 percent.
View Full Paper PDF
Working Paper

An Economist's Primer on Survey Samples

September 2000

Authors: William J Carrington, John L Eltinge, Kristin McCue

Working Paper Number:

CES-00-15

Survey data underlie most empirical work in economics, yet economists typically have little familiarity with survey sample design and its effects on inference. This paper describes how sample designs depart from the simple random sampling model implicit in most econometrics textbooks, points out where the effects of this departure are likely to be greatest, and describes the relationship between design-based estimators developed by survey statisticians and related econometric methods for regression. Its intent is to provide empirical economists with enough background in survey methods to make informed use of design-based estimators. It emphasizes surveys of households (the source of most public-use files), but also considers how surveys of businesses differ. Examples from the National Longitudinal Survey of Youth of 1979 and the Current Population Survey illustrate practical aspects of design-based estimation.
View Full Paper PDF
Working Paper

Incorporating Administrative Data in Survey Weights for the 2018-2022 Survey of Income and Program Participation

October 2024

Authors: Jonathan Eggleston, Julia Yang

Working Paper Number:

CES-24-58

Response rates to the Survey of Income and Program Participation (SIPP) have declined over time, raising the potential for nonresponse bias in survey estimates. A potential solution is to leverage administrative data from government agencies and third-party data providers when constructing survey weights. In this paper, we modify various parts of the SIPP weighting algorithm to incorporate such data. We create these new weights for the 2018 through 2022 SIPP panels and examine how the new weights affect survey estimates. Our results show that before weighting adjustments, SIPP respondents in these panels have higher socioeconomic status than the general population. Existing weighting procedures reduce many of these differences. Comparing SIPP estimates between the production weights and the administrative data-based weights yields changes that are not uniform across the joint income and program participation distribution. Unlike other Census Bureau household surveys, there is no large increase in nonresponse bias in SIPP due to the COVID-19 Pandemic. In summary, the magnitude and sign of nonresponse bias in SIPP is complicated, and the existing weighting procedures may change the sign of nonresponse bias for households with certain incomes and program benefit statuses.
View Full Paper PDF
Working Paper

Nonresponse and Coverage Bias in the Household Pulse Survey: Evidence from Administrative Data

October 2024

Authors: Jonathan Eggleston, Carl Lieberman

Working Paper Number:

CES-24-60

The Household Pulse Survey (HPS) conducted by the U.S. Census Bureau is a unique survey that provided timely data on the effects of the COVID-19 Pandemic on American households and continues to provide data on other emergent social and economic issues. Because the survey has a response rate in the single digits and only has an online response mode, there are concerns about nonresponse and coverage bias. In this paper, we match administrative data from government agencies and third-party data to HPS respondents to examine how representative they are of the U.S. population. For comparison, we create a benchmark of American Community Survey (ACS) respondents and nonrespondents and include the ACS respondents as another point of reference. Overall, we find that the HPS is less representative of the U.S. population than the ACS. However, performance varies across administrative variables, and the existing weighting adjustments appear to greatly improve the representativeness of the HPS. Additionally, we look at household characteristics by their email domain to examine the effects on coverage from limiting email messages in 2023 to addresses from the contact frame with at least 90% deliverability rates, finding no clear change in the representativeness of the HPS afterwards.
View Full Paper PDF
Working Paper

The Impact of Household Surveys on 2020 Census Self-Response

July 2022

Authors: Jonathan Eggleston

Working Paper Number:

CES-22-24

Households who were sampled in 2019 for the American Community Survey (ACS) had lower self-response rates to the 2020 Census. The magnitude varied from -1.5 percentage point for household sampled in January 2019 to -15.1 percent point for households sampled in December 2019. Similar effects are found for the Current Population Survey (CPS) as well.
View Full Paper PDF
Working Paper

EITC Participation Results and IRS-Census Match Methodology, Tax Year 2021

December 2024

Authors: Charles Hokayem, Ethan Krohn, Ciyata Coleman, Sanghun (Eric) Kim, Krishnan Patel, Dean Plueger

Working Paper Number:

CES-24-75

The Earned Income Tax Credit (EITC), enacted in 1975, offers a refundable tax credit to low income working families. This paper provides taxpayer and dollar participation estimates for the EITC covering tax year 2021. The estimates derive from an approach that relies on linking the 2022 Current Population Survey Annual Social and Economic Supplement (CPS ASEC) to IRS administrative data. This approach, called the Exact Match, uses survey data to identify EITC eligible taxpayers and IRS administrative data to indicate which eligible taxpayers claimed and received the credit. Overall in tax year 2021 eligible taxpayers participated in the EITC program at a rate of 78 percent while dollar participation was 81 percent.
View Full Paper PDF
Working Paper

BIAS IN FOOD STAMPS PARTICIPATION ESTIMATES IN THE PRESENCE OF MISREPORTING ERROR

March 2013

Authors: Cathleen Li

Working Paper Number:

CES-13-13

This paper focuses on how survey misreporting of food stamp receipt can bias demographic estimation of program participation. Food stamps is a federally funded program which subsidizes the nutrition of low-income households. In order to improve the reach of this program, studies on how program participation varies by demographic groups have been conducted using census data. Census data are subject to a lot of misreporting error, both underreporting and over-reporting, which can bias the estimates. The impact of misreporting error on estimate bias is examined by calculating food stamp participation rates, misreporting rates, and bias for select household characteristics (covariates).
View Full Paper PDF
Working Paper

Gradient Boosting to Address Statistical Problems Arising from Non-Linkage of Census Bureau Datasets

June 2024

Authors: Narayan Sastry, Todd Gardner, Matthew Cefalu, John Sullivan, Elizabeth Fussell

Working Paper Number:

CES-24-27

This article introduces the twangRDC package, which contains functions to address non-linkage in US Census Bureau datasets. The Census Bureau's Person Identification Validation System facilitates data linkage by assigning unique person identifiers to federal, third party, decennial census, and survey data. Not all records in these datasets can be linked to the reference file and as such not all records will be assigned an identifier. This article is a tutorial for using the twangRDC to generate nonresponse weights to account for non-linkage of person records across US Census Bureau datasets.
View Full Paper PDF
Working Paper

Interactions, Neighborhood Selection, and Housing Demand

August 2002

Authors: Yannis M Ioannides, Jeffrey Zabel

Working Paper Number:

CES-02-19

This paper contributes to the growing literature that identifies and measures the impact of social context on individual economic behavior. We develop a model of housing demand with neighborhood e'ects and neighborhood choice. Modelling neighborhood choice is of fundamental importance in estimating and understanding endogenous and exogenous neighborhood effects. That is, to obtain unbiased estimates of neighborhood effects, it is necessary to control for non-random sorting into neighborhoods. Estimation of the model exploits a unique data set of household data that has been augmented with contextual information at two di'erent levels ('scales') of aggregation. One is at the neighborhood cluster level, of about ten neighbors, with the data coming from a special sample of the American Housing Survey. A second level is the census tract to which these dwelling units belong. Tract-level data are available in the Summary Tape Files of the decennial Census data. We merge these two data sets by gaining access to confidential data of the U.S. Bureau of the Census. We overcome some limitations of these data by implementing some significant methodological advances in estimating discrete choice models. Our results for the neighborhood choice model indicate that individuals prefer to live near others like themselves. This can perpetuate income inequality since those with the best opportunities at economic success will cluster together. The results for the housing demand equation are similar to those in our earlier work [Ioannides and Zabel (2000] where we find evidence of significant endogenous and contextual neighborhood effects.
View Full Paper PDF
Working Paper

A METHOD OF CORRECTING FOR MISREPORTING APPLIED TO THE FOOD STAMP PROGRAM

May 2013

Authors: Nikolas Mittag

Working Paper Number:

CES-13-28

Survey misreporting is known to be pervasive and bias common statistical analyses. In this paper, I first use administrative data on SNAP receipt and amounts linked to American Community Survey data from New York State to show that survey data can misrepresent the program in important ways. For example, more than 1.4 billion dollars received are not reported in New York State alone. 46 percent of dollars received by house- holds with annual income above the poverty line are not reported in the survey data, while only 19 percent are missing below the poverty line. Standard corrections for measurement error cannot remove these biases. I then develop a method to obtain consistent estimates by combining parameter estimates from the linked data with publicly available data. This conditional density method recovers the correct estimates using public use data only, which solves the problem that access to linked administrative data is usually restricted. I examine the degree to which this approach can be used to extrapolate across time and geography, in order to solve the problem that validation data is often based on a convenience sample. I present evidence from within New York State that the extent of heterogeneity is small enough to make extrapolation work well across both time and geography. Extrapolation to the entire U.S. yields substantive differences to survey data and reduces deviations from official aggregates by a factor of 4 to 9 compared to survey aggregates.
View Full Paper PDF

Optimal Stratified Sampling for Probability-Based Online Panels

September 2025

Working Paper Number:

CES-25-69

Abstract

Document Tags and Keywords

The 10 most similar working papers to the working paper 'Optimal Stratified Sampling for Probability-Based Online Panels' are listed below in order of similarity.

December 2024

Working Paper Number:

CES-24-76

September 2000

Working Paper Number:

CES-00-15

October 2024

Working Paper Number:

CES-24-58

October 2024

Working Paper Number:

CES-24-60

July 2022

Working Paper Number:

CES-22-24

December 2024

Working Paper Number:

CES-24-75

March 2013

Working Paper Number:

CES-13-13

June 2024

Working Paper Number:

CES-24-27

August 2002

Working Paper Number:

CES-02-19

May 2013

Working Paper Number:

CES-13-28