-
Effects of a Government-Academic Partnership: Has the NSF-Census Bureau Research Network Helped Improve the U.S. Statistical System?
January 2017
Authors:
Lars Vilhuber,
John M. Abowd,
Daniel Weinberg,
Jerome P. Reiter,
Matthew D. Shapiro,
Robert F. Belli,
Noel Cressie,
David C. Folch,
Scott H. Holan,
Margaret C. Levenstein,
Kristen M. Olson,
Jolene Smyth,
Leen-Kiat Soh,
Bruce D. Spencer,
Seth E. Spielman,
Christopher K. Wikle
Working Paper Number:
CES-17-59R
The National Science Foundation-Census Bureau Research Network (NCRN) was established in 2011 to create interdisciplinary research nodes on methodological questions of interest and significance to the broader research community and to the Federal Statistical System (FSS), particularly the Census Bureau. The activities to date have covered both fundamental and applied statistical research and have focused at least in part on the training of current and future generations of researchers in skills of relevance to surveys and alternative measurement of economic units, households, and persons. This paper discusses some of the key research findings of the eight nodes, organized into six topics: (1) Improving census and survey data collection methods; (2) Using alternative sources of data; (3) Protecting privacy and confidentiality by improving disclosure avoidance; (4) Using spatial and spatio-temporal statistical modeling to improve estimates; (5) Assessing data cost and quality tradeoffs; and (6) Combining information from multiple sources. It also reports on collaborations across nodes and with federal agencies, new software developed, and educational activities and outcomes. The paper concludes with an evaluation of the ability of the FSS to apply the NCRN's research outcomes and suggests some next steps, as well as the implications of this research-network model for future federal government renewal initiatives.
View Full
Paper PDF
-
The Potential for Using Combined Survey and Administrative Data Sources to Study Internal Labor Migration
January 2017
Working Paper Number:
CES-17-55
This paper introduces a novel data set combining survey data from the American Community Survey (ACS) with administrative data on employment from the Longitudinal Employer-Household Dynamics program, in order to study geographic labor mobility. With its rich set of information about individuals at the time of the migration decision, large sample size, and near-comprehensive ability to detect labor mobility, the new combined ACS-LEHD data offers several advantages over the existing data sets that are typically used in the study of migration, such as the Decennial Census, Current Population Survey, and Internal Revenue Service data. An overview of how these different data sets can be employed, and examples demonstrating the usefulness of the newly proposed data set, are provided.
Aggregate statistics and stylized facts are generated from the ACS-LEHD data which reveal many of the same features as the existing data sets, including the decline of aggregate mobility throughout the past decade, as well as many of the known demographic differences in migration propensity.
View Full
Paper PDF
-
The Annual Survey of Entrepreneurs: An Update
January 2017
Working Paper Number:
CES-17-46
We provide an update on the Annual Survey of Entrepreneurs (ASE), which is a relatively new Census Bureau business survey. About 290,000 employer firms in the private, non-agricultural U.S. economy are in the ASE sample. Its content is relatively constant over collections, allowing for comparability over time; however, each year there are approximately ten new questions in a changing topical module. Earlier topical modules covered innovation (2014) and management practices (2015). The topical module for reference year 2016 covers business advice and planning, finance, and regulations. The ASE is collected through a partnership of the Census Bureau with the Kauffman Foundation and the Minority Business Development Agency. Qualified researchers on approved projects may request access to the ASE micro data through the Federal Statistical Research Data Center (FSRDC) network.
View Full
Paper PDF
-
Decennial Census Return Rates: The Role of Social Capital
January 2017
Working Paper Number:
CES-17-39
This paper explores how useful information about social and civic engagement (social capital)
might be to the U.S. Census Bureau in their efforts to improve predictions of mail return rates for the Decennial Census (DC) at the census tract level. Through construction of Hard-to-count (HRC) scores and multivariate analysis, we find that if information about social capital were available, predictions of response rates would be marginally improved.
View Full
Paper PDF
-
Revisiting the Economics of Privacy: Population Statistics and Confidentiality Protection as Public Goods
January 2017
Working Paper Number:
CES-17-37
We consider the problem of determining the optimal accuracy of public statistics when increased accuracy requires a loss of privacy. To formalize this allocation problem, we use tools from statistics and computer science to model the publication technology used by a public statistical agency. We derive the demand for accurate statistics from first principles to generate interdependent preferences that account for the public-good nature of both data accuracy and privacy loss. We first show data accuracy is inefficiently undersupplied by a private provider. Solving the appropriate social planner's problem produces an implementable publication strategy. We implement the socially optimal publication plan for statistics on income and health status using data from the American Community Survey, National Health Interview Survey, Federal Statistical System Public Opinion Survey and Cornell National Social Survey. Our analysis indicates that welfare losses from providing too much privacy protection and, therefore, too little accuracy can be substantial.
View Full
Paper PDF
-
Public-Use vs. Restricted-Use:
An Analysis Using the American Community Survey
January 2017
Working Paper Number:
CES-17-12
Statistical agencies frequently publish microdata that have been altered to protect confidentiality. Such data retain utility for many types of broad analyses but can yield biased or Insufficiently precise results in others. Research access to de-identified versions of the restricted-use data with little or no alteration is often possible, albeit costly and time-consuming. We investigate the the advantages and disadvantages of public-use and restricted-use data from the American Community
Survey (ACS) in constructing a wage index. The public-use data used were Public Use Microdata Samples, while the restricted-use data were accessed via a Federal Statistical Research Data Center. We discuss the advantages and disadvantages of each data source and compare estimated CWIs and standard errors at the state and labor market levels.
View Full
Paper PDF
-
Medicare Coverage and Reporting
December 2016
Working Paper Number:
carra-2016-12
Medicare coverage of the older population in the United States is widely recognized as being nearly universal. Recent statistics from the Current Population Survey Annual Social and Economic Supplement (CPS ASEC) indicate that 93 percent of individuals aged 65 and older were covered by Medicare in 2013. Those without Medicare include those who are not eligible for the public health program, though the CPS ASEC estimate may also be impacted by misreporting. Using linked data from the CPS ASEC and Medicare Enrollment Database (i.e., the Medicare administrative data), we estimate the extent to which individuals misreport their Medicare coverage. We focus on those who report having Medicare but are not enrolled (false positives) and those who do not report having Medicare but are enrolled (false negatives). We use regression analyses to evaluate factors associated with both types of misreporting including socioeconomic, demographic, and household characteristics. We then provide estimates of the implied Medicare-covered, insured, and uninsured older population, taking into account misreporting in the CPS ASEC. We find an undercount in the CPS ASEC estimates of the Medicare covered population of 4.5 percent. This misreporting is not random - characteristics associated with misreporting include citizenship status, year of entry, labor force participation, Medicare coverage of others in the household, disability status, and imputation of Medicare responses. When we adjust the CPS ASEC estimates to account for misreporting, Medicare coverage of the population aged 65 and older increases from 93.4 percent to 95.6 percent while the uninsured rate decreases from 1.4 percent to 1.3 percent.
View Full
Paper PDF
-
Response Error & the Medicaid undercount in the CPS
December 2016
Working Paper Number:
carra-2016-11
The Current Population Survey Annual Social and Economic Supplement (CPS ASEC) is an important source for estimates of the uninsured population. Previous research has shown that survey estimates produce an undercount of beneficiaries compared to Medicaid enrollment records. We extend past work by examining the Medicaid undercount in the 2007-2011 CPS ASEC compared to enrollment data from the Medicaid Statistical Information System for calendar years 2006-2010. By linking individuals across datasets, we analyze two types of response error regarding Medicaid enrollment - false negative error and false positive error. We use regression analysis to identify factors associated with these two types of response error in the 2011 CPS ASEC. We find that the Medicaid undercount was between 22 and 31 percent from 2007 to 2011. In 2011, the false negative rate was 40 percent, and 27 percent of Medicaid reports in CPS ASEC were false positives. False negative error is associated with the duration of enrollment in Medicaid, enrollment in Medicare and private insurance, and Medicaid enrollment in the survey year. False positive error is associated with enrollment in Medicare and shared Medicaid coverage in the household. We discuss implications for survey reports of health insurance coverage and for estimating the uninsured population.
View Full
Paper PDF
-
Evaluating the Use of Commercial Data to Improve Survey Estimates of Property Taxes
August 2016
Working Paper Number:
carra-2016-06
While commercial data sources offer promise to statistical agencies for use in production of official statistics, challenges can arise as the data are not collected for statistical purposes. This paper evaluates the use of 2008-2010 property tax data from CoreLogic, Inc. (CoreLogic), aggregated from county and township governments from around the country, to improve 2010 American Community Survey (ACS) estimates of property tax amounts for single-family homes. Particularly, the research evaluates the potential to use CoreLogic to reduce respondent burden, to study survey response error and to improve adjustments for survey nonresponse. The research found that the coverage of the CoreLogic data varies between counties as does the correspondence between ACS and CoreLogic property taxes. This geographic variation implies that different approaches toward using CoreLogic are needed in different areas of the country. Further, large differences between CoreLogic and ACS property taxes in certain counties seem to be due to conceptual differences between what is collected in the two data sources. The research examines three counties, Clark County, NV, Philadelphia County, PA and St. Louis County, MO, and compares how estimates would change with different approaches using the CoreLogic data. Mean county property tax estimates are highly sensitive to whether ACS or CoreLogic data are used to construct estimates. Using CoreLogic data in imputation modeling for nonresponse adjustment of ACS estimates modestly improves the predictive power of imputation models, although estimates of county property taxes and property taxes by mortgage status are not very sensitive to the imputation method.
View Full
Paper PDF
-
Playing with Matches: An Assessment of Accuracy in Linked Historical Data
June 2016
Working Paper Number:
carra-2016-05
This paper evaluates linkage quality achieved by various record linkage techniques used in historical demography. I create benchmark, or truth, data by linking the 2005 Current Population Survey Annual Social and Economic Supplement to the Social Security Administration's Numeric Identification System by Social Security Number. By comparing simulated linkages to the benchmark data, I examine the value added (in terms of number and quality of links) from incorporating text-string comparators, adjusting age, and using a probabilistic matching algorithm. I find that text-string comparators and probabilistic approaches are useful for increasing the linkage rate, but use of text-string comparators may decrease accuracy in some cases. Overall, probabilistic matching offers the best balance between linkage rates and accuracy.
View Full
Paper PDF