CREAT: Census Research Exploration and Analysis Tool

Papers Containing Keywords(s): 'use census'

The following papers contain search terms that you selected. From the papers listed below, you can navigate to the PDF, the profile page for that working paper, or see all the working papers written by an author. You can also explore tags, keywords, and authors that occur frequently within these papers.
Click here to search again

Frequently Occurring Concepts within this Search

American Community Survey - 23

Internal Revenue Service - 19

Census Bureau Disclosure Review Board - 16

Protected Identification Key - 15

Social Security Administration - 15

2010 Census - 15

Center for Economic Studies - 14

Decennial Census - 13

Social Security Number - 13

Current Population Survey - 12

Master Address File - 12

Person Validation System - 11

Personally Identifiable Information - 11

Person Identification Validation System - 9

Longitudinal Employer Household Dynamics - 9

Service Annual Survey - 9

Bureau of Labor Statistics - 8

Social Security - 8

Cornell University - 8

Federal Statistical Research Data Center - 8

North American Industry Classification System - 8

Survey of Income and Program Participation - 7

Some Other Race - 7

Disclosure Review Board - 7

1940 Census - 7

Supplemental Nutrition Assistance Program - 6

Census Numident - 6

Census Edited File - 6

Individual Taxpayer Identification Numbers - 6

Research Data Center - 6

Business Register - 6

National Science Foundation - 6

Temporary Assistance for Needy Families - 5

Office of Management and Budget - 5

Indian Health Service - 5

Department of Housing and Urban Development - 5

Census Bureau Business Register - 5

Annual Survey of Manufactures - 5

Standard Industrial Classification - 5

Longitudinal Business Database - 5

Economic Census - 5

Medicaid Services - 4

Census Household Composition Key - 4

Population Estimates Program - 4

Housing and Urban Development - 4

National Opinion Research Center - 4

MAFID - 4

Census Bureau Master Address File - 4

Computer Assisted Personal Interview - 4

Selective Service System - 4

Metropolitan Statistical Area - 4

Standard Statistical Establishment List - 4

Employer Identification Numbers - 4

Public Use Micro Sample - 4

Administrative Records - 4

Postal Service - 4

American Housing Survey - 4

Quarterly Workforce Indicators - 4

DOB - 4

Center for Administrative Records Research and Applications - 3

National Academy of Sciences - 3

Stanford University - 3

Adjusted Gross Income - 3

Indian Housing Information Center - 3

Quarterly Census of Employment and Wages - 3

Census Bureau Person Identification Validation System - 3

American Economic Association - 3

Statistics Canada - 3

NUMIDENT - 3

Department of Homeland Security - 3

Cornell Institute for Social and Economic Research - 3

Urban Institute - 3

Business Dynamics Statistics - 3

PIKed - 3

Federal Reserve Bank - 3

Federal Tax Information - 3

Census of Manufactures - 3

Census Bureau Longitudinal Business Database - 3

General Accounting Office - 3

National Center for Health Statistics - 3

Special Sworn Status - 3

Local Employment Dynamics - 3

Agency for Healthcare Research and Quality - 3

Viewing papers 11 through 20 of 37


  • Working Paper

    A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census: Full Technical Report

    December 2023

    Working Paper Number:

    CES-23-63R

    For the last half-century, it has been a common and accepted practice for statistical agencies, including the United States Census Bureau, to adopt different strategies to protect the confidentiality of aggregate tabular data products from those used to protect the individual records contained in publicly released microdata products. This strategy was premised on the assumption that the aggregation used to generate tabular data products made the resulting statistics inherently less disclosive than the microdata from which they were tabulated. Consistent with this common assumption, the 2010 Census of Population and Housing in the U.S. used different disclosure limitation rules for its tabular and microdata publications. This paper demonstrates that, in the context of disclosure limitation for the 2010 Census, the assumption that tabular data are inherently less disclosive than their underlying microdata is fundamentally flawed. The 2010 Census published more than 150 billion aggregate statistics in 180 table sets. Most of these tables were published at the most detailed geographic level'individual census blocks, which can have populations as small as one person. Using only 34 of the published table sets, we reconstructed microdata records including five variables (census block, sex, age, race, and ethnicity) from the confidential 2010 Census person records. Using only published data, an attacker using our methods can verify that all records in 70% of all census blocks (97 million people) are perfectly reconstructed. We further confirm, through reidentification studies, that an attacker can, within census blocks with perfect reconstruction accuracy, correctly infer the actual census response on race and ethnicity for 3.4 million vulnerable population uniques (persons with race and ethnicity different from the modal person on the census block) with 95% accuracy. Having shown the vulnerabilities inherent to the disclosure limitation methods used for the 2010 Census, we proceed to demonstrate that the more robust disclosure limitation framework used for the 2020 Census publications defends against attacks that are based on reconstruction. Finally, we show that available alternatives to the 2020 Census Disclosure Avoidance System would either fail to protect confidentiality, or would overly degrade the statistics' utility for the primary statutory use case: redrawing the boundaries of all of the nation's legislative and voting districts in compliance with the 1965 Voting Rights Act. You are reading the full technical report. For the summary paper see https://doi.org/10.1162/99608f92.4a1ebf70.
    View Full Paper PDF
  • Working Paper

    The 2010 Census Confidentiality Protections Failed, Here's How and Why

    December 2023

    Working Paper Number:

    CES-23-63

    Using only 34 published tables, we reconstruct five variables (census block, sex, age, race, and ethnicity) in the confidential 2010 Census person records. Using the 38-bin age variable tabulated at the census block level, at most 20.1% of reconstructed records can differ from their confidential source on even a single value for these five variables. Using only published data, an attacker can verify that all records in 70% of all census blocks (97 million people) are perfectly reconstructed. The tabular publications in Summary File 1 thus have prohibited disclosure risk similar to the unreleased confidential microdata. Reidentification studies confirm that an attacker can, within blocks with perfect reconstruction accuracy, correctly infer the actual census response on race and ethnicity for 3.4 million vulnerable population uniques (persons with nonmodal characteristics) with 95% accuracy, the same precision as the confidential data achieve and far greater than statistical baselines. The flaw in the 2010 Census framework was the assumption that aggregation prevented accurate microdata reconstruction, justifying weaker disclosure limitation methods than were applied to 2010 Census public microdata. The framework used for 2020 Census publications defends against attacks that are based on reconstruction, as we also demonstrate here. Finally, we show that alternatives to the 2020 Census Disclosure Avoidance System with similar accuracy (enhanced swapping) also fail to protect confidentiality, and those that partially defend against reconstruction attacks (incomplete suppression implementations) destroy the primary statutory use case: data for redistricting all legislatures in the country in compliance with the 1965 Voting Rights Act.
    View Full Paper PDF
  • Working Paper

    Noncitizen Coverage and Its Effects on U.S. Population Statistics

    August 2023

    Working Paper Number:

    CES-23-42

    We produce population estimates with the same reference date, April 1, 2020, as the 2020 Census of Population and Housing by combining 31 types of administrative record (AR) and third-party sources, including several new to the Census Bureau with a focus on noncitizens. Our AR census national population estimate is higher than other Census Bureau official estimates: 1.8% greater than the 2020 Demographic Analysis high estimate, 3.0% more than the 2020 Census count, and 3.6% higher than the vintage-2020 Population Estimates Program estimate. Our analysis suggests that inclusion of more noncitizens, especially those with unknown legal status, explains the higher AR census estimate. About 19.8% of AR census noncitizens have addresses that cannot be linked to an address in the 2020 Census collection universe, compared to 5.7% of citizens, raising the possibility that the 2020 Census did not collect data for a significant fraction of noncitizens residing in the United States under the residency criteria used for the census. We show differences in estimates by age, sex, Hispanic origin, geography, and socioeconomic characteristics symptomatic of the differences in noncitizen coverage.
    View Full Paper PDF
  • Working Paper

    Estimating the U.S. Citizen Voting-Age Population (CVAP) Using Blended Survey Data, Administrative Record Data, and Modeling: Technical Report

    April 2023

    Working Paper Number:

    CES-23-21

    This report develops a method using administrative records (AR) to fill in responses for nonresponding American Community Survey (ACS) housing units rather than adjusting survey weights to account for selection of a subset of nonresponding housing units for follow-up interviews and for nonresponse bias. The method also inserts AR and modeling in place of edits and imputations for ACS survey citizenship item nonresponses. We produce Citizen Voting-Age Population (CVAP) tabulations using this enhanced CVAP method and compare them to published estimates. The enhanced CVAP method produces a 0.74 percentage point lower citizen share, and it is 3.05 percentage points lower for voting-age Hispanics. The latter result can be partly explained by omissions of voting-age Hispanic noncitizens with unknown legal status from ACS household responses. Weight adjustments may be less effective at addressing nonresponse bias under those conditions.
    View Full Paper PDF
  • Working Paper

    Improving Estimates of Neighborhood Change with Constant Tract Boundaries

    May 2022

    Working Paper Number:

    CES-22-16

    Social scientists routinely rely on methods of interpolation to adjust available data to their research needs. This study calls attention to the potential for substantial error in efforts to harmonize data to constant boundaries using standard approaches to areal and population interpolation. We compare estimates from a standard source (the Longitudinal Tract Data Base) to true values calculated by re-aggregating original 2000 census microdata to 2010 tract areas. We then demonstrate an alternative approach that allows the re-aggregated values to be publicly disclosed, using 'differential privacy' (DP) methods to inject random noise to protect confidentiality of the raw data. The DP estimates are considerably more accurate than the interpolated estimates. We also examine conditions under which interpolation is more susceptible to error. This study reveals cause for greater caution in the use of interpolated estimates from any source. Until and unless DP estimates can be publicly disclosed for a wide range of variables and years, research on neighborhood change should routinely examine data for signs of estimation error that may be substantial in a large share of tracts that experienced complex boundary changes.
    View Full Paper PDF
  • Working Paper

    Redesigning the Longitudinal Business Database

    May 2021

    Working Paper Number:

    CES-21-08

    In this paper we describe the U.S. Census Bureau's redesign and production implementation of the Longitudinal Business Database (LBD) first introduced by Jarmin and Miranda (2002). The LBD is used to create the Business Dynamics Statistics (BDS), tabulations describing the entry, exit, expansion, and contraction of businesses. The new LBD and BDS also incorporate information formerly provided by the Statistics of U.S. Businesses program, which produced similar year-to-year measures of employment and establishment flows. We describe in detail how the LBD is created from curation of the input administrative data, longitudinal matching, retiming of economic census-year births and deaths, creation of vintage consistent industry codes and noise factors, and the creation and cleaning of each year of LBD data. This documentation is intended to facilitate the proper use and understanding of the data by both researchers with approved projects accessing the LBD microdata and those using the BDS tabulations.
    View Full Paper PDF
  • Working Paper

    Determination of the 2020 U.S. Citizen Voting Age Population (CVAP) Using Administrative Records and Statistical Methodology Technical Report

    October 2020

    Working Paper Number:

    CES-20-33

    This report documents the efforts of the Census Bureau's Citizen Voting-Age Population (CVAP) Internal Expert Panel (IEP) and Technical Working Group (TWG) toward the use of multiple data sources to produce block-level statistics on the citizen voting-age population for use in enforcing the Voting Rights Act. It describes the administrative, survey, and census data sources used, and the four approaches developed for combining these data to produce CVAP estimates. It also discusses other aspects of the estimation process, including how records were linked across the multiple data sources, and the measures taken to protect the confidentiality of the data.
    View Full Paper PDF
  • Working Paper

    The Management and Organizational Practices Survey (MOPS): Collection and Processing

    December 2018

    Working Paper Number:

    CES-18-51

    The U.S. Census Bureau partnered with a team of external researchers to conduct the first-ever large-scale survey of management practices in the United States, the Management and Organizational Practices Survey (MOPS), for reference year 2010. With the help of the research team, the Census Bureau expanded and improved the survey for a second wave for reference year 2015. The MOPS is a supplement to the Annual Survey of Manufacturing (ASM), and so the collection and processing strategy for the MOPS built on the methodology for the ASM, while differing on key dimensions to address the unique nature of management relative to other business data. This paper provides detail on the mail strategy pursued for the MOPS, the collection methods for paper and electronic responses, the processing and estimation procedures, and the official Census Bureau data releases. This detail is useful for all those who have interest in using the MOPS for research purposes, those wishing to understand the MOPS data more deeply, and those with an interest in survey methodology.
    View Full Paper PDF
  • Working Paper

    Reservation Nonemployer and Employer Establishments: Data from U.S. Census Longitudinal Business Databases

    December 2018

    Working Paper Number:

    CES-18-50

    The presence of businesses on American Indian reservations has been difficult to analyze due to limited data. Akee, Mykerezi, and Todd (AMT; 2017) geocoded confidential data from the U.S. Census Longitudinal Business Database to identify whether employer establishments were located on or off American Indian reservations and then compared federally recognized reservations and nearby county areas with respect to their per capita number of employers and jobs. We use their methods and the U.S. Census Integrated Longitudinal Business Database to develop parallel results for nonemployer establishments and for the combination of employer and nonemployer establishments. Similar to AMT's findings, we find that reservations and nearby county areas have a similar sectoral distribution of nonemployer and nonemployer-plus-employer establishments, but reservations have significantly fewer of them in nearly all sectors, especially when the area population is below 15,000. By contrast to AMT, the average size of reservation nonemployer establishments, as measured by revenue (instead of the jobs measure AMT used for employers), is smaller than the size of nonemployers in nearby county areas, and this is true in most industries as well. The most significant exception is in the retail sector. Geographic and demographic factors, such as population density and per capita income, statistically account for only a small portion of these differences. However, when we assume that nonemployer establishments create the equivalent of one job and use combined employer-plus-nonemployer jobs to measure establishment size, the employer job numbers dominate and we parallel AMT's finding that, due to large job counts in the Arts/Entertainment/Recreation and Public Administration sectors, reservations on average have slightly more jobs per resident than nearby county areas.
    View Full Paper PDF
  • Working Paper

    Foreign-Born and Native-Born Migration in the U.S.: Evidence from IRS Administrative and Census Survey Records

    July 2018

    Working Paper Number:

    carra-2018-07

    This paper details efforts to link administrative records from the Internal Revenue Service (IRS) to American Community Survey (ACS) and 2010 Census microdata for the study of migration among foreign-born and native-born populations in the United States. Specifically, we (1) document our linkage strategy and methodology for inferring migration in IRS records; (2) model selection into and survival across IRS records to determine suitability for research applications; and (3) gauge the efficacy of the IRS records by demonstrating how they can be used to validate and potentially improve migration responses for native-born and foreign-born respondents in ACS microdata. Our results show little evidence of selection or survival bias in the IRS records, suggesting broad generalizability to the nation as a whole. Moreover, we find that the combined IRS 1040, 1099, and W2 records may provide important information on populations, such as the foreign-born, that may be difficult to reach with traditional Census Bureau surveys. Finally, while preliminary, the results of our comparison of IRS and ACS migration responses shows that IRS records may be useful in improving ACS migration measurement for respondents whose migration response is proxy, allocated, or imputed. Taking these results together, we discuss the potential application of our longitudinal IRS dataset to innovations in migration research on both the native-born and foreign-born populations of the United States.
    View Full Paper PDF