CREAT: Census Research Exploration and Analysis Tool

Papers Containing Keywords(s): 'statistical'

The following papers contain search terms that you selected. From the papers listed below, you can navigate to the PDF, the profile page for that working paper, or see all the working papers written by an author. You can also explore tags, keywords, and authors that occur frequently within these papers.
Click here to search again

Frequently Occurring Concepts within this Search

Center for Economic Studies - 43

National Science Foundation - 39

Internal Revenue Service - 35

Bureau of Labor Statistics - 33

American Community Survey - 31

Cornell University - 31

Current Population Survey - 28

Census Bureau Disclosure Review Board - 27

Social Security Administration - 24

North American Industry Classification System - 24

Survey of Income and Program Participation - 23

Longitudinal Employer Household Dynamics - 23

Standard Industrial Classification - 18

Research Data Center - 18

Longitudinal Business Database - 17

Service Annual Survey - 17

Employer Identification Numbers - 16

Social Security Number - 15

Protected Identification Key - 15

Federal Statistical Research Data Center - 15

Annual Survey of Manufactures - 15

Bureau of Economic Analysis - 15

Alfred P Sloan Foundation - 15

Longitudinal Research Database - 15

Economic Census - 14

Decennial Census - 13

Quarterly Census of Employment and Wages - 13

Disclosure Review Board - 13

Quarterly Workforce Indicators - 13

Business Register - 13

Social Security - 12

Census of Manufactures - 12

Total Factor Productivity - 12

Ordinary Least Squares - 12

County Business Patterns - 12

Metropolitan Statistical Area - 11

2010 Census - 11

Special Sworn Status - 11

Unemployment Insurance - 10

Person Validation System - 9

Office of Management and Budget - 9

Business Dynamics Statistics - 9

Person Identification Validation System - 8

Master Address File - 8

National Longitudinal Survey of Youth - 8

Statistics Canada - 8

National Center for Health Statistics - 8

Cornell Institute for Social and Economic Research - 8

LEHD Program - 8

Personally Identifiable Information - 7

Local Employment Dynamics - 7

National Bureau of Economic Research - 7

Standard Statistical Establishment List - 7

Public Use Micro Sample - 7

Duke University - 7

American Statistical Association - 7

Chicago Census Research Data Center - 7

Census Bureau Longitudinal Business Database - 7

Social and Economic Supplement - 6

Housing and Urban Development - 6

Federal Statistical System - 6

National Academy of Sciences - 6

Census Bureau Business Register - 6

Department of Labor - 6

Detailed Earnings Records - 6

Federal Reserve Bank - 6

PSID - 6

Department of Agriculture - 5

Temporary Assistance for Needy Families - 5

Supplemental Nutrition Assistance Program - 5

Department of Housing and Urban Development - 5

Bureau of Labor - 5

1940 Census - 5

Census Edited File - 5

Characteristics of Business Owners - 5

W-2 - 5

Computer Assisted Personal Interview - 5

Census of Manufacturing Firms - 5

National Institutes of Health - 5

Department of Commerce - 5

Sloan Foundation - 5

Permanent Plant Number - 5

Department of Health and Human Services - 4

Centers for Disease Control and Prevention - 4

National Research Council - 4

MAFID - 4

Department of Education - 4

Cobb-Douglas - 4

Department of Economics - 4

United States Census Bureau - 4

Some Other Race - 4

Financial, Insurance and Real Estate Industries - 4

Health and Retirement Study - 4

American Economic Association - 4

Small Business Administration - 4

Individual Characteristics File - 4

National Health Interview Survey - 4

National Institute on Aging - 4

Summary Earnings Records - 4

Company Organization Survey - 4

Journal of Economic Literature - 4

Economic Research Service - 3

Food and Nutrition Service - 3

COVID-19 - 3

Establishment Micro Properties - 3

Employment History File - 3

MAF-ARF - 3

Stanford University - 3

Annual Business Survey - 3

University of Texas - 3

Federal Insurance Contribution Act - 3

CPS ASEC - 3

Postal Service - 3

COVID - 3

Agency for Healthcare Research and Quality - 3

Urban Institute - 3

American Housing Survey - 3

LEHD Origin-Destination Employment Statistics - 3

University of Michigan - 3

Employer Characteristics File - 3

North American Industry Classi - 3

Securities and Exchange Commission - 3

Multiple Worksite Report - 3

Review of Economics and Statistics - 3

Organization for Economic Cooperation and Development - 3

University of Maryland - 3

survey - 49

data - 43

respondent - 41

population - 34

estimating - 33

census bureau - 32

report - 28

agency - 28

microdata - 27

statistician - 24

data census - 22

datasets - 20

analysis - 19

economist - 18

census data - 18

aggregate - 18

percentile - 17

estimation - 17

use census - 16

earnings - 16

disclosure - 16

researcher - 15

confidentiality - 15

database - 14

privacy - 14

socioeconomic - 12

study - 12

research - 12

quarterly - 12

payroll - 12

employed - 12

workforce - 12

record - 12

public - 12

econometric - 12

statistical agencies - 12

longitudinal - 12

employ - 11

estimator - 11

federal - 10

imputation - 10

expenditure - 10

labor - 10

salary - 10

aggregation - 10

information - 10

poverty - 9

macroeconomic - 9

resident - 9

manufacturing - 9

economic census - 9

statistical disclosure - 9

publicly - 9

recession - 9

employee - 9

research census - 8

prevalence - 8

censuses surveys - 8

paper census - 8

census employment - 8

census survey - 8

census disclosure - 8

labor statistics - 8

sector - 8

market - 8

inference - 8

eligibility - 7

income data - 7

www census - 7

individuals census - 7

revenue - 7

sample - 7

survey data - 7

sampling - 7

enterprise - 7

sale - 7

census research - 7

employee data - 7

2010 census - 6

disadvantaged - 6

yearly - 6

ethnicity - 6

average - 6

assessed - 6

trend - 6

production - 6

company - 6

analyst - 6

empirical - 6

establishment - 6

industrial - 6

available census - 5

hispanic - 5

enrollment - 5

minority - 5

discrepancy - 5

gdp - 5

employment statistics - 5

census years - 5

social - 5

reporting - 5

model - 5

incorporated - 5

business data - 5

aging - 5

growth - 5

welfare - 4

medicaid - 4

eligible - 4

estimates census - 4

sample census - 4

enrolled - 4

state - 4

irs - 4

productivity growth - 4

productivity measures - 4

measures productivity - 4

imputation model - 4

census responses - 4

ssa - 4

survey income - 4

demand - 4

household surveys - 4

mobility - 4

income year - 4

citizen - 4

regression - 4

policymakers - 4

assessing - 4

employment data - 4

metropolitan - 4

tenure - 4

measure - 4

employer household - 4

corporation - 4

census use - 3

provided census - 3

enrollee - 3

housing - 3

residential - 3

neighborhood - 3

ethnic - 3

disparity - 3

disability - 3

efficiency - 3

bias - 3

decade - 3

population survey - 3

matching - 3

racial - 3

intergenerational - 3

residence - 3

regressing - 3

information census - 3

corporate - 3

unobserved - 3

competitor - 3

surveys censuses - 3

economic statistics - 3

department - 3

linked census - 3

employment count - 3

work census - 3

regressors - 3

coverage - 3

produce - 3

family - 3

establishments data - 3

longitudinal employer - 3

workforce indicators - 3

poorer - 3

worker - 3

merger - 3

classified - 3

industrial classification - 3

classification - 3

classifying - 3

Viewing papers 41 through 50 of 100


  • Working Paper

    Response Error & the Medicaid undercount in the CPS

    December 2016

    Working Paper Number:

    carra-2016-11

    The Current Population Survey Annual Social and Economic Supplement (CPS ASEC) is an important source for estimates of the uninsured population. Previous research has shown that survey estimates produce an undercount of beneficiaries compared to Medicaid enrollment records. We extend past work by examining the Medicaid undercount in the 2007-2011 CPS ASEC compared to enrollment data from the Medicaid Statistical Information System for calendar years 2006-2010. By linking individuals across datasets, we analyze two types of response error regarding Medicaid enrollment - false negative error and false positive error. We use regression analysis to identify factors associated with these two types of response error in the 2011 CPS ASEC. We find that the Medicaid undercount was between 22 and 31 percent from 2007 to 2011. In 2011, the false negative rate was 40 percent, and 27 percent of Medicaid reports in CPS ASEC were false positives. False negative error is associated with the duration of enrollment in Medicaid, enrollment in Medicare and private insurance, and Medicaid enrollment in the survey year. False positive error is associated with enrollment in Medicare and shared Medicaid coverage in the household. We discuss implications for survey reports of health insurance coverage and for estimating the uninsured population.
    View Full Paper PDF
  • Working Paper

    Using Partially Synthetic Microdata to Protect Sensitive Cells in Business Statistics

    February 2016

    Working Paper Number:

    CES-16-10

    We describe and analyze a method that blends records from both observed and synthetic microdata into public-use tabulations on establishment statistics. The resulting tables use synthetic data only in potentially sensitive cells. We describe different algorithms, and present preliminary results when applied to the Census Bureau's Business Dynamics Statistics and Synthetic Longitudinal Business Database, highlighting accuracy and protection afforded by the method when compared to existing public-use tabulations (with suppressions).
    View Full Paper PDF
  • Working Paper

    Food and Agricultural Industries: Opportunities for Improving Measurement and Reporting

    January 2016

    Working Paper Number:

    CES-16-58

    We measure one component of off-farm food and agricultural industries using establishment level microdata in the federal statistical system. We focus on services for crop production, and compare measures of firm and employment dynamics in this sector during the period 1992-2012 with county-level publicly available data for the same measures. Based on differences across data sources, we establish new facts regarding the evolution of food and agricultural industries, and demonstrate the value of working with confidential microdata. In addition to the data and results we present, we highlight possibilities for collaboration across universities and federal agencies to improve reporting in other segments of food and agricultural industries.
    View Full Paper PDF
  • Working Paper

    Measuring Cross-Country Differences in Misallocation

    January 2016

    Working Paper Number:

    CES-16-50R

    We describe differences between the commonly used version of the U.S. Census of Manufactures available at the RDCs and what establishments themselves report. The originally reported data has substantially more dispersion in measured establishment productivity. Measured allocative efficiency is substantially higher in the cleaned data than the raw data: 4x higher in 2002, 20x in 2007, and 80x in 2012. Many of the important editing strategies at the Census, including industry analysts' manual edits and edits using tax records, are infeasible in non-U.S. datasets. We describe a new Bayesian approach for editing and imputation that can be used across contexts.
    View Full Paper PDF
  • Working Paper

    The Timing of Teenage Births: Estimating the Effect on High School Graduation and Later Life Outcomes

    January 2016

    Working Paper Number:

    CES-16-39R

    We examine the long-term outcomes for a population of teenage mothers who give birth to their children around the end of their high school year. We compare the mothers whose high school education was interrupted by childbirth, because the child was born before her expected graduation date to mothers who did not experience the same disruption to their education. We find that mothers who give birth during the school year are seven percent less likely to graduate from high school, are less likely to be married, and have more children than their counterparts who gave birth just a few months later. The labor market outcomes for these two sets of teenage mothers are not statistically different, but with a lower likelihood of marriage and more children, the households of the treated mothers are more likely to fall below the poverty threshold. While differences in educational attainment have narrowed over time, the differences in labor market outcomes and family structure have remained stable.
    View Full Paper PDF
  • Working Paper

    Simultaneous Edit-Imputation for Continuous Microdata

    December 2015

    Working Paper Number:

    CES-15-44

    Many statistical organizations collect data that are expected to satisfy linear constraints; as examples, component variables should sum to total variables, and ratios of pairs of variables should be bounded by expert-specified constants. When reported data violate constraints, organizations identify and replace values potentially in error in a process known as edit-imputation. To date, most approaches separate the error localization and imputation steps, typically using optimization methods to identify the variables to change followed by hot deck imputation. We present an approach that fully integrates editing and imputation for continuous microdata under linear constraints. Our approach relies on a Bayesian hierarchical model that includes (i) a flexible joint probability model for the underlying true values of the data with support only on the set of values that satisfy all editing constraints, (ii) a model for latent indicators of the variables that are in error, and (iii) a model for the reported responses for variables in error. We illustrate the potential advantages of the Bayesian editing approach over existing approaches using simulation studies. We apply the model to edit faulty data from the 2007 U.S. Census of Manufactures. Supplementary materials for this article are available online.
    View Full Paper PDF
  • Working Paper

    USING IMPUTATION TECHNIQUES TO EVALUATE STOPPING RULES IN ADAPTIVE SURVEY DESIGN

    October 2014

    Working Paper Number:

    CES-14-40

    Adaptive Design methods for social surveys utilize the information from the data as it is collected to make decisions about the sampling design. In some cases, the decision is either to continue or stop the data collection. We evaluate this decision by proposing measures to compare the collected data with follow-up samples. The options are assessed by imputation of the nonrespondents under different missingness scenarios, including Missing Not at Random. The variation in the utility measures is compared to the cost induced by the follow-up sample sizes. We apply the proposed method to the 2007 U.S. Census of Manufacturers.
    View Full Paper PDF
  • Working Paper

    NOISE INFUSION AS A CONFIDENTIALITY PROTECTION MEASURE FOR GRAPH-BASED STATISTICS

    September 2014

    Working Paper Number:

    CES-14-30

    We use the bipartite graph representation of longitudinally linked em-ployer-employee data, and the associated projections onto the employer and em-ployee nodes, respectively, to characterize the set of potential statistical summar-ies that the trusted custodian might produce. We consider noise infusion as the primary confidentiality protection method. We show that a relatively straightfor-ward extension of the dynamic noise-infusion method used in the U.S. Census Bureau's Quarterly Workforce Indicators can be adapted to provide the same confidentiality guarantees for the graph-based statistics: all inputs have been modified by a minimum percentage deviation (i.e., no actual respondent data are used) and, as the number of entities contributing to a particular statistic increases, the accuracy of that statistic approaches the unprotected value. Our method also ensures that the protected statistics will be identical in all releases based on the same inputs.
    View Full Paper PDF
  • Working Paper

    Within and Across County Variation in SNAP Misreporting: Evidence from Linked ACS and Administrative Records

    July 2014

    Working Paper Number:

    carra-2014-05

    This paper examines sub-state spatial and temporal variation in misreporting of participation in the Supplemental Nutrition Assistance Program (SNAP) using several years of the American Community Survey linked to SNAP administrative records from New York (2008-2010) and Texas (2006-2009). I calculate county false-negative (FN) and false-positive (FP) rates for each year of observation and find that, within a given state and year, there is substantial heterogeneity in FN rates across counties. In addition, I find evidence that FN rates (but not FP rates) persist over time within counties. This persistence in FN rates is strongest among more populous counties, suggesting that when noise from sampling variation is not an issue, some counties have consistently high FN rates while others have consistently low FN rates. This finding is important for understanding how misreporting might bias estimates of sub-state SNAP participation rates, changes in those participation rates, and effects of program participation. This presentation was given at the CARRA Seminar, June 27, 2013
    View Full Paper PDF
  • Working Paper

    A FIRST STEP TOWARDS A GERMAN SYNLBD: CONSTRUCTING A GERMAN LONGITUDINAL BUSINESS DATABASE

    February 2014

    Working Paper Number:

    CES-14-13

    One major criticism against the use of synthetic data has been that the efforts necessary to generate useful synthetic data are so in- tense that many statistical agencies cannot afford them. We argue many lessons in this evolving field have been learned in the early years of synthetic data generation, and can be used in the development of new synthetic data products, considerably reducing the required in- vestments. The final goal of the project described in this paper will be to evaluate whether synthetic data algorithms developed in the U.S. to generate a synthetic version of the Longitudinal Business Database (LBD) can easily be transferred to generate a similar data product for other countries. We construct a German data product with infor- mation comparable to the LBD - the German Longitudinal Business Database (GLBD) - that is generated from different administrative sources at the Institute for Employment Research, Germany. In a fu- ture step, the algorithms developed for the synthesis of the LBD will be applied to the GLBD. Extensive evaluations will illustrate whether the algorithms provide useful synthetic data without further adjustment. The ultimate goal of the project is to provide access to multiple synthetic datasets similar to the SynLBD at Cornell to enable comparative studies between countries. The Synthetic GLBD is a first step towards that goal.
    View Full Paper PDF