CREAT: Census Research Exploration and Analysis Tool

Papers Containing Keywords(s): 'statistical'

The following papers contain search terms that you selected. From the papers listed below, you can navigate to the PDF, the profile page for that working paper, or see all the working papers written by an author. You can also explore tags, keywords, and authors that occur frequently within these papers.
Click here to search again

Frequently Occurring Concepts within this Search

Center for Economic Studies - 43

National Science Foundation - 39

Internal Revenue Service - 35

Bureau of Labor Statistics - 33

American Community Survey - 31

Cornell University - 31

Current Population Survey - 28

Census Bureau Disclosure Review Board - 27

Social Security Administration - 24

North American Industry Classification System - 24

Survey of Income and Program Participation - 23

Longitudinal Employer Household Dynamics - 23

Standard Industrial Classification - 18

Research Data Center - 18

Longitudinal Business Database - 17

Service Annual Survey - 17

Employer Identification Numbers - 16

Social Security Number - 15

Protected Identification Key - 15

Federal Statistical Research Data Center - 15

Annual Survey of Manufactures - 15

Bureau of Economic Analysis - 15

Alfred P Sloan Foundation - 15

Longitudinal Research Database - 15

Economic Census - 14

Decennial Census - 13

Quarterly Census of Employment and Wages - 13

Disclosure Review Board - 13

Quarterly Workforce Indicators - 13

Business Register - 13

Social Security - 12

Census of Manufactures - 12

Total Factor Productivity - 12

Ordinary Least Squares - 12

County Business Patterns - 12

Metropolitan Statistical Area - 11

2010 Census - 11

Special Sworn Status - 11

Unemployment Insurance - 10

Person Validation System - 9

Office of Management and Budget - 9

Business Dynamics Statistics - 9

Person Identification Validation System - 8

Master Address File - 8

National Longitudinal Survey of Youth - 8

Statistics Canada - 8

National Center for Health Statistics - 8

Cornell Institute for Social and Economic Research - 8

LEHD Program - 8

Personally Identifiable Information - 7

Local Employment Dynamics - 7

National Bureau of Economic Research - 7

Standard Statistical Establishment List - 7

Public Use Micro Sample - 7

Duke University - 7

American Statistical Association - 7

Chicago Census Research Data Center - 7

Census Bureau Longitudinal Business Database - 7

Social and Economic Supplement - 6

Housing and Urban Development - 6

Federal Statistical System - 6

National Academy of Sciences - 6

Census Bureau Business Register - 6

Department of Labor - 6

Detailed Earnings Records - 6

Federal Reserve Bank - 6

PSID - 6

Department of Agriculture - 5

Temporary Assistance for Needy Families - 5

Supplemental Nutrition Assistance Program - 5

Department of Housing and Urban Development - 5

Bureau of Labor - 5

1940 Census - 5

Census Edited File - 5

Characteristics of Business Owners - 5

W-2 - 5

Computer Assisted Personal Interview - 5

Census of Manufacturing Firms - 5

National Institutes of Health - 5

Department of Commerce - 5

Sloan Foundation - 5

Permanent Plant Number - 5

Department of Health and Human Services - 4

Centers for Disease Control and Prevention - 4

National Research Council - 4

MAFID - 4

Department of Education - 4

Cobb-Douglas - 4

Department of Economics - 4

United States Census Bureau - 4

Some Other Race - 4

Financial, Insurance and Real Estate Industries - 4

Health and Retirement Study - 4

American Economic Association - 4

Small Business Administration - 4

Individual Characteristics File - 4

National Health Interview Survey - 4

National Institute on Aging - 4

Summary Earnings Records - 4

Company Organization Survey - 4

Journal of Economic Literature - 4

Economic Research Service - 3

Food and Nutrition Service - 3

COVID-19 - 3

Establishment Micro Properties - 3

Employment History File - 3

MAF-ARF - 3

Stanford University - 3

Annual Business Survey - 3

University of Texas - 3

Federal Insurance Contribution Act - 3

CPS ASEC - 3

Postal Service - 3

COVID - 3

Agency for Healthcare Research and Quality - 3

Urban Institute - 3

American Housing Survey - 3

LEHD Origin-Destination Employment Statistics - 3

University of Michigan - 3

Employer Characteristics File - 3

North American Industry Classi - 3

Securities and Exchange Commission - 3

Multiple Worksite Report - 3

Review of Economics and Statistics - 3

Organization for Economic Cooperation and Development - 3

University of Maryland - 3

survey - 49

data - 43

respondent - 41

population - 34

estimating - 33

census bureau - 32

report - 28

agency - 28

microdata - 27

statistician - 24

data census - 22

datasets - 20

analysis - 19

economist - 18

census data - 18

aggregate - 18

percentile - 17

estimation - 17

use census - 16

earnings - 16

disclosure - 16

researcher - 15

confidentiality - 15

database - 14

privacy - 14

socioeconomic - 12

study - 12

research - 12

quarterly - 12

payroll - 12

employed - 12

workforce - 12

record - 12

public - 12

econometric - 12

statistical agencies - 12

longitudinal - 12

employ - 11

estimator - 11

federal - 10

imputation - 10

expenditure - 10

labor - 10

salary - 10

aggregation - 10

information - 10

poverty - 9

macroeconomic - 9

resident - 9

manufacturing - 9

economic census - 9

statistical disclosure - 9

publicly - 9

recession - 9

employee - 9

research census - 8

prevalence - 8

censuses surveys - 8

paper census - 8

census employment - 8

census survey - 8

census disclosure - 8

labor statistics - 8

sector - 8

market - 8

inference - 8

eligibility - 7

income data - 7

www census - 7

individuals census - 7

revenue - 7

sample - 7

survey data - 7

sampling - 7

enterprise - 7

sale - 7

census research - 7

employee data - 7

2010 census - 6

disadvantaged - 6

yearly - 6

ethnicity - 6

average - 6

assessed - 6

trend - 6

production - 6

company - 6

analyst - 6

empirical - 6

establishment - 6

industrial - 6

available census - 5

hispanic - 5

enrollment - 5

minority - 5

discrepancy - 5

gdp - 5

employment statistics - 5

census years - 5

social - 5

reporting - 5

model - 5

incorporated - 5

business data - 5

aging - 5

growth - 5

welfare - 4

medicaid - 4

eligible - 4

estimates census - 4

sample census - 4

enrolled - 4

state - 4

irs - 4

productivity growth - 4

productivity measures - 4

measures productivity - 4

imputation model - 4

census responses - 4

ssa - 4

survey income - 4

demand - 4

household surveys - 4

mobility - 4

income year - 4

citizen - 4

regression - 4

policymakers - 4

assessing - 4

employment data - 4

metropolitan - 4

tenure - 4

measure - 4

employer household - 4

corporation - 4

census use - 3

provided census - 3

enrollee - 3

housing - 3

residential - 3

neighborhood - 3

ethnic - 3

disparity - 3

disability - 3

efficiency - 3

bias - 3

decade - 3

population survey - 3

matching - 3

racial - 3

intergenerational - 3

residence - 3

regressing - 3

information census - 3

corporate - 3

unobserved - 3

competitor - 3

surveys censuses - 3

economic statistics - 3

department - 3

linked census - 3

employment count - 3

work census - 3

regressors - 3

coverage - 3

produce - 3

family - 3

establishments data - 3

longitudinal employer - 3

workforce indicators - 3

poorer - 3

worker - 3

merger - 3

classified - 3

industrial classification - 3

classification - 3

classifying - 3

Viewing papers 51 through 60 of 100


  • Working Paper

    IMPROVING THE SYNTHETIC LONGITUDINAL BUSINESS DATABASE

    February 2014

    Working Paper Number:

    CES-14-12

    In most countries, national statistical agencies do not release establishment-level business microdata, because doing so represents too large a risk to establishments' confidentiality. Agencies potentially can manage these risks by releasing synthetic microdata, i.e., individual establishment records simulated from statistical models de- signed to mimic the joint distribution of the underlying observed data. Previously, we used this approach to generate a public-use version'now available for public use'of the U. S. Census Bureau's Longitudinal Business Database (LBD), a longitudinal cen- sus of establishments dating back to 1976. While the synthetic LBD has proven to be a useful product, we now seek to improve and expand it by using new synthesis models and adding features. This article describes our efforts to create the second generation of the SynLBD, including synthesis procedures that we believe could be replicated in other contexts.
    View Full Paper PDF
  • Working Paper

    LOOKING BACK ON THREE YEARS OF USING THE SYNTHETIC LBD BETA

    February 2014

    Working Paper Number:

    CES-14-11

    Distributions of business data are typically much more skewed than those for household or individual data and public knowledge of the underlying units is greater. As a results, national statistical offices (NSOs) rarely release establishment or firm-level business microdata due to the risk to respondent confidentiality. One potential approach for overcoming these risks is to release synthetic data where the establishment data are simulated from statistical models designed to mimic the distributions of the real underlying microdata. The US Census Bureau's Center for Economic Studies in collaboration with Duke University, the National Institute of Statistical Sciences, and Cornell University made available a synthetic public use file for the Longitudinal Business Database (LBD) comprising more than 20 million records for all business establishment with paid employees dating back to 1976. The resulting product, dubbed the SynLBD, was released in 2010 and is the first-ever comprehensive business microdata set publicly released in the United States including data on establishments employment and payroll, birth and death years, and industrial classification. This pa- per documents the scope of projects that have requested and used the SynLBD.
    View Full Paper PDF
  • Working Paper

    EXPANDING THE ROLE OF SYNTHETIC DATA AT THE U.S. CENSUS BUREAU

    February 2014

    Working Paper Number:

    CES-14-10

    National Statistical offices (NSOs) create official statistics from data collected from survey respondents, government administrative records and other sources. The raw source data is usually considered to be confidential. In the case of the U.S. Census Bureau, confidentiality of survey and administrative records microdata is mandated by statute, and this mandate to protect confidentiality is often at odds with the needs of users to extract as much information from the data as possible. Traditional disclosure protection techniques result in official data products that do not fully utilize the information content of the underlying microdata. Typically, these products take the form of simple aggregate tabulations. In a few cases anonymized public- use micro samples are made available, but these face a growing risk of re-identification by the increasing amounts of information about individuals and firms available in the public domain. One approach for overcoming these risks is to release products based on synthetic data where values are simulated from statistical models designed to mimic the (joint) distributions of the underlying microdata. We discuss re- cent Census Bureau work to develop and deploy such products. We discuss the benefits and challenges involved with extending the scope of synthetic data products in official statistics.
    View Full Paper PDF
  • Working Paper

    SYNTHETIC DATA FOR SMALL AREA ESTIMATION IN THE AMERICAN COMMUNITY SURVEY

    April 2013

    Working Paper Number:

    CES-13-19

    Small area estimates provide a critical source of information used to study local populations. Statistical agencies regularly collect data from small areas but are prevented from releasing detailed geographical identifiers in public-use data sets due to disclosure concerns. Alternative data dissemination methods used in practice include releasing summary/aggregate tables, suppressing detailed geographic information in public-use data sets, and accessing restricted data via Research Data Centers. This research examines an alternative method for disseminating microdata that contains more geographical details than are currently being released in public-use data files. Specifically, the method replaces the observed survey values with imputed, or synthetic, values simulated from a hierarchical Bayesian model. Confidentiality protection is enhanced because no actual values are released. The method is demonstrated using restricted data from the 2005-2009 American Community Survey. The analytic validity of the synthetic data is assessed by comparing small area estimates obtained from the synthetic data with those obtained from the observed data.
    View Full Paper PDF
  • Working Paper

    Estimation of Job-to-Job Flow Rates under Partially Missing Geography

    September 2012

    Working Paper Number:

    CES-12-29

    Integration of data from different regions presents challenges for the calculation of entitylevel longitudinal statistics with a strong geographic component: for example, movements between employers, migration, business dynamics, and health statistics. In this paper, we consider the estimation of worker-level employment statistics when the geographies (in our application, US states) over which such measures are defined are partially missing. We focus on the recent pilot set of job-to-job flow statistics produced by the US Census Bureau's Longitudinal Employer- Household Dynamics (LEHD) program, which measure the frequency of worker movements between jobs and into and out of nonemployment. LEHD's coverage of the labor force gradually increases during the 1990s and 2000s because some states have a longer time series than others, so employment transitions involving missing states are only partially or not at all observed. We propose and implement a method for estimating national-level job-to-job flow statistics that involves dropping observed states to recover the relationship between missing states and directly tabulated job-to-job flow rates. Using the estimated relationship between the observable characteristics of the missing states and changes in the employment measures, we provide estimates of the rates of job-to-job, and job-to-nonemployment, job-to-nonemploymentto- job flows were all states uniformly available.
    View Full Paper PDF
  • Working Paper

    Dynamically Consistent Noise Infusion and Partially Synthetic Data as Confidentiality Protection Measures for Related Time Series

    July 2012

    Working Paper Number:

    CES-12-13

    The Census Bureau's Quarterly Workforce Indicators (QWI) provide detailed quarterly statistics on employment measures such as worker and job flows, tabulated by worker characteristics in various combinations. The data are released for several levels of NAICS industries and geography, the lowest aggregation of the latter being counties. Disclosure avoidance methods are required to protect the information about individuals and businesses that contribute to the underlying data. The QWI disclosure avoidance mechanism we describe here relies heavily on the use of noise infusion through a permanent multiplicative noise distortion factor, used for magnitudes, counts, differences and ratios. There is minimal suppression and no complementary suppressions. To our knowledge, the release in 2003 of the QWI was the first large-scale use of noise infusion in any official statistical product. We show that the released statistics are analytically valid along several critical dimensions { measures are unbiased and time series properties are preserved. We provide an analysis of the degree to which confidentiality is protected. Furthermore, we show how the judicious use of synthetic data, injected into the tabulation process, can completely eliminate suppressions, maintain analytical validity, and increase the protection of the underlying confidential data.
    View Full Paper PDF
  • Working Paper

    LEHD Data Documentation LEHD-OVERVIEW-S2008-rev1

    December 2011

    Working Paper Number:

    CES-11-43

    View Full Paper PDF
  • Working Paper

    Firm Market Power and the Earnings Distribution

    December 2011

    Authors: Douglas Webber

    Working Paper Number:

    CES-11-41

    Using the Longitudinal Employer Household Dynamics (LEHD) data from the United States Census Bureau, I compute firm-level measures of labor market (monopsony) power. To generate these measures, I extend the dynamic model proposed by Manning (2003) and estimate the labor supply elasticity facing each private non-farm firm in the US. While a link between monopsony power and earnings has traditionally been assumed, I provide the first direct evidence of the positive relationship between a firm\'s labor supply elasticity and the earnings of its workers. I also contrast the semistructural method with the more traditional use of concentration ratios to measure a firm\'s labor market power. In addition, I provide several alternative measures of labor market power which account for potential threats to identification such as endogenous mobility. Finally, I construct a counterfactual earnings distribution which allows the effects of firm market power to vary across the earnings distribution. I estimate the average firm\'s labor supply elasticity to be 1.08, however my findings suggest there to be significant variability in the distribution of firm market power across US firms, and that dynamic monopsony models are superior to the use of concentration ratios in evaluating a firm\'s labor market power. I find that a one-unit increase in the labor supply elasticity to the firm is associated with wage gains of between 5 and 18 percent. While nontrivial, these estimates imply that firms do not fully exercise their labor market power over their workers. Furthermore, I find that the negative earnings impact of a firm\'s market power is strongest in the lower half of the earnings distribution, and that a one standard deviation increase in firms\' labor supply elasticities reduces the variance of the earnings distribution by 9 percent.
    View Full Paper PDF
  • Working Paper

    Black-White Differences in Intergenerational Economic Mobility in the U.S.

    December 2011

    Working Paper Number:

    CES-11-40

    Traditional measures of intergenerational mobility such as the intergenerational elasticity are not useful for inferences concerning group differences in mobility with respect to the pooled income distribution. This paper uses transition probabilities and measures of 'directional rank mobility' that can identify inter-racial differences in intergenerational mobility. The study uses two data sources including one that contains social security earnings for a large intergenerational sample. I find that recent cohorts of blacks are not only significantly less upwardly mobile but also significantly more downwardly mobile than whites. This implies a steady-state distribution in which there is no racial convergence in income. A descriptive analysis using covariates reveals that test scores in adolescence can explain much of the racial difference in both upward and downward mobility. Family structure can account for some of the racial gap in upward mobility but not downward mobility. Completed schooling and parental wealth also appear to account for some of the racial gaps in intergenerational mobility.
    View Full Paper PDF
  • Working Paper

    Using the Survey of Plant Capacity to Measure Capital Utilization

    July 2011

    Working Paper Number:

    CES-11-19

    Most capital in the United States is idle much of the time. By some measures, the average workweek of capital in U.S. manufacturing is as low as 55 hours per 168 hour week. The level and variability of capital utilization has important implications for understanding both the level of production and its cyclical fluctuations. This paper investigates a number of issues relating to aggregation of capital utilization measures from the Survey of Plant Capacity and makes recommendations on expanding and improving the published statistics deriving from the Survey of Plant Capacity. The paper documents a number of facts about properties of capital utilization. First, after growing for decades, capital utilization started to fall in mid 1990s. Second, capital utilization is a useful predictor of changes in capacity utilization and other factors of production. Third, adjustment of productivity measures for variable capital utilization improves statistical and economic properties of these measures. Fourth, the paper constructs weights to aggregate firm level capital utilization rates to industry and economy level, which is the major enhancement to available data.
    View Full Paper PDF