CREAT: Census Research Exploration and Analysis Tool

Papers Containing Keywords(s): 'data'

The following papers contain search terms that you selected. From the papers listed below, you can navigate to the PDF, the profile page for that working paper, or see all the working papers written by an author. You can also explore tags, keywords, and authors that occur frequently within these papers.
Click here to search again

Frequently Occurring Concepts within this Search

National Science Foundation - 36

Internal Revenue Service - 36

American Community Survey - 35

Center for Economic Studies - 35

Social Security Administration - 29

Service Annual Survey - 29

Research Data Center - 27

Current Population Survey - 24

Protected Identification Key - 22

Bureau of Labor Statistics - 22

Longitudinal Employer Household Dynamics - 20

North American Industry Classification System - 20

Cornell University - 20

Survey of Income and Program Participation - 19

Census Bureau Disclosure Review Board - 18

2010 Census - 18

Decennial Census - 17

Economic Census - 17

Social Security Number - 16

Person Validation System - 16

Master Address File - 16

Business Register - 16

Longitudinal Business Database - 16

Social Security - 15

Employer Identification Numbers - 14

Standard Industrial Classification - 14

Quarterly Workforce Indicators - 13

Disclosure Review Board - 12

Center for Administrative Records Research and Applications - 12

Special Sworn Status - 12

Person Identification Validation System - 11

Personally Identifiable Information - 11

Administrative Records - 11

Bureau of Economic Analysis - 11

Housing and Urban Development - 10

Census Bureau Business Register - 10

Alfred P Sloan Foundation - 10

Annual Survey of Manufactures - 10

Longitudinal Research Database - 10

National Opinion Research Center - 10

Department of Housing and Urban Development - 9

Indian Health Service - 9

National Center for Health Statistics - 9

Standard Statistical Establishment List - 9

County Business Patterns - 9

Business Dynamics Statistics - 9

Chicago Census Research Data Center - 9

MAFID - 8

SSA Numident - 8

Federal Statistical Research Data Center - 8

Computer Assisted Personal Interview - 7

Statistics Canada - 7

Quarterly Census of Employment and Wages - 7

Metropolitan Statistical Area - 7

Duke University - 7

American Statistical Association - 7

Public Use Micro Sample - 7

Census Bureau Master Address File - 6

Individual Taxpayer Identification Numbers - 6

Indian Housing Information Center - 6

Agency for Healthcare Research and Quality - 6

American Housing Survey - 6

Company Organization Survey - 6

DOB - 6

Unemployment Insurance - 6

Medicaid Services - 6

Census of Manufactures - 6

Postal Service - 6

LEHD Program - 6

Supplemental Nutrition Assistance Program - 5

Sloan Foundation - 5

Census Numident - 5

Census Bureau Person Identification Validation System - 5

Some Other Race - 5

National Institute on Aging - 5

University of Michigan - 5

Small Business Administration - 5

Office of Management and Budget - 5

Cornell Institute for Social and Economic Research - 5

PIKed - 5

University of Chicago - 5

American Economic Association - 5

Federal Reserve Bank - 5

National Bureau of Economic Research - 5

Local Employment Dynamics - 5

Permanent Plant Number - 5

Journal of Economic Literature - 5

Ordinary Least Squares - 4

1940 Census - 4

W-2 - 4

Census Edited File - 4

National Institutes of Health - 4

Health and Retirement Study - 4

National Longitudinal Survey of Youth - 4

Census of Manufacturing Firms - 4

Probability Density Function - 4

Minnesota Population Center - 4

Center for Administrative Records Research - 4

Organization for Economic Cooperation and Development - 4

Characteristics of Business Owners - 4

Total Factor Productivity - 4

Federal Insurance Contribution Act - 3

Social and Economic Supplement - 3

ASEC - 3

Adjusted Gross Income - 3

Temporary Assistance for Needy Families - 3

Geographic Information Systems - 3

Department of Economics - 3

COVID-19 - 3

National Income and Product Accounts - 3

Bureau of Labor - 3

Centers for Medicare - 3

Census Bureau Longitudinal Business Database - 3

Centers for Disease Control and Prevention - 3

Employer Characteristics File - 3

Department of Health and Human Services - 3

National Research Council - 3

Computer Assisted Telephone Interviews and Computer Assisted Personal Interviews - 3

CATI - 3

Census Bureau Center for Economic Studies - 3

Census 2000 - 3

Office of Personnel Management - 3

Census Bureau Business Dynamics Statistics - 3

COMPUSTAT - 3

Securities and Exchange Commission - 3

survey - 53

respondent - 44

statistical - 43

microdata - 41

datasets - 38

census bureau - 36

record - 36

agency - 35

data census - 31

census data - 26

estimating - 23

population - 23

database - 23

report - 22

disclosure - 19

analysis - 19

statistician - 17

confidentiality - 16

matching - 16

research - 16

survey data - 15

information - 15

privacy - 15

imputation - 15

researcher - 15

aggregate - 14

use census - 12

census survey - 12

census research - 12

statistical agencies - 12

payroll - 11

study - 11

estimation - 10

earnings - 10

sampling - 10

sample - 10

coverage - 10

public - 10

records census - 10

linkage - 10

employee - 10

workforce - 10

research census - 10

publicly - 9

resident - 9

identifier - 9

census records - 9

quarterly - 9

economic census - 9

business data - 9

matched - 9

economist - 9

sector - 9

2010 census - 8

assessed - 8

federal - 8

statistical disclosure - 8

employed - 8

enterprise - 8

longitudinal - 8

employment data - 8

employee data - 8

ssa - 7

census years - 7

residential - 7

residence - 7

household surveys - 7

reporting - 7

census use - 7

aggregation - 7

inference - 7

associate - 7

econometric - 7

estimator - 6

enrollment - 6

irs - 6

income data - 6

ethnicity - 6

race census - 6

census employment - 6

department - 6

work census - 6

information census - 6

recession - 6

surveys censuses - 6

percentile - 6

censuses surveys - 6

sale - 6

expenditure - 6

employ - 6

model - 6

census file - 6

industrial - 6

minority - 5

salary - 5

census linked - 5

citizen - 5

provided census - 5

race - 5

state - 5

housing - 5

assessing - 5

housing survey - 5

establishments data - 5

market - 5

analyst - 5

social - 5

worker - 5

manufacturing - 5

macroeconomic - 5

average - 4

survey income - 4

population survey - 4

census disclosure - 4

income individuals - 4

tax - 4

taxpayer - 4

geographic - 4

linked census - 4

survey households - 4

hispanic - 4

census 2020 - 4

home - 4

individuals census - 4

imputation model - 4

incorporated - 4

policymakers - 4

gdp - 4

employment statistics - 4

establishment - 4

investment - 4

labor - 4

trend - 4

earner - 3

household income - 3

1040 - 3

environmental - 3

impact - 3

disparity - 3

discrepancy - 3

racial - 3

empirical - 3

classification - 3

prevalence - 3

apartment - 3

unobserved - 3

organizational - 3

acquisition - 3

economic statistics - 3

classifying - 3

employer household - 3

imputed - 3

ancestry - 3

ethnic - 3

bias - 3

census responses - 3

worker demographics - 3

production - 3

manufacturer - 3

inventory - 3

employment dynamics - 3

workforce indicators - 3

classified - 3

measures employment - 3

employment measures - 3

firm data - 3

company - 3

Viewing papers 71 through 80 of 94


  • Working Paper

    Measuring Inequality Using Censored Data: A Multiple Imputation Approach

    April 2009

    Working Paper Number:

    CES-09-05

    To measure income inequality with right censored (topcoded) data, we propose multiple imputation for censored observations using draws from Generalized Beta of the Second Kind distributions to provide partially synthetic datasets analyzed using complete data methods. Estimation and inference uses Reiter's (Survey Methodology 2003) formulae. Using Current Population Survey (CPS) internal data, we find few statistically significant differences in income inequality for pairs of years between 1995 and 2004. We also show that using CPS public use data with cell mean imputations may lead to incorrect inferences about inequality differences. Multiply-imputed public use data provide an intermediate solution.
    View Full Paper PDF
  • Working Paper

    Measuring Labor Earnings Inequality Using Public-Use March Current Population Survey Data: The Value of Including Variances and Cell Means When Imputing Topcoded Values

    November 2008

    Working Paper Number:

    CES-08-38

    Using the Census Bureau's internal March Current Population Surveys (CPS) file, we construct and make available variances and cell means for all topcoded income values in the publicuse version of these data. We then provide a procedure that allows researchers with access only to the public-use March CPS data to take advantage of this added information when imputing its topcoded income values. As an example of its value we show how our new procedure improves on existing imputation methods in the labor earnings inequality literature.
    View Full Paper PDF
  • Working Paper

    Health-Related Research Using Confidential U.S. Census Bureau Data

    August 2008

    Working Paper Number:

    CES-08-21

    Economic studies on health-related issues have the potential to benefit all Americans. The approaches for dealing with the growth of health care costs and health insurance coverage are ever changing and information is needed on their efficacy. Research on health-related topics has been conducted for about a decade at the Census Bureau\u2019s Center for Economic Studies and the Research Data Centers. This paper begins by describing the confidential business and demographic Census Bureau data products used in this research. The discussion continues with summaries of nearly 30 papers, including how this work has benefited the Census Bureau and its research findings. Some focus on data linkages and assessing data quality, while others address important questions in the employer, public, and individual insurance markets. This research could not have been accomplished with public-use data. The newly available data from the Agency for Healthcare Research and Quality and National Center for Health Statistics, as well as additional Census Bureau data now available in the Research Data Centers are also discussed.
    View Full Paper PDF
  • Working Paper

    Consistent Cell Means for Topcoded Incomes in the Public Use March CPS (1976-2007)

    March 2008

    Working Paper Number:

    CES-08-06

    Using the internal March CPS, we create and in this paper distribute to the larger research community a cell mean series that provides the mean of all income values above the topcode for any income source of any individual in the public use March CPS that has been topcoded since 1976. We also describe our construction of this series. When we use this series together with the public use March CPS, we closely match the yearly mean income levels and income inequalities of the U.S. population found using the internal March CPS data.
    View Full Paper PDF
  • Working Paper

    Access Methods for United States Microdata

    August 2007

    Working Paper Number:

    CES-07-25

    Beyond the traditional methods of tabulations and public-use microdata samples, statistical agencies have developed four key alternatives for providing non-government researchers with access to confidential microdata to improve statistical modeling. The first, licensing, allows qualified researchers access to confidential microdata at their own facilities, provided certain security requirements are met. The second, statistical data enclaves, offer qualified researchers restricted access to confidential economic and demographic data at specific agency-controlled locations. Third, statistical agencies can offer remote access, through a computer interface, to the confidential data under automated or manual controls. Fourth, synthetic data developed from the original data but retaining the correlations in the original data have the potential for allowing a wide range of analyses.
    View Full Paper PDF
  • Working Paper

    Using the P90/P10 Index to Measure U.S. Inequality Trends with Current Population Survey Data: A View From Inside the Census Bureau Vaults

    June 2007

    Working Paper Number:

    CES-07-17

    The March Current Population Survey (CPS) is the primary data source for estimation of levels and trends in labor earnings and income inequality in the USA. Time-inconsistency problems related to top coding in theses data have led many researchers to use the ratio of the 90th and 10th percentiles of these distributions (P90/P10) rather than a more traditional summary measure of inequality. With access to public use and restricted-access internal CPS data, and bounding methods, we show that using P90/P10 does not completely obviate time inconsistency problems, especially for household income inequality trends. Using internal data, we create consistent cell mean values for all top-coded public use values that, when used with public use data, closely track inequality trends in labor earnings and household income using internal data. But estimates of longer-term inequality trends with these corrected data based on P90/P10 differ from those based on the Gini coefficient. The choice of inequality measure matters.
    View Full Paper PDF
  • Working Paper

    Distribution Preserving Statistical Disclosure Limitation

    September 2006

    Working Paper Number:

    tp-2006-04

    One approach to limiting disclosure risk in public-use microdata is to release multiply-imputed, partially synthetic data sets. These are data on actual respondents, but with confidential data replaced by multiply-imputed synthetic values. A mis-specified imputation model can invalidate inferences because the distribution of synthetic data is completely determined by the model used to generate them. We present two practical methods of generating synthetic values when the imputer has only limited information about the true data generating process. One is applicable when the true likelihood is known up to a monotone transformation. The second requires only limited knowledge of the true likelihood, but nevertheless preserves the conditional distribution of the confidential data, up to sampling error, on arbitrary subdomains. Our method maximizes data utility and minimizes incremental disclosure risk up to posterior uncertainty in the imputation model and sampling error in the estimated transformation. We validate the approach with a simulation and application to a large linked employer-employee database.
    View Full Paper PDF
  • Working Paper

    The Impact of Hurricanes Katrina, Rita and Wilma on Business Establishments: A GIS Approach

    August 2006

    Working Paper Number:

    CES-06-23

    We use Geographic Information System tools to develop estimates of the economic impact of disaster events such as Hurricane Katrina. Our methodology relies on mapping establishments from the Census Bureau's Business Register into damage zones defined by remote sensing information provided by FEMA. The identification of damaged establishments by precisely locating them on a map provides a far more accurate characterization of affected businesses than those typically reported from readily available county level data. The need for prompt estimates is critical since they are more valuable the sooner they are released after a catastrophic event. Our methodology is based on pre-storm data. Therefore, estimates can be made available very quickly to inform the public as well as policy makers. Robustness tests using data from after the storms indicate our GIS estimates, while much smaller than those based on publicly available county-level data, still overstate actual observed losses. We discuss ways to refine and augment the GIS approach to provide even more accurate estimates of the impact of disasters on businesses.
    View Full Paper PDF
  • Working Paper

    Micro and Macro Data Integration: The Case of Capital

    May 2005

    Working Paper Number:

    CES-05-02

    Micro and macro data integration should be an objective of economic measurement as it is clearly advantageous to have internally consistent measurement at all levels of aggregation ' firm, industry and aggregate. In spite of the apparently compelling arguments, there are few measures of business activity that achieve anything close to micro/macro data internal consistency. The measures of business activity that are arguably the worst on this dimension are capital stocks and flows. In this paper, we document, quantify and analyze the widely different approaches to the measurement of capital from the aggregate (top down) and micro (bottom up) perspectives. We find that recent developments in data collection permit improved integration of the top down and bottom up approaches. We develop a prototype hybrid method that exploits these data to improve micro/macro data internal consistency in a manner that could potentially lead to substantially improved measures of capital stocks and flows at the industry level. We also explore the properties of the micro distribution of investment. In spite of substantial data and associated measurement limitations, we show that the micro distributions of investment exhibit properties that are of interest to both micro and macro analysts of investment behavior. These findings help highlight some of the potential benefits of micro/macro data integration.
    View Full Paper PDF
  • Working Paper

    New Approaches to Confidentiality Protection Synthetic Data, Remote Access and Research Data Centers

    June 2004

    Working Paper Number:

    tp-2004-03

    View Full Paper PDF