In the U.S. Census of Manufactures, the Census Bureau imputes missing values using a combination of mean imputation, ratio imputation, and conditional mean imputation. It is wellknown that imputations based on these methods can result in underestimation of variability and potential bias in multivariate inferences. We show that this appears to be the case for the existing imputations in the Census of Manufactures. We then present an alternative strategy for handling the missing data based on multiple imputation. Specifically, we impute missing values via sequences of classification and regression trees, which offer a computationally straightforward and flexible approach for semi-automatic, large-scale multiple imputation. We also present an approach to evaluating these imputations based on posterior predictive checks. We use the multiple imputations, and the imputations currently employed by the Census Bureau, to estimate production function parameters and productivity dispersions. The results suggest that the two approaches provide quite different answers about productivity.
-
RECOVERING THE ITEM-LEVEL EDIT AND IMPUTATION FLAGS IN THE 1977-1997 CENSUSES OF MANUFACTURES
September 2014
Working Paper Number:
CES-14-37
As part of processing the Census of Manufactures, the Census Bureau edits some data items and imputes for missing data and some data that is deemed erroneous. Until recently it was difficult for researchers using the plant-level microdata to determine which data items were changed or imputed during the editing and imputation process, because the edit/imputation processing flags were not available to researchers. This paper describes the process of reconstructing the edit/imputation flags for variables in the 1977, 1982, 1987, 1992, and 1997 Censuses of Manufactures using recently recovered Census Bureau files. Thepaper also reports summary statistics for the percentage of cases that are imputed for key variables. Excluding plants with fewer than 5 employees, imputation rates for several key variables range from 8% to 54% for the manufacturing sector as a whole, and from 1% to 72% at the 2-digit SIC industry level.
View Full
Paper PDF
-
Manufacturing Dispersion: How Data Cleaning Choices Affect Measured Misallocation and Productivity Growth in the Annual Survey of Manufactures
September 2025
Working Paper Number:
CES-25-67
Measurement of dispersion of productivity levels and productivity growth rates across businesses is a key input for answering a variety of important economic questions, such as understanding the allocation of economic inputs across businesses and over time. While item nonresponse is a readily quantifiable issue, we show there is also misreporting by respondents in the Annual Survey of Manufactures (ASM). Aware of these measurement issues, the Census Bureau edits and imputes survey responses before tabulation and dissemination. However, edit and imputation methods that are suitable for publishing aggregate totals may not be suitable for estimating other measures from the microdata. We show that the methods used dramatically affect estimates of productivity dispersion, allocative efficiency, and aggregate productivity growth. Using a Bayesian approach for editing and imputation, we model the joint distributions of all variables needed to estimate these measures, and we quantify the degree of uncertainty in the estimates due to imputations for faulty or missing data.
View Full
Paper PDF
-
USING IMPUTATION TECHNIQUES TO EVALUATE STOPPING RULES IN ADAPTIVE SURVEY DESIGN
October 2014
Working Paper Number:
CES-14-40
Adaptive Design methods for social surveys utilize the information from the data as it is collected to make decisions about the sampling design. In some cases, the decision is either to continue or stop the data collection. We evaluate this decision by proposing measures to compare the collected data with follow-up samples. The options are assessed by imputation of the nonrespondents under different missingness scenarios, including Missing Not at Random. The variation in the utility measures is compared to the cost induced by the follow-up sample sizes. We apply the proposed method to the 2007 U.S. Census of Manufacturers.
View Full
Paper PDF
-
Regulating Mismeasured Pollution: Implications of Firm Heterogeneity for Environmental Policy
August 2018
Working Paper Number:
CES-18-03R
This paper provides the first estimates of within-industry heterogeneity in energy and CO2 productivity for the entire U.S. manufacturing sector. We measure energy and CO2 productivity as output per dollar energy input or per ton CO2 emitted. Three findings emerge. First, within narrowly defined industries, heterogeneity in energy and CO2 productivity across plants is enormous. Second, heterogeneity in energy and CO2 productivity exceeds heterogeneity in most other productivity measures, like labor or total factor productivity. Third, heterogeneity in energy and CO2 productivity has important implications for environmental policies targeting industries rather than plants, including technology standards and carbon border adjustments.
View Full
Paper PDF
-
Simultaneous Edit-Imputation for Continuous Microdata
December 2015
Working Paper Number:
CES-15-44
Many statistical organizations collect data that are expected to satisfy linear constraints; as examples, component variables should sum to total variables, and ratios of pairs of variables should be bounded by expert-specified constants. When reported data violate constraints, organizations identify and replace values potentially in error in a process known as edit-imputation. To date, most approaches separate the error localization and imputation steps, typically using optimization methods to identify the variables to change followed by hot deck imputation. We present an approach that fully integrates editing and imputation for continuous microdata under linear constraints. Our approach relies on a Bayesian hierarchical model that includes (i) a flexible joint probability model for the underlying true values of the data with support only on the set of values that satisfy all editing constraints, (ii) a model for latent indicators of the variables that are in error, and (iii) a model for the reported responses for variables in error. We illustrate the potential advantages of the Bayesian editing approach over existing approaches using simulation studies. We apply the model to edit faulty data from the 2007 U.S. Census of Manufactures. Supplementary materials for this article are available online.
View Full
Paper PDF
-
File Matching with Faulty Continuous Matching Variables
January 2017
Working Paper Number:
CES-17-45
We present LFCMV, a Bayesian file linking methodology designed to link records using continuous matching variables in situations where we do not expect values of these matching variables to agree exactly across matched pairs. The method involves a linking model for the distance between the matching variables of records in one file and the matching variables of their linked records in the second. This linking model is conditional on a vector indicating the links. We specify a mixture model for the distance component of the linking model, as this latent structure allows the distance between matching variables in linked pairs to vary across types of linked pairs. Finally, we specify a model for the linking vector. We describe the Gibbs sampling algorithm for sampling from the posterior distribution of this linkage model and use artificial data to illustrate model performance. We also introduce a linking application using public survey information and data from the U.S. Census of Manufactures and use
LFCMV to link the records.
View Full
Paper PDF
-
Materials Prices and Productivity
June 2012
Working Paper Number:
CES-12-11
There is substantial within-industry variation, even within industries that use and produce homogeneous inputs and outputs, in the prices that plants pay for their material inputs. I explore, using plant-level data from the U.S. Census Bureau, the consequences and sources of this variation in materials prices. For a sample of industries with relatively homogeneous products, the standard deviation of plant-level productivities would be 7% lower if all plants faced the same materials prices. Moreover, plant-level materials prices are both persistent across time and predictive of exit. The contribution of net entry to aggregate productivity growth is smaller for productivity measures that strip out di'erences in materials prices. After documenting these patterns, I discuss three potential sources of materials price variation: geography, di'erences in suppliers. marginal costs, and suppliers. price discriminatory behavior. Together, these variables account for 13% of the dispersion of materials prices. Finally, I demonstrate that plants.marginal costs are correlated with the marginal costs of their intermediate input suppliers.
View Full
Paper PDF
-
Multiply-Imputing Confidential Characteristics and File Links in Longitudinal Linked Data
June 2004
Working Paper Number:
tp-2004-04
This paper describes ongoing research to protect confidentiality in longitudinal linked
data through creation of multiply-imputed, partially synthetic data. We present two enhancements to the methods
of [2]. The first is designed to preserve marginal distributions in the partially synthetic data. The second is
designed to protect confidential links between sampling frames.
View Full
Paper PDF
-
Empirical Distribution of the Plant-Level Components of Energy and Carbon Intensity at the Six-digit NAICS Level Using a Modified KAYA Identity
September 2024
Working Paper Number:
CES-24-46
Three basic pillars of industry-level decarbonization are energy efficiency, decarbonization of energy sources, and electrification. This paper provides estimates of a decomposition of these three components of carbon emissions by industry: energy intensity, carbon intensity of energy, and energy (fuel) mix. These estimates are constructed at the six-digit NAICS level from non-public, plant-level data collected by the Census Bureau. Four quintiles of the distribution of each of the three components are constructed, using multiple imputation (MI) to deal with non-reported energy variables in the Census data. MI allows the estimates to avoid non-reporting bias. MI also allows more six-digit NAICS to be estimated under Census non-disclosure rules, since dropping non-reported observations may have reduced the sample sizes unnecessarily. The estimates show wide variation in each of these three components of emissions (intensity) and provide a first empirical look into the plant-level variation that underlies carbon emissions.
View Full
Paper PDF
-
Collaborative Micro-productivity Project: Establishment-Level Productivity Dataset, 1972-2020
December 2023
Working Paper Number:
CES-23-65
We describe the process for building the Collaborative Micro-productivity Project (CMP) microdata and calculating establishment-level productivity numbers. The documentation is for version 7 and the data cover the years 1972-2020. These data have been used in numerous research papers and are used to create the experimental public-use data product Dispersion Statistics on Productivity (DiSP).
View Full
Paper PDF