CREAT - Census Bureau

Multiple Classification Systems For Economic Data: Can A Thousand Flowers Bloom? And Should They?

December 1991

Written by: Robert H Mcguckin

Working Paper Number:

CES-91-08

Abstract

The principle that the statistical system should provide flexibility-- possibilities for generating multiple groupings of data to satisfy multiple objectives--if it is to satisfy users is universally accepted. Yet in practice, this goal has not been achieved. This paper discusses the feasibility of providing flexibility in the statistical system to accommodate multiple uses of the industrial data now primarily examined within the Standard Industrial Classification (SIC) system. In one sense, the question of feasibility is almost trivial. With today's computer technology, vast amounts of data can be manipulated and stored at very low cost. Reconfigurations of the basic data are very inexpensive compared to the cost of collecting the data. Flexibility in the statistical system implies more than the technical ability to regroup data. It requires that the basic data are sufficiently detailed to support user needs and are processed and maintained in a fashion that makes the use of a variety of aggregation rules possible. For this to happen, statistical agencies must recognize the need for high quality microdata and build this into their planning processes. Agencies need to view their missions from a multiple use perspective and move away from use of a primary reporting and collection vehicle. Although the categories used to report data must be flexible, practical considerations dictate that data collection proceed within a fixed classification system. It is simply too expensive for both respondents and statistical agencies to process survey responses in the absence of standardized forms, data entry programs, etc. I argue for a basic classification centered on commodities--products, services, raw materials and labor inputs--as the focus of data collection. The idea is to make the principle variables of interest--the commodities--the vehicle for the collection and processing of the data. For completeness, the basic classification should include labor usage through some form of occupational classification. In most economic surveys at the Census Bureau, the reporting unit and the classified unit have been the establishment. But there is no need for this to be so. The basic principle to be followed in data collection is that the data should be collected in the most efficient way--efficiency being defined jointly in terms of statistical agency collection costs and respondent burdens.

Document Tags and Keywords

Keywords:

data, aggregation, statistical, report, industrial, microdata, statistical agencies, aggregate, agency, classified, industrial classification, classification, classifying, reporting

Tags:

Census of Manufactures, Annual Survey of Manufactures, Standard Industrial Classification, Bureau of Labor Statistics, Center for Economic Studies, Statistics Canada

Similar Working Papers

The 10 most similar working papers to the working paper 'Multiple Classification Systems For Economic Data: Can A Thousand Flowers Bloom? And Should They?' are listed below in order of similarity.

Working Paper

The Importance of Establishment Data in Economic Research

August 1993

Authors: Robert H Mcguckin

Working Paper Number:

CES-93-10

The importance and usefulness of establishment microdata for economic research and policy analysis is outlined and contrasted with traditional products of statistical agencies -- aggregate cross-section tabulations. It is argued that statistical agencies must begin to seriously rethink the way they view establishment data products.
View Full Paper PDF
Working Paper

Analytic Use Of Economic Microdata; A Model For Researcher Access With Confidentiality Protection

August 1992

Authors: Robert H Mcguckin

Working Paper Number:

CES-92-08

A primary responsibility of the Center for Economic Studies (CES) of the U.S. Bureau of the Census is to facilitate researcher access to confidential economic microdata files. Benefits from this program accrue not only to policy makers--there is a growing awareness of the importance of microdata for analyzing both the descriptive and welfare implications of regulatory and environmental changes--but also and importantly to the statistical agencies themselves. In fact, there is substantial recent literature arguing for the proposition that the largest single improvement that the U.S. statistical system could make is to improve its analytic capabilities. In this paper I briefly discuss these benefits to greater access for analytical work and ways to achieve them. Due to the nature of business data, public use databases and masking technologies are not available as vehicles for releasing useful microdata files. I conclude that a combination of outside and inside research programs, carefully coordinated and integrated is the best model for ensuring that statistical agencies reap the gains from analytic data users. For the United States, at least, this is fortuitous with respect to justifying access since any direct research with confidential data by outsiders must have a "statistical purpose". Until the advent of CES, it was virtually impossible for researchers to work with the economic microdata collected by the various economic censuses. While the CES program is quite large, as it now stands, researchers, or their representatives, must come to the Census Bureau in Washington, D.C. to access the data. The success of the program has led to increasing demands for data access in facilities outside of the Washington, D.C. area. Two options are considered: 1) Establish Census Bureau facilities in various universities or similar nonprofit research facilities and 2) Develop CES regional operations in existing Census Bureau regional offices.
View Full Paper PDF
Working Paper

Longitudinal Economic Data At The Census Bureau: A New Database Yields Fresh Insight On Some Old Issues

January 1990

Authors: Robert H Mcguckin

Working Paper Number:

CES-90-01

This paper has two goals. First, it illustrates the importance of panel data with examples taken from research in progress using the U.S. Census Bureau's Longitudinal Research Database ( LRD ). Although the LRD is not the result of a "true" longitudinal survey, it provides both balanced and unbalanced panel data sets for establishments, firms, and lines of business. The second goal is to integrate the results of recent research with the LRD and to draw conclusions about the importance of longitudinal microdata for econometric research and time series analysis. The advantages of panel data arise from both the micro and time series aspects of the observations. This also leads us to consider why panel data are necessary to understand and interpret the time series behavior of aggregate statistics produced in cross-section establishment surveys and censuses. We find that typical homogeneity assumptions are likely to be inappropriate in a wide variety of applications. In particular, the industry in which an establishment is located, the ownership of the establishment, and the existence of the establishment (births and deaths) are endogenous variables that cannot simply be taken as time invariant fixed effects in econometric modeling.
View Full Paper PDF
Working Paper

Unlocking the Information in Integrated Social Data

May 2002

Authors: John M. Abowd

Working Paper Number:

tp-2002-21

View Full Paper PDF
Working Paper

Exploring New Ways to Classify Industries for Energy Analysis and Modeling

November 2022

Authors: Gale Boyd, Matthew Doolin, Liz Wachs, Colin McMillan

Working Paper Number:

CES-22-49

Combustion, other emitting processes and fossil energy use outside the power sector have become urgent concerns given the United States' commitment to achieving net-zero greenhouse gas emissions by 2050. Industry is an important end user of energy and relies on fossil fuels used directly for process heating and as feedstocks for a diverse range of applications. Fuel and energy use by industry is heterogeneous, meaning even a single product group can vary broadly in its production routes and associated energy use. In the United States, the North American Industry Classification System (NAICS) serves as the standard for statistical data collection and reporting. In turn, data based on NAICS are the foundation of most United States energy modeling. Thus, the effectiveness of NAICS at representing energy use is a limiting condition for current expansive planning to improve energy efficiency and alternatives to fossil fuels in industry. Facility-level data could be used to build more detail into heterogeneous sectors and thus supplement data from Bureau of the Census and U.S Energy Information Administration reporting at NAICS code levels but are scarce. This work explores alternative classification schemes for industry based on energy use characteristics and validates an approach to estimate facility-level energy use from publicly available greenhouse gas emissions data from the U.S. Environmental Protection Agency (EPA). The approaches in this study can facilitate understanding of current, as well as possible future, energy demand. First, current approaches to the construction of industrial taxonomies are summarized along with their usefulness for industrial energy modeling. Unsupervised machine learning techniques are then used to detect clusters in data reported from the U.S. Department of Energy's Industrial Assessment Center program. Clusters of Industrial Assessment Center data show similar levels of correlation between energy use and explanatory variables as three-digit NAICS codes. Interestingly, the clusters each include a large cross section of NAICS codes, which lends additional support to the idea that NAICS may not be particularly suited for correlation between energy use and the variables studied. Fewer clusters are needed for the same level of correlation as shown in NAICS codes. Initial assessment shows a reasonable level of separation using support vector machines with higher than 80% accuracy, so machine learning approaches may be promising for further analysis. The IAC data is focused on smaller and medium-sized facilities and is biased toward higher energy users for a given facility type. Cladistics, an approach for classification developed in biology, is adapted to energy and process characteristics of industries. Cladistics applied to industrial systems seeks to understand the progression of organizations and technology as a type of evolution, wherein traits are inherited from previous systems but evolve due to the emergence of inventions and variations and a selection process driven by adaptation to pressures and favorable outcomes. A cladogram is presented for evolutionary directions in the iron and steel sector. Cladograms are a promising tool for constructing scenarios and summarizing directions of sectoral innovation. The cladogram of iron and steel is based on the drivers of energy use in the sector. Phylogenetic inference is similar to machine learning approaches as it is based on a machine-led search of the solution space, therefore avoiding some of the subjectivity of other classification systems. Our prototype approach for constructing an industry cladogram is based on process characteristics according to the innovation framework derived from Schumpeter to capture evolution in a given sector. The resulting cladogram represents a snapshot in time based on detailed study of process characteristics. This work could be an important tool for the design of scenarios for more detailed modeling. Cladograms reveal groupings of emerging or dominant processes and their implications in a way that may be helpful for policymakers and entrepreneurs, allowing them to see the larger picture, other good ideas, or competitors. Constructing a cladogram could be a good first step to analysis of many industries (e.g. nitrogenous fertilizer production, ethyl alcohol manufacturing), to understand their heterogeneity, emerging trends, and coherent groupings of related innovations. Finally, validation is performed for facility-level energy estimates from the EPA Greenhouse Gas Reporting Program. Facility-level data availability continues to be a major challenge for industrial modeling. The method outlined by (McMillan et al. 2016; McMillan and Ruth 2019) allows estimating of facility level energy use based on mandatory greenhouse gas reporting. The validation provided here is an important step for further use of this data for industrial energy modeling.
View Full Paper PDF
Working Paper

Primary Versus Secondary Production Techniques in U.S. Manufacturing

October 1994

Authors: Joe Mattey, Thijs T Raa

Working Paper Number:

CES-94-12

In this paper we discuss and analyze a classical economic puzzle: whether differences in factor intensities reflect patterns of specialization or the co-existence of alternative techniques to produce output. We use observations on a large cross-section of U.S. manufacturing plants from the Census of Manufactures, including those that make goods primary to other industries, to study differences in production techniques. We find that in most cases material requirements do not depend on whether goods are made as primary products or as secondary products, which suggests that differences in factor intensities usually reflect patterns of specialization. A few cases where secondary production techniques do differ notably are discussed in more detail. However, overall the regression results support the neoclassical assumption that a single, best-practice technique is chosen for making each product.
View Full Paper PDF
Working Paper

Price Dispersion In U.S. Manufacturing: Implications For The Aggregation Of Products And Firms

March 1992

Authors: Thomas A Abbott Iii

Working Paper Number:

CES-92-03

This paper addresses the question of whether products in the U.S. Manufacturing sector sell at a single (common) price, or whether prices vary across producers. Price dispersion is interesting for at least two reasons. First, if output prices vary across producers, standard methods of using industry price deflators lead to errors in measuring real output at the industry, firm, and establishment level which may bias estimates of the production function and productivity growth. Second, price dispersion suggests product heterogeneity which, if consumers do not have identical preferences, could lead to market segmentation and price in excess of marginal cost, thus making the current (competitive) characterization of the Manufacturing sector inappropriate and invalidating many empirical studies. In the course of examining these issues, the paper develops a robust measure of price dispersion as well as new quantitative methods for testing whether observed price differences are the result of differences in product quality. Our results indicate that price dispersion is widespread throughout manufacturing and that for at least one industry, Hydraulic Cement, it is not the result of differences in product quality.
View Full Paper PDF
Working Paper

The Longitudinal Research Database (LRD): Status And Research Possibilities

July 1988

Authors: Robert H Mcguckin, George A Pascoe

Working Paper Number:

CES-88-02

This paper discusses the development and use of the Longitudinal Research Data available at the Center for Economic Studies of the Bureau of the Census in terms of what has been accomplished thus far, what projects are currently in progress, and what plans are in place for the near future. The major achievement to date is the construction of the database itself, which contains data for manufacturing establishments collected by the Census in 1963, 1967, 1972, 1977 and 1982, and the Annual Survey of Manufactures for non-Census years from 1973 to 1985. These data now reside in the Center's computer in a consistent format across all years. In addition, a large software development task that greatly simplifies the task of selecting subsets of the database for specific research projects is well underway. Finally, a number of powerful microcomputers have been purchased for use by researchers for their statistical analysis. Current efforts underway at the Center include research on such policy-relevant issues as mergers and their impact on profits and production, high technology trade, import competition, plant level productivity, entry and exit, and productivity differences between large and small firms. Due to the confidentiality requirements of the Census data, most of their research is performed by Center staff and Special Sworn Employees. Under certain circumstances, the Center accepts user-written programs from outside researchers. These routines are executed by Center staff, and the resultant output is reviewed thoroughly for disclosure problems. The Center is also an active member of a task force working on methods on release "masked" or "cloned" microdata in public-use files that will protect the confidentiality of the data while at the same time provide a research tool for outside users. The Center research program contributes directly to future research possibilities. The current batch of research projects is adding insight into the nature of the LRD database. This information is continually being incorporated into the Center's software system, thus facilitating yet more research activity. Moreover, since a good portion of the research involves linking the Longitudinal Research Data to other data files, such as the NSF/Census R&D data, the scope of the databases is continually being expanded. Furthermore, the Center is exploring the possibility of linking the demographic data collected by the Census Bureau to the LRD database.
View Full Paper PDF
Working Paper

Public Use Microdata: Disclosure And Usefulness

September 1988

Authors: Robert H Mcguckin, Sang V Nguyen

Working Paper Number:

CES-88-03

Official statistical agencies such as the Census Bureau and the Bureau of Labor Statistics collect enormous quantities of microdata in statistical surveys. These data are valuable for economic research and market and policy analysis. However, the data cannot be released to the public because of confidentiality commitments to individual respondents. These commitments, coupled with the strong research demand for microdata, have led the agencies to consider various proposals for releasing public use microdata. Most proposals for public use microdata call for the development of surrogate data that disguise the original data. Thus, they involve the addition of measurement errors to the data. In this paper, we examine disclosure issues and explore alternative masking methods for generating panels of useful economic microdata that can be released to researchers. While our analysis applies to all confidential microdata, applications using the Census Bureau's Longitudinal Research Data Base (LRD) are used for illustrative purposes throughout the discussion.
View Full Paper PDF
Working Paper

Measuring the Electronic Economy: Current Status and Next Steps

June 2000

Authors: Ron Jarmin, Barbara K Atrostic, John Gates

Working Paper Number:

CES-00-10

The recent growth of consumer retailing over the Internet draws attention to the electronic economy. However, businesses also conduct other business processes over computer networks, and many have been doing so for some time. Uses of computer networks attract attention because of assertions that they lead to new products and services, new delivery methods, streamlined or re-engineered business processes, new business structures, and enhanced business performance. These changes, in turn, potentially affect the performance of the entire economy, including economic growth, productivity, prices, employment, trade, and the structures of businesses, regions, and markets. Evaluating these assertions, and their effects on economic performance, requires solid statistical information about the electronic economy. This paper develops principles for identifying information critical to measuring the size and evaluating the potential effects of the electronic economy, relates that information to current data collection programs, and notes relevant measurement issues. Some of the required information about the electronic economy can be collected by adding questions to existing surveys, making the scope of existing surveys consistent, or developing new surveys. However, many key pieces of information pose significant challenges to economic measurement. While some of those challenges are specific to the electronic economy, others are long-standing ones. Interest in the electronic economy highlights the importance of continuing attempts to address these challenges. Improving and enhancing the statistical system to provide information about the electronic economy, therefore, would also substantially improve the baseline information available for evaluating the performance of the entire economy.
View Full Paper PDF

Multiple Classification Systems For Economic Data: Can A Thousand Flowers Bloom? And Should They?

December 1991

Working Paper Number:

CES-91-08

Abstract

Document Tags and Keywords

The 10 most similar working papers to the working paper 'Multiple Classification Systems For Economic Data: Can A Thousand Flowers Bloom? And Should They?' are listed below in order of similarity.

August 1993

Working Paper Number:

CES-93-10

August 1992

Working Paper Number:

CES-92-08

January 1990

Working Paper Number:

CES-90-01

May 2002

Working Paper Number:

tp-2002-21

November 2022

Working Paper Number:

CES-22-49

October 1994

Working Paper Number:

CES-94-12

March 1992

Working Paper Number:

CES-92-03

July 1988

Working Paper Number:

CES-88-02

September 1988

Working Paper Number:

CES-88-03

June 2000

Working Paper Number:

CES-00-10