Characterizing the work that people do on their jobs is a longstanding and core issue in labor economics. Traditionally, classification has been done manually. If it were possible to combine new computational tools and administrative wage records to generate an automated crosswalk between job titles and occupations, millions of dollars could be saved in labor costs, data processing could be sped up, data could become more consistent, and it might be possible to generate, without a lag, current information about the changing occupational composition of the labor market. This paper examines the potential to assign occupations to job titles contained in administrative data using automated, machine-learning approaches. We use a new extraordinarily rich and detailed set of data on transactional HR records of large firms (universities) in a relatively narrowly defined industry (public institutions of higher education) to identify the potential for machine-learning approaches to classify occupations.
-
A Task-based Approach to Constructing Occupational Categories
with Implications for Empirical Research in Labor Economics
September 2019
Working Paper Number:
CES-19-27
Most applied research in labor economics that examines returns to worker skills or differences in earnings across subgroups of workers typically accounts for the role of occupations by controlling for occupational categories. Researchers often aggregate detailed occupations into categories based on the Standard Occupation Classification (SOC) coding scheme, which is based largely on narratives or qualitative measures of workers' tasks. Alternatively, we propose two quantitative task-based approaches to constructing occupational categories by using factor analysis with O*NET job descriptors that provide a rich set of continuous measures of job tasks across all occupations. We find that our task-based approach outperforms the SOC-based approach in terms of lower occupation distance measures. We show that our task-based approach provides an intuitive, nuanced interpretation for grouping occupations and permits quantitative assessments of similarities in task compositions across occupations. We also replicate a recent analysis and find that our task-based occupational categories explain more of the gender wage gap than the SOC-based approaches explain. Our study enhances the Federal Statistical System's understanding of the SOC codes, investigates ways to use third-party data to construct useful research variables that can potentially be added to Census Bureau data products to improve their quality and versatility, and sheds light on how the use of alternative occupational categories in economics research may lead to different empirical results and deeper understanding in the analysis of labor market outcomes.
View Full
Paper PDF
-
NOISE INFUSION AS A CONFIDENTIALITY PROTECTION MEASURE FOR GRAPH-BASED STATISTICS
September 2014
Working Paper Number:
CES-14-30
We use the bipartite graph representation of longitudinally linked em-ployer-employee data, and the associated projections onto the employer and em-ployee nodes, respectively, to characterize the set of potential statistical summar-ies that the trusted custodian might produce. We consider noise infusion as the primary confidentiality protection method. We show that a relatively straightfor-ward extension of the dynamic noise-infusion method used in the U.S. Census Bureau's Quarterly Workforce Indicators can be adapted to provide the same confidentiality guarantees for the graph-based statistics: all inputs have been modified by a minimum percentage deviation (i.e., no actual respondent data are used) and, as the number of entities contributing to a particular statistic increases, the accuracy of that statistic approaches the unprotected value. Our method also ensures that the protected statistics will be identical in all releases based on the same inputs.
View Full
Paper PDF
-
Squeezing More Out of Your Data: Business Record Linkage with Python
November 2018
Working Paper Number:
CES-18-46
Integrating data from different sources has become a fundamental component of modern data analytics. Record linkage methods represent an important class of tools for accomplishing such integration. In the absence of common disambiguated identifiers, researchers often must resort to ''fuzzy" matching, which allows imprecision in the characteristics used to identify common entities across dfferent datasets. While the record linkage literature has identified numerous individually useful fuzzy matching techniques, there is little consensus on a way to integrate those techniques within a
single framework. To this end, we introduce the Multiple Algorithm Matching for Better Analytics (MAMBA), an easy-to-use, flexible, scalable, and transparent software platform for business record linkage applications using Census microdata. MAMBA leverages multiple string comparators to assess the similarity of records using a machine learning algorithm to disambiguate matches. This software represents a transparent tool for researchers seeking to link external business data to the Census Business Register files.
View Full
Paper PDF
-
A Tale of Two Fields? STEM Career Outcomes
October 2023
Working Paper Number:
CES-23-53
Is the labor market for US researchers experiencing the best or worst of times? This paper analyzes the market for recently minted Ph.D. recipients using supply-and-demand logic and data linking graduate students to their dissertations and W2 tax records. We also construct a new dissertation-industry 'relevance' measure, comparing dissertation and patent text and linking patents to assignee firms and industries. We find large disparities across research fields in placement (faculty, postdoc, and industry positions), earnings, and the use of specialized human capital. Thus, it appears to simultaneously be a good time for some fields and a bad time for others.
View Full
Paper PDF
-
Job Tasks, Worker Skills, and Productivity
September 2025
Authors:
John Haltiwanger,
Lucia Foster,
Cheryl Grim,
Zoltan Wolf,
Cindy Cunningham,
Sabrina Wulff Pabilonia,
Jay Stewart,
Cody Tuttle,
G. Jacob Blackwood,
Matthew Dey,
Rachel Nesbit
Working Paper Number:
CES-25-63
We present new empirical evidence suggesting that we can better understand productivity dispersion across businesses by accounting for differences in how tasks, skills, and occupations are organized. This aligns with growing attention to the task content of production. We link establishment-level data from the Bureau of Labor Statistics Occupational Employment and Wage Statistics survey with productivity data from the Census Bureau's manufacturing surveys. Our analysis reveals strong relationships between establishment productivity and task, skill, and occupation inputs. These relationships are highly nonlinear and vary by industry. When we account for these patterns, we can explain a substantial share of productivity dispersion across establishments.
View Full
Paper PDF
-
Person Matching in Historical Files using the Census Bureau's Person Validation System
September 2014
Working Paper Number:
carra-2014-11
The recent release of the 1940 Census manuscripts enables the creation of longitudinal data spanning the whole of the twentieth century. Linked historical and contemporary data would allow unprecedented analyses of the causes and consequences of health, demographic, and economic change. The Census Bureau is uniquely equipped to provide high quality linkages of person records across datasets. This paper summarizes the linkage techniques employed by the Census Bureau and discusses utilization of these techniques to append protected identification keys to the 1940 Census.
View Full
Paper PDF
-
An Evaluation of the Gender Wage Gap Using Linked Survey and Administrative Data
November 2020
Working Paper Number:
CES-20-34
The narrowing of the gender wage gap has slowed in recent decades. However, current estimates show that, among full-time year-round workers, women earn approximately 18 to 20 percent less than men at the median. Women's human capital and labor force characteristics that drive wages increasingly resemble men's, so remaining differences in these characteristics explain less of the gender wage gap now than in the past. As these factors wane in importance, studies show that others like occupational and industrial segregation explain larger portions of the gender wage gap. However, a major limitation of these studies is that the large datasets required to analyze occupation and industry effectively lack measures of labor force experience. This study combines survey and administrative data to analyze and improve estimates of the gender wage gap within detailed occupations, while also accounting for gender differences in work experience. We find a gender wage gap of 18 percent among full-time, year-round workers across 316 detailed occupation categories. We show the wage gap varies significantly by occupation: while wages are at parity in some occupations, gaps are as large as 45 percent in others. More competitive and hazardous occupations, occupations that reward longer hours of work, and those that have a larger proportion of women workers have larger gender wage gaps. The models explain less of the wage gap in occupations with these attributes. Occupational characteristics shape the conditions under which men and women work and we show these characteristics can make for environments that are more or less conducive to gender parity in earnings.
View Full
Paper PDF
-
The Gender Pay Gap and Its Determinants Across the Human Capital Distribution
June 2023
Working Paper Number:
CES-23-31R
This paper links American Community Survey data and postsecondary transcript records to examine how the gender pay gap varies across the distribution of education credentials for a sample of 2003-2013 graduates. Although recent literature emphasizes gender inequality among the most-educated, we find a smaller gender pay gap at higher education levels. Field-of-degree and occupation effects explain most of the gap among top bachelor's graduates, while work hours and unobserved channels matter more for less-competitive bachelor's, associate, and certificate graduates. We develop a novel decomposition of the child penalty to examine the role of children in explaining these results.
View Full
Paper PDF
-
The impact of manufacturing credentials on earnings and the probability of employment
May 2022
Working Paper Number:
CES-22-15
This paper examines the labor market returns to earning industry-certified credentials in the manufacturing sector. Specifically, we are interested in estimating the impact of a manufacturing credential on wages, probability of employment, and probability of employment specifically in the manufacturing sector post credential attainment. We link students who earned manufacturing credentials to their enrollment and completion records, and then further link them to their IRS tax records for earnings and employment (Form W2 and 1040) and to the American Community Survey and decennial census for demographic information. We present earnings trajectories for workers with credentials by type of credential, industry of employment, age, race and ethnicity, gender, and state. To obtain a more causal estimate of the impact of a credential on earnings, we implement a coarsened exact matching strategy to compare outcomes between otherwise similar people with and without a manufacturing credential. We find that the attainment of a manufacturing industry credential is associated with higher earnings and a higher likelihood of labor market participation when we compare attainers to a group of non-attainers who are otherwise similar.
View Full
Paper PDF
-
Business Dynamics of Innovating Firms: Linking U.S. Patents with Administrative Data on Workers and Firms
July 2015
Working Paper Number:
CES-15-19
This paper discusses the construction of a new longitudinal database tracking inventors and patent-owning firms over time. We match granted patents between 2000 and 2011 to administrative databases of firms and workers housed at the U.S. Census Bureau. We use inventor information in addition to the patent assignee firm name to and improve on previous efforts linking patents to firms. The triangulated database allows us to maximize match rates and provide validation for a large fraction of matches. In this paper, we describe the construction of the database and explore basic features of the data. We find patenting firms, particularly young patenting firms, disproportionally contribute jobs to the U.S. economy. We find patenting is a relatively rare event among small firms but that most patenting firms are nevertheless small, and that patenting is not as rare an event for the youngest firms compared to the oldest firms. While manufacturing firms are more likely to patent than firms in other sectors, we find most patenting firms are in the services and wholesale sectors. These new data are a product of collaboration within the U.S. Department of Commerce, between the U.S. Census Bureau and the U.S. Patent and Trademark Office.
View Full
Paper PDF