Big data offers potentially enormous benefits for improving economic measurement, but it also presents challenges (e.g., lack of representativeness and instability), implying that their value is not always clear. We propose a framework for quantifying the usefulness of these data sources for specific applications, relative to existing official sources. We specifically weigh the potential benefits of additional granularity and timeliness, while examining the accuracy associated with any new or improved estimates, relative to comparable accuracy produced in existing official statistics. We apply the methodology to employment estimates using data from a payroll processor, considering both the improvement of existing state-level estimates, but also the production of new, more timely, county-level estimates. We find that incorporating payroll data can improve existing state-level estimates by 11% based on out-of-sample mean absolute error, although the improvement is considerably higher for smaller state-industry cells. We also produce new county-level estimates that could provide more timely granular estimates than previously available. We develop a novel test to determine if these new county-level estimates have errors consistent with official series. Given the level of granularity, we cannot reject the hypothesis that the new county estimates have an accuracy in line with official measures, implying an expansion of the existing frontier. We demonstrate the practical importance of these experimental estimates by investigating a hypothetical application during the COVID-19 pandemic, a period in which more timely and granular information could have assisted in implementing effective policies. Relative to existing estimates, we find that the alternative payroll data series could help identify areas of the country where employment was lagging. Moreover, we also demonstrate the value of a more timely series.
-
Business Applications as a Leading Economic Indicator?
May 2021
Working Paper Number:
CES-21-09R
How are applications to start new businesses related to aggregate economic activity? This paper explores the properties of three monthly business application series from the U.S. Census Bureau's Business Formation Statistics as economic indicators: all business applications, business applications that are relatively likely to turn into new employer businesses ('likely employers'), and the residual series -- business applications that have a relatively low rate of becoming employers ('likely non-employers'). Growth in applications for likely employers significantly leads total nonfarm employment growth and has a strong positive correlation with it. Furthermore, growth in applications for likely employers leads growth in most of the monthly Principal Federal Economic Indicators (PFEIs). Motivated by our findings, we estimate a dynamic factor model (DFM) to forecast nonfarm employment growth over a 12-month period using the PFEIs and the likely employers series. The latter improves the model's forecast, especially in the years following the turning points of the Great Recession and the COVID-19 pandemic. Overall, applications for likely employers are a strong leading indicator of monthly PFEIs and aggregate economic activity, whereas applications for likely non-employers provide early information about changes in increasingly prevalent self-employment activity in the U.S. economy.
View Full
Paper PDF
-
Building the Census Bureau Index of Economic Activity (IDEA)
March 2023
Working Paper Number:
CES-23-15
The Census Bureau Index of Economic Activity (IDEA) is constructed from 15 of the Census Bureau's primary monthly economic time series. The index is intended to provide a single time series reflecting, to the extent possible, the variation over time in the whole set of component series. The component series provide monthly measures of activity in retail and wholesale trade, manufacturing, construction, international trade, and business formations. Most of the input series are Principal Federal Economic Indicators. The index is constructed by applying the method of principal components analysis (PCA) to the time series of monthly growth rates of the seasonally adjusted component series, after standardizing the growth rates to series with mean zero and variance 1. Similar PCA approaches have been used for the construction of other economic indices, including the Chicago Fed National Activity Index issued by the Federal Reserve Bank of Chicago, and the Weekly Economic Index issued by the Federal Reserve Bank of New York. While the IDEA is constructed from time series of monthly data, it is calculated and published every business day, and so is updated whenever a new monthly value is released for any of its component series. Since release dates of data values for a given month vary across the component series, with slight variations in the monthly release date for any one component series, updates to the index are frequent. It is unavoidably the case that, at almost all updates, some of the component series lack observations for the current (most recent) data month. To address this situation, component series that are one month behind are predicted (nowcast) for the current index month, using a multivariate autoregressive time series model. This report discusses the input series to the index, the construction of the index by PCA, and the nowcasting procedure used. The report then examines some properties of the index and its relation to quarterly U.S. Gross Domestic Product and to some monthly non-Census Bureau economic indicators.
View Full
Paper PDF
-
High-Growth Firms in the United States: Key Trends and New Data Opportunities
March 2024
Working Paper Number:
CES-24-11
Using administrative data from the U.S. Census Bureau, we introduce a new public-use database that tracks activities across firm growth distributions over time and by firm and establishment characteristics. With these new data, we uncover several key trends on high-growth firms'critical engines of innovation and economic growth. First, the share of firms that are high-growth has steadily decreased over the past four decades, driven not only by falling firm entry rates but also languishing growth among existing firms. Second, this decline is particularly pronounced among young and small firms, while the share of high-growth firms has been relatively stable among large and old firms. Third, the decline in high-growth firms is found in all sectors, but the information sector has shown a modest rebound beginning in 2010. Fourth, there is significant variation in high-growth firm activity across states, with California, Texas, and Florida having high shares of high-growth firms. We highlight several areas for future research enabled by these new data.
View Full
Paper PDF
-
JOB-TO-JOB (J2J) Flows: New Labor Market Statistics From Linked Employer-Employee Data
September 2014
Working Paper Number:
CES-14-34
Flows of workers across jobs are a principal mechanism by which labor markets allocate workers to optimize productivity. While these job flows are both large and economically important, they represent a significant gap in available economic statistics. A soon to be released data product from the U.S. Census Bureau will fill this gap. The Job-to-Job (J2J) flow statistics provide estimates of worker flows across jobs, across different geographic labor markets, by worker and firm characteristics, including direct job-to-job flows as well as job changes with intervening nonemployment. In this paper, we describe the creation of the public-use data product on job-to-job flows. The data underlying the statistics are the matched employer-employee data from the U.S. Census Bureau's Longitudinal Employer-Household Dynamics program. We describe definitional issues and the identification strategy for tracing worker movements between employers in administrative data. We then compare our data with related series and discuss similarities and differences. Lastly, we describe disclosure avoidance techniques for the public use file, and our methodology for estimating national statistics when there is partially missing geography.
View Full
Paper PDF
-
Estimating A Multivariate Arma Model with Mixed-Frequency Data: An Application to Forecasting U.S. GNP at Monthly Intervals
July 1990
Working Paper Number:
CES-90-05
This paper develops and applies a method for directly estimating a multivariate, autoregressive moving-average (ARMA) model with mixed-frequency, time-series data. Unlike standard, single-frequency methods, the method does not require the data to be transformed to a single frequency (by temporally aggregating higher-frequency data to lower frequencies for interpolating lower-frequency data to higher frequencies) or the model to be restricted by frequency. Subject to computational constraints, the method can handle any number of variable and frequencies. In addition, variable can be treated as temporally aggregated and observed with errors and delays. The key to the method is to view lower-frequency data as periodically missing and to use the missing-data variant of the Kalman filter.
In the application, a bivariate, ARMA model is estimated with monthly observations on total employment and quarterly observations on real GNP, in the U.S., for January 1958 to December 1978. The estimated model is, then, used to compute monthly forecasts of the variables for 1 to 12 months ahead, for January 1979 to December 1988. Compared with GNP forecasts, in particular, for similar periods produced by established econometric and time series models, present GNP forecasts are generally more accurate for 1 to 4 months ahead and about equally or slightly less accurate for 5 to 12 months ahead. The application, thus, shows that the present method is tractable and able to effectively exploit cross-frequency sample information, in ARMA estimate and forecasting, which standard methods cannot exploit at all.
View Full
Paper PDF
-
Public-Use vs. Restricted-Use:
An Analysis Using the American Community Survey
January 2017
Working Paper Number:
CES-17-12
Statistical agencies frequently publish microdata that have been altered to protect confidentiality. Such data retain utility for many types of broad analyses but can yield biased or Insufficiently precise results in others. Research access to de-identified versions of the restricted-use data with little or no alteration is often possible, albeit costly and time-consuming. We investigate the the advantages and disadvantages of public-use and restricted-use data from the American Community
Survey (ACS) in constructing a wage index. The public-use data used were Public Use Microdata Samples, while the restricted-use data were accessed via a Federal Statistical Research Data Center. We discuss the advantages and disadvantages of each data source and compare estimated CWIs and standard errors at the state and labor market levels.
View Full
Paper PDF
-
Incorporating Administrative Data in Survey Weights for the Basic Monthly Current Population Survey
January 2024
Working Paper Number:
CES-24-02
Response rates to the Current Population Survey (CPS) have declined over time, raising the potential for nonresponse bias in key population statistics. A potential solution is to leverage administrative data from government agencies and third-party data providers when constructing survey weights. In this paper, we take two approaches. First, we use administrative data to build a non-parametric nonresponse adjustment step while leaving the calibration to population estimates unchanged. Second, we use administratively linked data in the calibration process, matching income data from the Internal Return Service and state agencies, demographic data from the Social Security Administration and the decennial census, and industry data from the Census Bureau's Business Register to both responding and nonresponding households. We use the matched data in the household nonresponse adjustment of the CPS weighting algorithm, which changes the weights of respondents to account for differential nonresponse rates among subpopulations.
After running the experimental weighting algorithm, we compare estimates of the unemployment rate and labor force participation rate between the experimental weights and the production weights. Before March 2020, estimates of the labor force participation rates using the experimental weights are 0.2 percentage points higher than the original estimates, with minimal effect on unemployment rate. After March 2020, the new labor force participation rates are similar, but the unemployment rate is about 0.2 percentage points higher in some months during the height of COVID-related interviewing restrictions. These results are suggestive that if there is any nonresponse bias present in the CPS, the magnitude is comparable to the typical margin of error of the unemployment rate estimate. Additionally, the results are overall similar across demographic groups and states, as well as using alternative weighting methodology. Finally, we discuss how our estimates compare to those from earlier papers that calculate estimates of bias in key CPS labor force statistics.
This paper is for research purposes only. No changes to production are being implemented at this time.
View Full
Paper PDF
-
High Frequency Business Dynamics in the United States During the COVID-19 Pandemic
March 2021
Working Paper Number:
CES-21-06
Existing small businesses experienced very sharp declines in activity, business sentiment, and expectations early in the pandemic. While there has been some recovery since the early days of the pandemic, small businesses continued to exhibit indicators of negative growth, business sentiment, and expectations through the first week of January 2021. These findings are from a unique high frequency, real time survey of small employer businesses, the Census Bureau's Small Business Pulse Survey (SBPS). Findings from the SBPS show substantial variation across sectors in the outcomes for small businesses. Small businesses in Accommodation and Food Services have been hit especially hard relative to those Finance and Insurance. However, even in Finance and Insurance small businesses exhibit indicators of negative growth, business sentiment, and expectations for all weeks from late April 2020 through the first week of 2021. While existing small businesses have fared poorly, after an initial decline, there has been a surge in new business applications based on the high frequency, real time Business Formation Statistics (BFS). Most of these applications are for likely nonemployers that are out of scope for the SBPS. However, there has also been a surge in new applications for likely employers. The surge in applications has been especially apparent in Retail Trade (and especially Non-store Retailers). We compare and contrast the patterns from these two new high frequency data products that provide novel insights into the distinct patterns of dynamics for existing small businesses relative to new business formations.
View Full
Paper PDF
-
NOISE INFUSION AS A CONFIDENTIALITY PROTECTION MEASURE FOR GRAPH-BASED STATISTICS
September 2014
Working Paper Number:
CES-14-30
We use the bipartite graph representation of longitudinally linked em-ployer-employee data, and the associated projections onto the employer and em-ployee nodes, respectively, to characterize the set of potential statistical summar-ies that the trusted custodian might produce. We consider noise infusion as the primary confidentiality protection method. We show that a relatively straightfor-ward extension of the dynamic noise-infusion method used in the U.S. Census Bureau's Quarterly Workforce Indicators can be adapted to provide the same confidentiality guarantees for the graph-based statistics: all inputs have been modified by a minimum percentage deviation (i.e., no actual respondent data are used) and, as the number of entities contributing to a particular statistic increases, the accuracy of that statistic approaches the unprotected value. Our method also ensures that the protected statistics will be identical in all releases based on the same inputs.
View Full
Paper PDF
-
An Analysis of Key Differences in Micro Data: Results from the Business List Comparison Project
September 2008
Working Paper Number:
CES-08-28
The Bureau of Labor Statistics and the Bureau of the Census each maintain a business register, a universe of all U.S. business establishments and their characteristics, created from independent sources. Both registers serve critical functions such as supplying aggregate data inputs for certain national statistics generated by the Bureau of Economic Analysis. This paper examines key micro-level differences across these two business registers.
View Full
Paper PDF