CREAT: Census Research Exploration and Analysis Tool

Integrating Multiple U.S. Census Bureau Data Assets to Create Standardized Profiles of Program Participants

January 2026

Abstract

The Foundations for Evidence-Based Policymaking Act of 2018 (Evidence Act) directed federal agencies to systematically use data when making policy decisions. In response, the U.S. Census Bureau established the Evidence Group within its Center for Economic Studies (CES). With an interdisciplinary team of economists, sociologists, and statisticians, the Evidence Group can support the broader federal government in their efforts to use existing data to improve program operations without increasing respondent burden. For federal agencies administering social safety net and business assistance programs in particular, the team provides a no-cost evidence-building service that links program records to Census Bureau data assets and creates a series of standardized tables describing participants, their economic outcomes prior to program entry, and the communities where they live. These tables provide partner agencies with the detailed information they need to better understand their participants and potentially make their programs more accountable and effective in reaching their target populations. In this working paper, we describe the standardized tables themselves as well as the data assets available at the Census Bureau to create these tables, the data files produced by the table production process, and the methodology used to merge and harmonize data on participants and subsequently calculate unbiased and accurate estimates. We conclude with a brief discussion of steps taken to ensure confidentiality and data security. This documentation is intended to facilitate proper use and understanding of the standardized tables by partner agencies as well as researchers who are interested in leveraging these tools to explore characteristics of their samples of interest.

Document Tags and Keywords

Keywords Keywords are automatically generated using KeyBERT, a powerful and innovative keyword extraction tool that utilizes BERT embeddings to ensure high-quality and contextually relevant keywords.

By analyzing the content of working papers, KeyBERT identifies terms and phrases that capture the essence of the text, highlighting the most significant topics and trends. This approach not only enhances searchability but provides connections that go beyond potentially domain-specific author-defined keywords.
:
analysis, economist, researcher, statistical, data census, census data, survey, study, agency, respondent, research, yearly, statistician, federal, population, census bureau, census employment, census survey, paper census, individuals census


Similar Working Papers Similarity between working papers are determined by an unsupervised neural network model know as Doc2Vec.

Doc2Vec is a model that represents entire documents as fixed-length vectors, allowing for the capture of semantic meaning in a way that relates to the context of words within the document. The model learns to associate a unique vector with each document while simultaneously learning word vectors, enabling tasks such as document classification, clustering, and similarity detection by preserving the order and structure of words. The document vectors are compared using cosine similarity/distance to determine the most similar working papers. Papers identified with 🔥 are in the top 20% of similarity.

The 10 most similar working papers to the working paper 'Integrating Multiple U.S. Census Bureau Data Assets to Create Standardized Profiles of Program Participants' are listed below in order of similarity.