CREAT - Census Bureau

Evaluation of Commercial School and Teacher Lists to Enhance Survey Frames

July 2014

Written by: Quentin Brummet, Mark Masterto, Damon Smith

Working Paper Number:

carra-2014-07

Abstract

This report summarizes the potential for teacher lists obtained from commercial vendors for enhancing sampling frames for the National Teacher and Principal Survey (NTPS). We investigate three separate vendor lists, and compare coverage rates across a range of school and teacher characteristics. Across all vendors, coverage rates are higher for regular, non-charter schools. Vendor A stands out as having higher coverage rates than the other two, and we recommend further evaluating Vendor A's teacher lists during the upcoming 2014-2015 NTPS Field Test.

Document Tags and Keywords

Keywords:

data, database, survey, respondent, education, student, district, sampling, sample, school, grade, educational

Tags:

Department of Defense, National Center for Health Statistics

Similar Working Papers

The 10 most similar working papers to the working paper 'Evaluation of Commercial School and Teacher Lists to Enhance Survey Frames' are listed below in order of similarity.

Working Paper

Assessing Coverage and Quality of the 2007 Prototype Census Kidlink Database

September 2015

Authors: Adela Luque, Deborah Wagner

Working Paper Number:

carra-2015-07

The Census Bureau is conducting research to expand the use of administrative records data in censuses and surveys to decrease respondent burden and reduce costs while improving data quality. Much of this research (e.g., Rastogi and O''Hara (2012), Luque and Bhaskar (2014)) hinges on the ability to integrate multiple data sources by linking individuals across files. One of the Census Bureau's record linkage methodologies for data integration is the Person Identification Validation System or PVS. PVS assigns anonymous and unique IDs (Protected Identification Keys or PIKs) that serve as linkage keys across files. Prior research showed that integrating 'known associates' information into PVS's reference files could potentially enhance PVS's PIK assignment rates. The term 'known associates' refers to people that are likely to be associated with each other because of a known common link (such as family relationships or people sharing a common address), and thus, to be observed together in different files. One of the results from this prior research was the creation of the 2007 Census Kidlink file, a child-level file linking a child's Social Security Number (SSN) record to the SSN of those identified as the child's parents. In this paper, we examine to what extent the 2007 Census Kidlink methodology was able to link parents SSNs to children SSN records, and also evaluate the quality of those links. We find that in approximately 80 percent of cases, at least one parent was linked to the child's record. Younger children and noncitizens have a higher percentage of cases where neither parent could be linked to the child. Using 2007 tax data as a benchmark, our quality evaluation results indicate that in at least 90 percent of the cases, the parent-child link agreed with those found in the tax data. Based on our findings, we propose improvements to the 2007 Kidlink methodology to increase child-parent links, and discuss how the creation of the file could be operationalized moving forward.
View Full Paper PDF
Working Paper

Comparison of Survey, Federal, and Commercial Address Data Quality

June 2014

Authors: Quentin Brummet

Working Paper Number:

carra-2014-06

This report summarizes matching of survey, commercial, and administrative records housing units to the Census Bureau Master Address File (MAF). We document overall MAF match rates in each data set and evaluate differences in match rates across a variety of housing characteristics. Results show that over 90 percent of records in survey data from the American Housing Survey (AHS) match to the MAF. Commercial data from CoreLogic matches at much lower rates, in part due to missing address information and poor match rates for multi-unit buildings. MAF match rates for administrative records from the Department of Housing and Urban Development are also high, and open the possibility of using this information in surveys such as the AHS.
View Full Paper PDF
Working Paper

Person Matching in Historical Files using the Census Bureau's Person Validation System

September 2014

Authors: Amy B. O'Hara, Catherine G. Massey, Amy OHara

Working Paper Number:

carra-2014-11

The recent release of the 1940 Census manuscripts enables the creation of longitudinal data spanning the whole of the twentieth century. Linked historical and contemporary data would allow unprecedented analyses of the causes and consequences of health, demographic, and economic change. The Census Bureau is uniquely equipped to provide high quality linkages of person records across datasets. This paper summarizes the linkage techniques employed by the Census Bureau and discusses utilization of these techniques to append protected identification keys to the 1940 Census.
View Full Paper PDF
Working Paper

The Design of Sampling Strata for the National Household Food Acquisition and Purchase Survey

February 2025

Authors: Jonathan Eggleston, Mark Klee, Linden McBride

Working Paper Number:

CES-25-13

The National Household Food Acquisition and Purchase Survey (FoodAPS), sponsored by the United States Department of Agriculture's (USDA) Economic Research Service (ERS) and Food and Nutrition Service (FNS), examines the food purchasing behavior of various subgroups of the U.S. population. These subgroups include participants in the Supplemental Nutrition Assistance Program (SNAP) and the Special Supplemental Nutrition Program for Women, Infants, and Children (WIC), as well as households who are eligible for but don't participate in these programs. Participants in these social protection programs constitute small proportions of the U.S. population; obtaining an adequate number of such participants in a survey would be challenging absent stratified sampling to target SNAP and WIC participating households. This document describes how the U.S. Census Bureau (which is planning to conduct future versions of the FoodAPS survey on behalf of USDA) created sampling strata to flag the FoodAPS targeted subpopulations using machine learning applications in linked survey and administrative data. We describe the data, modeling techniques, and how well the sampling flags target low-income households and households receiving WIC and SNAP benefits. We additionally situate these efforts in the nascent literature on the use of big data and machine learning for the improvement of survey efficiency.
View Full Paper PDF
Working Paper

The Effect of Class Size on Teacher Attrition: Evidence from Class Size Reduction Policies in New York State

February 2010

Authors: Emily Isenberg

Working Paper Number:

CES-10-05

Starting in 1999, New York State implemented class size reduction policies targeted at early elementary grades, but due to funding limitations, most schools reduced class size in some grades and not others. I use class size variation within a school induced by the policies to construct instrumental variable estimates of the effect of class size on teacher attrition. Teachers with smaller classes were not significantly less likely to leave schools in the full sample of districts but were less likely to leave a school in districts that targeted the same grade across schools. District-wide class size reduction policies were more likely to persist in the same grade in the next year, suggesting that teacher expectations of continued smaller classes played a role in their decision whether or not to leave a school. A decrease in class size from 23 to 20 students (a decrease of one standard deviation) under a district-wide policy decreases the probability that a teacher leaves a school by 4.2 percentage points.
View Full Paper PDF
Working Paper

Matching Addresses between Household Surveys and Commercial Data

July 2015

Authors: Quentin Brummet

Working Paper Number:

carra-2015-04

Matching third-party data sources to household surveys can benefit household surveys in a number of ways, but the utility of these new data sources depends critically on our ability to link units between data sets. To understand this better, this report discusses potential modifications to the existing match process that could potentially improve our matches. While many changes to the matching procedure produce marginal improvements in match rates, substantial increases in match rates can only be achieved by relaxing the definition of a successful match. In the end, the results show that the most important factor determining the success of matching procedures is the quality and composition of the data sets being matched.
View Full Paper PDF
Working Paper

Methodology on Creating the U.S. Linked Retail Health Clinic (LiRHC) Database

March 2023

Authors: Alice Zawacki, Joey Marshall, Donald Cherry, Xianghua Yin, Brian W. Ward

Working Paper Number:

CES-23-10

Retail health clinics (RHCs) are a relatively new type of health care setting and understanding the role they play as a source of ambulatory care in the United States is important. To better understand these settings, a joint project by the Census Bureau and National Center for Health Statistics used data science techniques to link together data on RHCs from Convenient Care Association, County Business Patterns Business Register, and National Plan and Provider Enumeration System to create the Linked RHC (LiRHC, pronounced 'lyric') database of locations throughout the United States during the years 2018 to 2020. The matching methodology used to perform this linkage is described, as well as the benchmarking, match statistics, and manual review and quality checks used to assess the resulting matched data. The large majority (81%) of matches received quality scores at or above 75/100, and most matches were linked in the first two (of eight) matching passes, indicating high confidence in the final linked dataset. The LiRHC database contained 2,000 RHCs and found that 97% of these clinics were in metropolitan statistical areas and 950 were in the South region of the United States. Through this collaborative effort, the Census Bureau and National Center for Health Statistics strive to understand how RHCs can potentially impact population health as well as the access and provision of health care services across the nation.
View Full Paper PDF
Working Paper

Creating Linked Historical Data: An Assessment of the Census Bureau's Ability to Assign Protected Identification Keys to the 1960 Census

September 2014

Authors: Catherine G. Massey

Working Paper Number:

carra-2014-12

In order to study social phenomena over the course of the 20th century, the Census Bureau is investigating the feasibility of digitizing historical census records and linking them to contemporary data. However, historical censuses have limited personally identifiable information available to match on. In this paper, I discuss the problems associated with matching older censuses to contemporary data files, and I describe the matching process used to match a small sample of the 1960 census to the Social Security Administration Numeric Identification System.
View Full Paper PDF
Working Paper

Comparing the 2019 American Housing Survey to Contemporary Sources of Property Tax Records: Implications for Survey Efficiency and Quality

June 2022

Authors: John Voorheis, Ariel J. Binder, Emily Molfino

Working Paper Number:

CES-22-22

Given rising nonresponse rates and concerns about respondent burden, government statistical agencies have been exploring ways to supplement household survey data collection with administrative records and other sources of third-party data. This paper evaluates the potential of property tax assessment records to improve housing surveys by comparing these records to responses from the 2019 American Housing Survey. Leveraging the U.S. Census Bureau's linkage infrastructure, we compute the fraction of AHS housing units that could be matched to a unique property parcel (coverage rate), as well as the extent to which survey and property tax data contain the same information (agreement rate). We analyze heterogeneity in coverage and agreement across states, housing characteristics, and 11 AHS items of interest to housing researchers. Our results suggest that partial replacement of AHS data with property data, targeted toward certain survey items or single-family detached homes, could reduce respondent burden without altering data quality. Further research into partial-replacement designs is needed and should proceed on an item-by-item basis. Our work can guide this research as well as those who wish to conduct independent research with property tax records that is representative of the U.S. housing stock.
View Full Paper PDF
Working Paper

Some Open Questions on Multiple-Source Extensions of Adaptive-Survey Design Concepts and Methods

February 2023

Authors: Stephanie Coffey, PhD., Jaya Damineni, John Eltinge, PhD., Anup Mathur, PhD., Kayla Varela, Allison Zotti

Working Paper Number:

CES-23-03

Adaptive survey design is a framework for making data-driven decisions about survey data collection operations. This paper discusses open questions related to the extension of adaptive principles and capabilities when capturing data from multiple data sources. Here, the concept of 'design' encompasses the focused allocation of resources required for the production of high-quality statistical information in a sustainable and cost-effective way. This conceptual framework leads to a discussion of six groups of issues including: (i) the goals for improvement through adaptation; (ii) the design features that are available for adaptation; (iii) the auxiliary data that may be available for informing adaptation; (iv) the decision rules that could guide adaptation; (v) the necessary systems to operationalize adaptation; and (vi) the quality, cost, and risk profiles of the proposed adaptations (and how to evaluate them). A multiple data source environment creates significant opportunities, but also introduces complexities that are a challenge in the production of high-quality statistical information.
View Full Paper PDF

Evaluation of Commercial School and Teacher Lists to Enhance Survey Frames

July 2014

Working Paper Number:

carra-2014-07

Abstract

Document Tags and Keywords

The 10 most similar working papers to the working paper 'Evaluation of Commercial School and Teacher Lists to Enhance Survey Frames' are listed below in order of similarity.

September 2015

Working Paper Number:

carra-2015-07

June 2014

Working Paper Number:

carra-2014-06

September 2014

Working Paper Number:

carra-2014-11

February 2025

Working Paper Number:

CES-25-13

February 2010

Working Paper Number:

CES-10-05

July 2015

Working Paper Number:

carra-2015-04

March 2023

Working Paper Number:

CES-23-10

September 2014

Working Paper Number:

carra-2014-12

June 2022

Working Paper Number:

CES-22-22

February 2023

Working Paper Number:

CES-23-03