Scholars deploy census-based measures of neighborhood context throughout the social sciences and epidemiology. Decades of research confirm that variation in how individuals are aggregated into geographic units to create variables that control for social, economic or political contexts can dramatically alter analyses. While most researchers are aware of the problem, they have lacked the tools to determine its magnitude in the literature and in their own projects. By using confidential access to the complete 2010 U.S. Decennial Census, we are able to construct'for all persons in the US'individual-specific contexts, which we group according to the Census-assigned block, block group, and tract. We compare these individual-specific measures to the published statistics at each scale, and we then determine the magnitude of variation in context for an individual with respect to the published measures using a simple statistic, the standard deviation of individual context (SDIC). For three key measures (percent Black, percent Hispanic, and Entropy'a measure of ethno-racial diversity), we find that block-level Census statistics frequently do not capture the actual context of individuals within them. More problematic, we uncover systematic spatial patterns in the contextual variables at all three scales. Finally, we show that within-unit variation is greater in some parts of the country than in others. We publish county-level estimates of the SDIC statistics that enable scholars to assess whether mis-specification in context variables is likely to alter analytic findings when measured at any of the three common Census units.
-
Improving Estimates of Neighborhood Change with Constant Tract Boundaries
May 2022
Working Paper Number:
CES-22-16
Social scientists routinely rely on methods of interpolation to adjust available data to their research needs. This study calls attention to the potential for substantial error in efforts to harmonize data to constant boundaries using standard approaches to areal and population interpolation. We compare estimates from a standard source (the Longitudinal Tract Data Base) to true values calculated by re-aggregating original 2000 census microdata to 2010 tract areas. We then demonstrate an alternative approach that allows the re-aggregated values to be publicly disclosed, using 'differential privacy' (DP) methods to inject random noise to protect confidentiality of the raw data. The DP estimates are considerably more accurate than the interpolated estimates. We also examine conditions under which interpolation is more susceptible to error. This study reveals cause for greater caution in the use of interpolated estimates from any source. Until and unless DP estimates can be publicly disclosed for a wide range of variables and years, research on neighborhood change should routinely examine data for signs of estimation error that may be substantial in a large share of tracts that experienced complex boundary changes.
View Full
Paper PDF
-
Metropolitan Segregation: No Breakthrough in Sight
May 2022
Working Paper Number:
CES-22-14
The 2020 Census offers new information on changes in residential segregation in metropolitan regions across the country as they continue to become more diverse. We take a long view, assessing trends since 1980 and extrapolating to the future. These new data mostly reinforce patterns that were observed a decade ago: high but slowly declining black-white segregation, and less intense but hardly changing segregation of Hispanics and Asians from whites. Enough time has passed since the civil rights era of the 1960s and 1970s to draw this conclusion: segregation will continue to divide Americans well into the 21st Century.
View Full
Paper PDF
-
Peer Income Exposure Across the Income Distribution
February 2025
Working Paper Number:
CES-25-16
Children from families across the income distribution attend public schools, making schools and classrooms potential sites for interaction between more- and less-affluent children. However, limited information exists regarding the extent of economic integration in these contexts. We merge educational administrative data from Oregon with measures of family income derived from IRS records to document student exposure to economically diverse school and classroom peers. Our findings indicate that affluent children in public schools are relatively isolated from their less affluent peers, while low- and middle-income students experience relatively even peer income distributions. Students from families in the top percentile of the income distribution attend schools where 20 percent of their peers, on average, come from the top five income percentiles. A large majority of the differences in peer exposure that we observe arise from the sorting of students across schools; sorting across classrooms within schools plays a substantially smaller role.
View Full
Paper PDF
-
WHITE-LATINO RESIDENTIAL ATTAINMENTS AND SEGREGATION
IN SIX CITIES: ASSESSING THE ROLE OF MICRO-LEVEL FACTORS
January 2016
Working Paper Number:
CES-16-51
This study examines the residential outcomes of Latinos in major metropolitan areas using new methods to connect micro-level analyses of residential attainments to overall patterns of segregation in the metropolitan area. Drawing on new formulations of standard measures of evenness, we conduct micro-level multivariate analyses using the restricted-use census microdata files to predict segregation-relevant neighborhood outcomes for individuals by race. We term the dependent variables segregation-relevant neighborhood outcomes because the differences in average outcomes for each group on these variables determine the values of the aggregate measures of evenness. This approach allows me to use standardization and components analysis to quantitatively assess the separate contributions that differences in social characteristics and differences in rates of return make towards determining the overall disparity in residential outcomes ' that is, the level of segregation ' between Whites and Latinos. Based on our micro-level residential attainment analyses we find that for Latinos, acculturation and gains in socioeconomic status are associated with greater residential contact with Whites, in agreement with spatial assimilation theory, which promotes lower segregation. However, our standardization and components analyses reveals that a substantial portion of White-Latino disparities in residential contact with Whites can be attributed to differences in rates of return; that is White-Latino differences in the ability to translate acculturation and gains in socioeconomic status into more residential contact with Whites. This is further elaborated upon by assessing the changes in contact with Whites for Whites and Latinos after manipulating single variables while holding all others constant. This can be interpreted as the role of discrimination which is emphasized by place stratification theory. Therefore we conclude that while members of minority groups make gains in residential outcomes that reduce segregation by attaining parity with Whites on social characteristics as spatial assimilation theory would predict, a substantial disparity will persist as Latinos cannot translate those gains into greater contact with Whites at the rate that Whites can. At the aggregate level of analysis, this means that White-Latino segregation remains substantial even when groups are equalized on social and economic characteristics.
View Full
Paper PDF
-
Dynamics of Race: Joining, Leaving, and Staying in the American Indian/Alaska Native Race Category between 2000 and 2010
August 2014
Working Paper Number:
carra-2014-10
Each census for decades has seen the American Indian and Alaska Native population increase substantially more than expected. Changes in racial reporting seem to play an important role in the observed net increases, though research has been hampered by data limitations. We address previously unanswerable questions about race response change among American Indian and Alaska Natives (hereafter 'American Indians') using uniquely-suited (but not nationally representative) linked data from the 2000 and 2010 decennial censuses (N = 3.1 million) and the 2006-2010 American Community Survey (N = 188,131). To what extent do people change responses to include or exclude American Indian? How are people who change responses similar to or different from those who do not? How are people who join a group similar to or different from those who leave it? We find considerable race response change by people in our data, especially by multiple-race and/or Hispanic American Indians. This turnover is hidden in cross-sectional comparisons because people joining the group are similar in number and characteristics to those who leave the group. People in our data who changed their race response to add or drop American Indian differ from those who kept the same race response in 2000 and 2010 and from those who moved between a single-race and multiple-race American Indian response. Those who consistently reported American Indian (including those who added or dropped another race response) were relatively likely to report a tribe, live in an American Indian area, report American Indian ancestry, and live in the West. There are significant differences between those who joined and those who left a specific American Indian response group, but poor model fit indicates general similarity between joiners and leavers. Response changes should be considered when conceptualizing and operationalizing 'the American Indian and Alaska Native population.'
View Full
Paper PDF
-
Factors that Influence Change in Hispanic Identification: Evidence from Linked Decennial Census and American Community Survey Data
October 2018
Working Paper Number:
CES-18-45
This study explores patterns of ethnic boundary crossing as evidenced by changes in Hispanic origin responses across decennial census and survey data. We identify socioeconomic, cultural, and demographic factors associated with Hispanic response change. In addition, we assess whether changes in the Hispanic origin question between the 2000 and 2010 censuses influenced changes in Hispanic reporting. We use a unique large dataset that links a person's unedited responses to the Hispanic origin question across Census 2000, the 2010 Census and the 2006-2010 American Community Survey five-year file. We find that most of the individuals in the sample identified consistently as Hispanic regardless of changes in the wording of the Hispanic origin question. Individuals who changed in or out of a Hispanic identification, as well as those who consistently identified as non-Hispanic (of Hispanic ancestry), differed in socioeconomic and cultural characteristics from individuals who consistently reported as Hispanic. The likelihood of changing their Hispanic origin response is higher among U.S.-born individuals, those reporting mixed Hispanic and non-Hispanic ancestries, those who speak only English at home, and those who live in tracts that are predominantly non-Hispanic. Racial identification and detailed Hispanic background also influence changes in Hispanic origin responses. Finally, changes in mode and relationship to the reference person in the household are associated with changes in Hispanic origin responses, suggesting that data collection elements also can influence Hispanic origin response change.
View Full
Paper PDF
-
A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census
August 2025
Authors:
Lars Vilhuber,
John M. Abowd,
Ethan Lewis,
Nathan Goldschlag,
Michael B. Hawes,
Robert Ashmead,
Daniel Kifer,
Philip Leclerc,
Rolando A. RodrÃguez,
Tamara Adams,
David Darais,
Sourya Dey,
Simson L. Garfinkel,
Scott Moore,
Ramy N. Tadros
Working Paper Number:
CES-25-57
For the last half-century, it has been a common and accepted practice for statistical agencies, including the United States Census Bureau, to adopt different strategies to protect the confidentiality of aggregate tabular data products from those used to protect the individual records contained in publicly released microdata products. This strategy was premised on the assumption that the aggregation used to generate tabular data products made the resulting statistics inherently less disclosive than the microdata from which they were tabulated. Consistent with this common assumption, the 2010 Census of Population and Housing in the U.S. used different disclosure limitation rules for its tabular and microdata publications. This paper demonstrates that, in the context of disclosure limitation for the 2010 Census, the assumption that tabular data are inherently less disclosive than their underlying microdata is fundamentally flawed. The 2010 Census published more than 150 billion aggregate statistics in 180 table sets. Most of these tables were published at the most detailed geographic level'individual census blocks, which can have populations as small as one person. Using only 34 of the published table sets, we reconstructed microdata records including five variables (census block, sex, age, race, and ethnicity) from the confidential 2010 Census person records. Using only published data, an attacker using our methods can verify that all records in 70% of all census blocks (97 million people) are perfectly reconstructed. We further confirm, through reidentification studies, that an attacker can, within census blocks with perfect reconstruction accuracy, correctly infer the actual census response on race and ethnicity for 3.4 million vulnerable population uniques (persons with race and ethnicity different from the modal person on the census block) with 95% accuracy. Having shown the vulnerabilities inherent to the disclosure limitation methods used for the 2010 Census, we proceed to demonstrate that the more robust disclosure limitation framework used for the 2020 Census publications defends against attacks that are based on reconstruction. Finally, we show that available alternatives to the 2020 Census Disclosure Avoidance System would either fail to protect confidentiality, or would overly degrade the statistics' utility for the primary statutory use case: redrawing the boundaries of all of the nation's legislative and voting districts in compliance with the 1965 Voting Rights Act.
View Full
Paper PDF
-
Structural versus Ethnic Dimensions of Housing Segregation
March 2016
Working Paper Number:
CES-16-22
Racial residential segregation is still very high in many American cities. Some portion of segregation is attributable to socioeconomic differences across racial lines; some portion is caused by purely racial factors, such as preferences about the racial composition of one's neighborhood or discrimination in the housing market. Social scientists have had great difficulty disaggregating segregation into a portion that can be explained by interracial differences in socioeconomic characteristics (what we call structural factors) versus a portion attributable to racial and ethnic factors. What would such a measure look like? In this paper, we draw on a new source of data to develop an innovative structural segregation measure that shows the amount of segregation that would remain if we could assign households to housing units based only on non-racial socioeconomic characteristics. This inquiry provides vital building blocks for the broader enterprise of understanding and remedying housing segregation.
View Full
Paper PDF
-
SYNTHETIC DATA FOR SMALL AREA ESTIMATION IN THE AMERICAN COMMUNITY SURVEY
April 2013
Working Paper Number:
CES-13-19
Small area estimates provide a critical source of information used to study local populations. Statistical agencies regularly collect data from small areas but are prevented from releasing detailed geographical identifiers in public-use data sets due to disclosure concerns. Alternative data dissemination methods used in practice include releasing summary/aggregate tables, suppressing detailed geographic information in public-use data sets, and accessing restricted data via Research Data Centers. This research examines an alternative method for disseminating microdata that contains more geographical details than are currently being released in public-use data files. Specifically, the method replaces the observed survey values with imputed, or synthetic, values simulated from a hierarchical Bayesian model. Confidentiality protection is enhanced because no actual values are released. The method is demonstrated using restricted data from the 2005-2009 American Community Survey. The analytic validity of the synthetic data is assessed by comparing small area estimates obtained from the synthetic data with those obtained from the observed data.
View Full
Paper PDF
-
Associations Between Public Housing and Individual Earnings in New Orleans
October 2015
Working Paper Number:
CES-15-32
This study uses a sample of the civilian labor force aged 16-64 constructed from the Decennial Census and American Community Survey, along with data from the HUD dataset Picture of Subsidized Households, to compare the likelihood for job earnings in relation to public housing developments in the New Orleans MSA before and after Hurricane Katrina. Results from a series of hierarchical linear models (HLM) indicate significant relationships are altered between time periods, including those from public and mixed-income developments, suggesting a fluid relationship between neighborhoods and economic outcomes during physical, demographic and economic restructuring.
View Full
Paper PDF