This paper describes a novel database and an associated suicide event prediction model that surmount longstanding barriers in suicide risk factor research. The database comingles person-level records from the National Violent Death Reporting System (NVDRS) and the American Community Survey (ACS) to establish a case-control study sample that includes all identified suicide cases, while faithfully reflecting general population sociodemographics, in sixteen USA states during the years 2005 2011. It supports a statistical model of individual suicide risk that accommodates person-level factors and the moderation of these factors by their community rates. Named the United States Multi-Level Suicide Data Set (US-MSDS), the database was developed outside the RDC laboratory using publicly available ACS microdata, and reconstructed inside the laboratory using restricted access ACS microdata. Analyses of the latter version yielded findings that largely amplified but also extended those obtained from analyses of the former. This experience shows that the analytic precision achievable using restricted access ACS data can play an important role in conducting social research, although it also indicates that publicly available ACS data have considerable value in conducting preliminary analyses and preparing to use an RDC laboratory. The database development strategy may interest scientists investigating sociodemographic risk factors for other types of low-frequency mortality.
-
Mortality in a Multi-State Cohort of Former State Prisoners, 2010-2015
February 2022
Working Paper Number:
CES-22-06
Previous studies report that individuals who have been imprisoned have higher mortality rates than their demographic counterparts in the general population, particularly non-Hispanic white former prisoners. Most of these studies have been based on a single state's prison system, and the extent to which their findings can be generalized has not been established. In this study we explore the role that race/Hispanic origin, other demographic characteristics, and custodial/ criminal history factors have on post-release mortality, including on the timing of deaths. We also assess whether conditional release to community supervision or reimprisonment may explain the higher post-release mortality found among non-Hispanic whites. In the second part of the analysis, we estimate standardized mortality ratios (SMRs) by sex, age group, and race/Hispanic origin using as reference the U.S. general population. The data come from state prison releases from the Bureau of Justice Statistics' (BJS) National Corrections Reporting Program (NCRP). The NCRP records were linked to the Census Numident to identify deaths occurring within five years from prison release. We also linked NCRP records to previous decennial censuses and survey responses to obtain self-reported race and Hispanic origin if available. We found that non-Hispanic white former prisoners were more likely to die within five years after prison release and more likely to die in the initial weeks after release compared to racial minorities and Hispanics. Reimprisonment, age at release, and a history of multiple prison terms had a similar influence on the odds of dying across all race/Hispanic origin groups. Other factors, such as the type of release and the duration of the last term in prison, were associated with higher risks of mortality for some groups but not for others.
View Full
Paper PDF
-
Leaving Home: Modeling the Effect of Civic and Economic Structure on Individual Migration Patterns
June 2002
Working Paper Number:
CES-02-16
This research analyzes the effect of community structure upon individuals' probabilities of moving between 1985 and 1990. Using the full Census sample long form microdata for 1990, we re-allocate adult persons in 1990 to their 1985 county of residence. Then, using origin county macro-structural variables (derived from the Economic Census microdata) and individual characteristics (from Decennial Census microdata), we develop a two level hierarchical linear model. In level 1, we construct a logistic equation modeling individual probabilities of moving. In level 2, we model the contextual effects of origin community structure on these models. These contextual effects fall into two categories: 1) economic conditions that comprise the usual aggregate 'push' factors and 2) civic community factors that act to retain people in their community. Results specify the relationship between community context and individual migration patterns, and demonstrate effects of local economic structure and local civic structure on these individual probabilities. Most notably, we find that civic attributes of communities are associated with a propensity to stay in place, net of community economic factors and individual characteristics.
View Full
Paper PDF
-
Gradient Boosting to Address Statistical Problems Arising from Non-Linkage of Census Bureau Datasets
June 2024
Working Paper Number:
CES-24-27
This article introduces the twangRDC package, which contains functions to address non-linkage in US Census Bureau datasets. The Census Bureau's Person Identification Validation System facilitates data linkage by assigning unique person identifiers to federal, third party, decennial census, and survey data. Not all records in these datasets can be linked to the reference file and as such not all records will be assigned an identifier. This article is a tutorial for using the twangRDC to generate nonresponse weights to account for non-linkage of person records across US Census Bureau datasets.
View Full
Paper PDF
-
Who are the people in my neighborhood? The 'contextual fallacy' of measuring individual context with census geographies
February 2018
Working Paper Number:
CES-18-11
Scholars deploy census-based measures of neighborhood context throughout the social sciences and epidemiology. Decades of research confirm that variation in how individuals are aggregated into geographic units to create variables that control for social, economic or political contexts can dramatically alter analyses. While most researchers are aware of the problem, they have lacked the tools to determine its magnitude in the literature and in their own projects. By using confidential access to the complete 2010 U.S. Decennial Census, we are able to construct'for all persons in the US'individual-specific contexts, which we group according to the Census-assigned block, block group, and tract. We compare these individual-specific measures to the published statistics at each scale, and we then determine the magnitude of variation in context for an individual with respect to the published measures using a simple statistic, the standard deviation of individual context (SDIC). For three key measures (percent Black, percent Hispanic, and Entropy'a measure of ethno-racial diversity), we find that block-level Census statistics frequently do not capture the actual context of individuals within them. More problematic, we uncover systematic spatial patterns in the contextual variables at all three scales. Finally, we show that within-unit variation is greater in some parts of the country than in others. We publish county-level estimates of the SDIC statistics that enable scholars to assess whether mis-specification in context variables is likely to alter analytic findings when measured at any of the three common Census units.
View Full
Paper PDF
-
Geographic Disparities in Alzheimer's Disease and Related Dementia Mortality in the US: Comparing Impacts of Place of Birth and Place of Residence
January 2025
Working Paper Number:
CES-25-11
Objective: Building on the hypothesis that early-life exposures might influence the onset of Alzheimer's Disease and Related Dementia (ADRD), this study delves into geographic variations in ADRD mortality in the US. By considering both state of residence and state of birth, we aim to discern the comparative significance of these geospatial factors.
Methods: We conducted a secondary data analysis of the National Longitudinal Mortality Study (NLMS), that has 3.5 million records from 1973-2011 and over 0.5 million deaths. We focused on individuals born in or before 1930, tracked in NLMS cohorts from 1979-2000. Employing multi-level logistic regression, with individuals nested within states of residence and/or states of birth, we assessed the role of geographical factors in ADRD mortality variation.
Results: We found that both state of birth and state of residence account for a modest portion of ADRD mortality variation. Specifically, state of residence explains 1.19% of the total variation in ADRD mortality, whereas state of birth explains only 0.6%. When combined, both state of residence and state of birth account for only 1.05% of the variation, suggesting state of residence could matter more in ADRD mortality outcomes.
Conclusion: Findings of this study suggest that state of residence explains more variation in ADRD mortality than state of birth. These results indicate that factors in later life may present more impactful intervention points for curbing ADRD mortality. While early-life environmental exposures remain relevant, their role as primary determinants of ADRD in later life appears to be less pronounced in this study.
View Full
Paper PDF
-
Estimating the Impact of the Age of Criminal Majority: Decomposing Multiple Treatments in a Regression Discontinuity Framework
January 2023
Working Paper Number:
CES-23-01
This paper studies the impact of adult prosecution on recidivism and employment trajectories for adolescent, first-time felony defendants. We use extensive linked Criminal Justice Administrative Record System and socio-economic data from Wayne County, Michigan (Detroit). Using the discrete age of majority rule and a regression discontinuity design, we find that adult prosecution reduces future criminal charges over 5 years by 0.48 felony cases (? 20%) while also worsening labor market outcomes: 0.76 fewer employers (? 19%) and $674 fewer earnings (? 21%) per year. We develop a novel econometric framework that combines standard regression discontinuity methods with predictive machine learning models to identify mechanism-specific treatment effects that underpin the overall impact of adult prosecution. We leverage these estimates to consider four policy counterfactuals: (1) raising the age of majority, (2) increasing adult dismissals to match the juvenile disposition rates, (3) eliminating adult incarceration, and (4) expanding juvenile record sealing opportunities to teenage adult defendants. All four scenarios generate positive returns for government budgets. When accounting for impacts to defendants as well as victim costs borne by society stemming from increases in recidivism, we find positive social returns for juvenile record sealing expansions and dismissing marginal adult charges; raising the age of majority breaks even. Eliminating prison for first-time adult felony defendants, however, increases net social costs. Policymakers may still find this attractive if they are willing to value beneficiaries (taxpayers and defendants) slightly higher (124%) than potential victims.
View Full
Paper PDF
-
Evaluating Race and Hispanic Origin Responses of Medicaid Participants Using Census Data
April 2015
Working Paper Number:
carra-2015-01
Health and health care disparities associated with race or Hispanic origin are complex and continue to challenge researchers and policy makers. With the intention of improving the measurement and monitoring of these disparities, provisions of the Patient Protection and Affordable Care Act (ACA) of 2010 require states to collect, report and analyze data on demographic characteristics of applicants and participants in Medicaid and other federally supported programs. By linking Medicaid records to 2010 Census, American Community Survey, and Census 2000, this new large-scale study examines and documents the extent to which pre-ACA Medicaid administrative records match self-reported race and Hispanic origin in Census data. Linked records allow comparisons between individuals with matching and non-matching race and Hispanic origin data across several demographic, socioeconomic and neighborhood characteristics, such as age, gender, language proficiency, education and Census tract variables. Identification of the groups most likely to have non-matching and missing race and Hispanic origin data in Medicaid relative to Census data can inform strategies to improve the quality of demographic data collected from Medicaid populations.
View Full
Paper PDF
-
Shift or replenishment? Reassessing the prospect of stable Spanish bilingualism across contexts of ethnic change
June 2023
Working Paper Number:
CES-23-28
Much of the existing literature on Latinos' use of Spanish claims that a general pattern of intergenerational decline in the use of Spanish will produce an overall shift away from Spanish use in the U.S. (Rumbaut, Massey, and Bean 2006; Veltman 1983b, 1990). In contrast, recent works emphasize the importance of the social and linguistic context in reinforcing the use of Spanish as well as (pan)ethnic identities among U.S.-born Latinos (Linton 2004; Linton and Jim'nez 2009; Stevens 1992). This literature suggests conditions under which Spanish-English bilingualism might become stable at the level of metropolitan areas; however, such conditions depend on how immigration shapes the context of language use for native-born Latinos. Given the declining levels of immigration from Latin America, will bilingualism subside in the U.S., or have certain communities created conditions in which bilingualism can be stable? Using geocoded data from restricted access versions of the Survey of Income and Program Participation (SIPP) and the American Community Survey (ACS), we model the probability of Spanish-English bilingualism among second- and third-generation Latinos using multilevel models with contextual measures of immigration and language use at both the neighborhood and metropolitan levels. We find evidence that U.S.-born Latinos are heavily influenced by the prevalence of Spanish use among U.S. born Latinos at both the metropolitan and neighborhood levels. Further, the proportion of foreign-born Latinos has little effect on the native born, after controlling for Spanish use among U.S,-born Latinos. These results are a first step in understanding the link between ethnic or panethnic contexts and language practices, and also in producing a better characterization of stable bilingualism that can be tested quantitatively.
View Full
Paper PDF
-
A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census
August 2025
Authors:
Lars Vilhuber,
John M. Abowd,
Ethan Lewis,
Nathan Goldschlag,
Michael B. Hawes,
Robert Ashmead,
Daniel Kifer,
Philip Leclerc,
Rolando A. RodrÃguez,
Tamara Adams,
David Darais,
Sourya Dey,
Simson L. Garfinkel,
Scott Moore,
Ramy N. Tadros
Working Paper Number:
CES-25-57
For the last half-century, it has been a common and accepted practice for statistical agencies, including the United States Census Bureau, to adopt different strategies to protect the confidentiality of aggregate tabular data products from those used to protect the individual records contained in publicly released microdata products. This strategy was premised on the assumption that the aggregation used to generate tabular data products made the resulting statistics inherently less disclosive than the microdata from which they were tabulated. Consistent with this common assumption, the 2010 Census of Population and Housing in the U.S. used different disclosure limitation rules for its tabular and microdata publications. This paper demonstrates that, in the context of disclosure limitation for the 2010 Census, the assumption that tabular data are inherently less disclosive than their underlying microdata is fundamentally flawed. The 2010 Census published more than 150 billion aggregate statistics in 180 table sets. Most of these tables were published at the most detailed geographic level'individual census blocks, which can have populations as small as one person. Using only 34 of the published table sets, we reconstructed microdata records including five variables (census block, sex, age, race, and ethnicity) from the confidential 2010 Census person records. Using only published data, an attacker using our methods can verify that all records in 70% of all census blocks (97 million people) are perfectly reconstructed. We further confirm, through reidentification studies, that an attacker can, within census blocks with perfect reconstruction accuracy, correctly infer the actual census response on race and ethnicity for 3.4 million vulnerable population uniques (persons with race and ethnicity different from the modal person on the census block) with 95% accuracy. Having shown the vulnerabilities inherent to the disclosure limitation methods used for the 2010 Census, we proceed to demonstrate that the more robust disclosure limitation framework used for the 2020 Census publications defends against attacks that are based on reconstruction. Finally, we show that available alternatives to the 2020 Census Disclosure Avoidance System would either fail to protect confidentiality, or would overly degrade the statistics' utility for the primary statutory use case: redrawing the boundaries of all of the nation's legislative and voting districts in compliance with the 1965 Voting Rights Act.
View Full
Paper PDF
-
Dynamics of Race: Joining, Leaving, and Staying in the American Indian/Alaska Native Race Category between 2000 and 2010
August 2014
Working Paper Number:
carra-2014-10
Each census for decades has seen the American Indian and Alaska Native population increase substantially more than expected. Changes in racial reporting seem to play an important role in the observed net increases, though research has been hampered by data limitations. We address previously unanswerable questions about race response change among American Indian and Alaska Natives (hereafter 'American Indians') using uniquely-suited (but not nationally representative) linked data from the 2000 and 2010 decennial censuses (N = 3.1 million) and the 2006-2010 American Community Survey (N = 188,131). To what extent do people change responses to include or exclude American Indian? How are people who change responses similar to or different from those who do not? How are people who join a group similar to or different from those who leave it? We find considerable race response change by people in our data, especially by multiple-race and/or Hispanic American Indians. This turnover is hidden in cross-sectional comparisons because people joining the group are similar in number and characteristics to those who leave the group. People in our data who changed their race response to add or drop American Indian differ from those who kept the same race response in 2000 and 2010 and from those who moved between a single-race and multiple-race American Indian response. Those who consistently reported American Indian (including those who added or dropped another race response) were relatively likely to report a tribe, live in an American Indian area, report American Indian ancestry, and live in the West. There are significant differences between those who joined and those who left a specific American Indian response group, but poor model fit indicates general similarity between joiners and leavers. Response changes should be considered when conceptualizing and operationalizing 'the American Indian and Alaska Native population.'
View Full
Paper PDF