The Quackery of Columbia’s Racialized Medical Research

A recent article by Chris Rufo and Hannah Grossman of the Manhattan Institute offers an unflattering profile of Dr. Jennifer Manly, a neuropsychologist and Columbia professor. Manly is an agitator within the pro-Hamas campus movement. Her activism spills into her research, much of it based on “the so-called social determinants of health thesis, which posits that racism, sexism, and homophobia can cause brain disease in ‘Black and Latinx communities.’”

Rufo and Grossman note that critics have described the thesis as “pseudo-science.” Manly and her defenders would surely disagree, pointing to her impressive collection of scholarly publications and citations. However, peer review in medical journals is often just as much a screen for ideology as it is one for rigor. A careful examination of a study that Manly recently co-authored makes for a useful illustration of just how faulty this line of research is, and why it is incumbent upon the NIH to stop funding it.

The study, “’Rest of the folks are tired and weary’: The impact of historical lynchings on biological and cognitive health for older adults racialized as Black,” was published in the journal, Social Science & Medicine, by ten authors, with Manly’s name appearing last.

From Social Science and Medicine, January 2025.

As the title suggests, the study examines the effect of lynchings on health outcomes. It claims to find that residing in states that historically had more lynchings of black victims causes black subjects to experience a greater increase in a measure of inflammation and a greater decline in cognitive function. The theory it offers to account for these findings is that experiencing lynchings activates a “stress response in early childhood” that contributes to adverse health outcomes later in life.

The article only supports this claim by using a convoluted and indefensible research design and by interpreting the results in implausible ways. Rather than devising a measure of lynching that might capture the likelihood that one was exposed to lynchings, they discard information and dichotomize the 37 states that had any lynchings of black people into having above median or below median number of lynchings. As they describe it, “we dichotomized the variable at the 50th percentile (median), where less than the median corresponded to states with 1 lynching and greater than the median corresponded to states that had 2 or more lynchings.”

As one of many examples of sloppiness, the median number of black victims of lynchings among the 37 states with at least one lynching is 16, not 1 or 2. And the median among all 44 states in the Tuskegee Institute’s data set is 5. Another egregious example of sloppiness is that the article mis-describes the Emancipation Proclamation as having been issued in 1864, when it was actually issued on January 1, 1863. And in another error, the text describes the main result inconsistently with how it is described in the tables. [i]

Regardless of how the researchers divided states into above and below median lynching categories, dichotomizing the data threw away information that prevented a more fine-grained examination of the health effects of lynchings. The effect of living in Mississippi, which had 539 black victims of lynchings between 1882 and1968, is treated the same as living in Virginia, with 83 lynchings, or Illinois, with 19. All would be above median, however they calculated it.

The researchers make no adjustment for the population of states, so that California, with 2 black victims of lynching, is treated the same as Montana, which also has 2. If the researchers are hoping to measure exposure to lynching, failing to differentiate between the size of states is a serious shortcoming.

But the most serious failure of the study is that the results of its fully specified model do not find that health outcomes are worse among black subjects in states classified as having a higher number of historic lynchings. As can be seen in Model 5 of Table 3, being in a state with above median historic lynchings is not significantly related to the CRP measure of inflammation among black subjects.

From “’Rest of the folks are tired and weary’: The impact of historical lynchings on biological and cognitive health for older adults racialized as Black.”

And in Model 5 of Table 4, residing in a state with an above median number of historic lynchings is not significantly related to the measure of cognitive performance among black subjects.

From “’Rest of the folks are tired and weary’: The impact of historical lynchings on biological and cognitive health for older adults racialized as Black.”

In model 5, the researchers add to their set of controls a variable for whether the state is in the Census definition of the South. So, what their research really shows is that black subjects who live in the South have worsening measures of inflammation and cognitive performance, not that historic lynchings contribute to adverse health outcomes. That is, within the South, being in a state that had above or below median numbers of lynchings is unrelated to health outcomes. And within the North, residing in a high or low lynching state also made no difference. The main factor driving their result is that health outcomes are worse in the South, even when controlling for a handful of other risk factors.

In the discussion section, the researchers attempt to dismiss these null results by noting that “the high overlap of participants living in U.S. South and a state with high lynching proportions may have induced some collinearity in our fully adjusted models.” Of course, there would be less collinearity if the researchers had not dichotomized the measure of lynchings. Regardless of their explanation, the fact remains that they have null results when controlling for region. This means that health outcomes for black subjects appear to vary based on the region in which they live, not the number of historic lynchings.

The claimed finding that exposure to lynchings of black victims by black subjects causes adverse health outcomes is further undermined by results that are inconsistent with their causal narrative. Specifically, they find that white subjects also have significantly lower cognitive performance in states with above median number of lynchings of black victims in all models except the one that controls for the South that also produces null results for black subjects.

The researchers attempt to explain this unexpected finding by arguing: “it is likely that structural racism against people racialized as Black co-occurred with structural racism-related factors, such as economic underdevelopment, which created a ‘universal harm’ (Brown and Homan, 2024) that also adversely impacted people racialized as White in the U.S. South.” By acknowledging that other factors, such as economic underdevelopment in the South may account for the adverse health outcomes of white subjects, it is unclear why this could also not be the explanation for the negative outcomes for black subjects.

Lastly, the claim that these subjects would have been exposed to lynchings strains credibility given how almost 99 percent of lynchings occurred before the average subject in the study could have been aware of them as a child. The average age for black subjects in the study was 67.68 when baseline data was collected in 2006 or 2008, meaning they were born around 1939. After 1942, when these subjects would have been three years-old, there were only 26 lynchings with black victims recorded nationwide in the Tuskegee Institute data set used by the researchers. That means that over 99 percent of the lynchings with black victims used for the study’s analysis occurred before the average subject could have been aware enough to have “experienced” it.

The only way that black subjects could have been affected is by living in the kind of state that had previously witnessed lynchings of black victims. But just as economic development is a different causal mechanism than the stress of having experienced lynchings, the conditions that facilitated lynchings in certain states before subjects could have been aware of those events are not the same thing as lynchings themselves.

All that this study demonstrates is that black and white subjects in the South have worse health outcomes than subjects in other regions. We have no way of knowing whether lynchings, economic conditions, or other factors in the past or present contributed to these negative health results. Claiming that the stress of experiencing lynchings caused these health problems is without any scientific basis. No credible scientists would use this evidence to make that claim and no credible scientific journal would publish and stand by these results. No credible health agency should be funding this pseudoscience, either.

[i] In another example of sloppiness, the article twice describes the effects as 18.5%: “Black that lived in states with higher proportions of lynchings (in midlife) experienced 18.5% (95% CI 3%, 36%) higher circulating CRP levels than participants racialized as Black that lived in states with lower proportions of lynchings.” But the coefficient in Table 3 is .17 and .185 does not appear anywhere in the tables of results.

Jay Greene, PhD is a senior fellow at Do No Harm.

Evidence Lacking for Claim That the Stress of Racism Shortens Lives

If researchers produced a study finding that poor and minority people tend to be more likely to have health problems and die at a younger age, it probably wouldn’t be published in a leading medical journal or covered with articles in national newspapers. It would rightly be seen as a restatement of the well-known, sad reality that for a variety of reasons poor and minority people tend to have worse diet and exercise and are more likely to use drugs and alcohol, contributing to worse health and earlier death.

But if researchers relabel the problems poor and minority people experience as “cumulative lifespan stress” and suggest those problems are the result of “systemic and explicit discrimination,” those same banal observations can earn a spot in one of the American Medical Association’s top journals and be covered in The Washington Post under the headline: “New evidence shows how discrimination shortens lives in Black communities.”

To be clear, the study published in JAMA Network Open does not demonstrate in any way that discrimination shortens lives in black communities. All it does is show that five measures, which they combine and call “cumulative lifespan stress,” are correlated with indicators of inflammation and are also correlated with dying younger. They also observe that black subjects scored higher on the index they called “stress,” had higher measures of inflammation, and also tended to die at an earlier age. The study’s research design does not allow them to identify whether the five measures they combine and label as “stress” caused inflammation or earlier death, nor can their study exclude whether other factors that they did not examine could have caused both the measures of inflammation and dying at a younger age.

Let’s consider the five measures the researchers use as an index for the physiological stress over one’s life to see how weak the study’s research design is. To capture this cumulative lifespan stress, researchers surveyed study participants to collect information on “(1) childhood maltreatment[…], (2) adult lifetime trauma exposure[…], (3) researcher-verified stressful life events[…], (4) discrimination[…], and (5) indices of socioeconomic status.”

The researchers combine these five measures into a single indicator that they call “cumulative lifespan stress,” but it is far from clear that these five measures actually capture physiological stress. In fact, many of these five measures include information on health problems or factors that could contribute to health problems. For example, the survey used to capture “adult lifetime trauma exposure” includes measures of whether subjects had “experienced a life threatening illness,” “experienced a miscarriage,” and was involved in an accident or otherwise received a serious injury. The measure of “stressful life events” includes information on serious illness or injury and whether a close relative had died.

These health challenges may be stressful, but it would be highly misleading to conclude that the stress associated with serious illnesses caused people to die at a younger age as opposed to the illnesses themselves. The researchers never control for the actual illnesses that subjects have when examining the correlation between their “cumulative lifespan stress” measure and the probability of early death. A subject could have chronic diabetes, uncontrolled blood pressure, or cancer and the researchers would conclude that they died of stress rather than these various diseases.

It is also important to note that only one of the five measures that they claim capture stress includes indicators of discrimination. And that measure asks whether subjects believe they had been treated “unfairly” in employment, housing, or other matters for a variety of reasons, only one of which is race. To conclude that this information, which is part of one of five measures that collectively are associated with early death, means that “discrimination shortens lives” would be completely irresponsible.

The reason this shoddy research receives such favorable treatment by a leading medical journal and alarmist coverage from national newspapers is that people wish to advance a political argument blaming racism for higher rates of health problems and early death in the black community. But nothing in this research demonstrates societal discrimination is to blame. By failing to control for the health challenges associated with diet, exercise, and alcohol and drug use, and by falsely relabeling reports of serious illness or risks of getting serious illnesses as “cumulative lifespan stress,” the study is attributing to racism what could easily be explained by medical comorbidities, individual choices, and community dysfunction.

If you are wondering who is paying for this shoddy research, the answer is you are.

Taxpayers funded this research through grants awarded by the National Institute on Aging, the National Science Foundation, and the National Institute on Alcohol Abuse and Alcoholism. The last source of funding is particularly ironic since the study did not examine the obvious possibility that alcohol abuse could be part of the explanation for the results they observe. It’s bad that the American people must be falsely blamed for causing their black neighbors to die because of stressful discrimination, but even worse that they have to pay for such chicanery. Perhaps paying to be falsely blamed is also dangerously stressful.

Why Medicine Should Avoid the ‘Money-cillin’ Treatment

Two professors at the University of Pennsylvania’s medical school have identified an exciting new treatment to improve health outcomes. It consists of basically giving people money. The marketing department has developed better sounding terms for this treatment, like cash transfers or guaranteed income, but if the marketing folks were really clever, they would be calling it “money-cillin.” That worked for penicillin, right?

Whatever we call it, the idea being advanced by Drs. Aaron Richterman and Harsha Thirumurthy in The Atlantic is that this new treatment is one of the most important interventions in medicine: “when designed with the right basic ingredients, cash transfers are one of the most powerful levers that governments have to alleviate poverty and improve health.” As prominent professors at an Ivy League medical school, we can imagine that Drs. Richterman and Thirumurthy think that “money-cillin” should be studied by leading medical researchers and taught to future doctors. Sure, anatomy and physiology are important for healing patients, but so is welfare policy.

The only problem is that the effects of cash transfers on health outcomes have been rigorously studied and the results have been very disappointing. There was a large-scale experiment in the U.S. in which low-income people were given $1,000 in cash per month for three years and compared to a randomized control group that was given $50 per month. The evaluators, including leading economists from the University of Michigan and the University of California, Berkeley, produced two studies, one describing results on health outcomes and another describing results on labor outcomes. It is worth quoting their findings at length. On health outcomes they found:

Over the three year time horizon that we study, we find no effect of the transfer across several measures of physical health, and we can rule out even very small improvements. The transfer also did not improve mental health after the first year and by year 2 we can again reject very small improvements. We also find precise null effects on self-reported access to health care, physical activity, sleep, and several other measures related to preventive care and health behaviors. We find no effect of the transfer on the health of participants’ children, but we do find that children in treated households were more likely to be up to date on their vaccinations.

Not only did “money-cillin” fail to improve health outcomes, but it actually harmed people’s motivation to work. Again, quoting the results at length, the economists found:

The transfer caused total individual income excluding the transfers to fall by about $1,800/year relative to the control group and a 4.1 percentage point decrease in labor market participation. Participants reduced their work hours as a result of the transfers by 1-2 hours/week and participants’ partners reduced their work hours by a comparable amount. Among other categories of time use, the greatest increase generated by the transfer was in time spent on leisure. Despite asking detailed questions about amenities, we find no impact on quality of employment, and our confidence intervals can rule out even small improvements. Treated participants broadly increase expenditures, led by spending on non-durable goods and services, with smaller increases in spending on durable goods and human capital. We observe no significant effects on degree attainment, though the magnitudes of the estimated effects generally appear larger among younger participants. Measures of subjective well-being are higher among treated participants in the first year of the transfers but then revert to control group levels. Overall, our results suggest a moderate labor supply effect that does not appear offset by other productive activities.

Not surprisingly, giving people cash allows them to work less and spend more on leisure.

In The Atlantic article, Drs. Richterman and Thirumurthy briefly mention that guaranteed income pilot programs in the U.S. “haven’t delivered the dramatic health improvements associated with cash-transfer programs elsewhere.” Without mentioning the large-scale randomized controlled trial described above, they (falsely) assert that these programs have “seemingly produced only modest health gains in the United States.”

But don’t worry. Even when the evidence is against them, these scientists are so smart that they still know they are right. They offer a variety of hypotheses, unsubstantiated by research, to rationalize why money-cillin has not produced the results they expected in the U.S.

First, they falsely describe past efforts as providing too little money to make a difference: “a few hundred dollars a month for a relatively short period of time, typical of guaranteed-income pilots, rarely matches the steep costs of housing, child care, and health care.” The experiment described above provided $1,000 per month for 3 years.

Second, they suggest that cash transfers are insufficient given deeper societal problems: “In the U.S., the dominant health problems are chronic diseases shaped by neighborhood environments, structural inequities in housing and health care, and years of accumulated risk from unhealthy diets and other long-term exposures. These problems are far less responsive to short-term financial boosts. Cash can reduce stress and improve stability, but it cannot, on its own, undo the deep roots of these conditions.”

But if this were true, then it would be a pretty convincing argument against expanding cash transfer programs. Short of revolutionary change, simply providing money would make no difference. Along these same lines, they argue that cash transfers won’t work unless millions get them: “U.S. pilots have been small, reaching only hundreds or thousands of families—too limited to change the broader conditions that shape health outcomes.”

Third, they suggest that cash transfers need to be conditioned on other behaviors to be effective: “cash works best when it is woven into social infrastructure that families already rely on. In many low- and middle-income countries, payments are linked with health visits and other essential services.”

But if that were true, then they are really advocating for traditional welfare programs that place conditions on support rather than arguing for cash transfers or guaranteed income programs. And even though they have no evidence to support this hypothesis, any success of conditional cash transfers would leave it unclear whether the cash or the required behavior was producing any benefits.

After offering all of these rationalizations (and falsehoods), they point to food stamps to prove that expanding cash transfers is essential for improving health outcomes. But they have no experimental evidence—the kind typically required for approval of a new drug—-to support this claim and instead point to observational evidence reliant on complicated and opaque research designs to draw this conclusion. The problem with this kind of evidence is that the results are highly sensitive to researcher choices about research design and model specification, making them very easy to manipulate. When presented with conflicting evidence, we should believe the results of experiments over observational studies.

Arguments like those made by Drs. Richterman and Thirumurthy in this Atlantic article are the equivalent of advocating for communism because it’s never really been tried. It reveals an ideological commitment that is impervious to empirical evidence or scientific examination. Medical research and education should not be wasting time on money-cillin and should instead be focused on scientifically backed medical interventions that doctors can use to help their patients.

Debunking Frakes and Gruber’s New Study on Racial Concordance

The claim that patients have better outcomes when they are treated by a doctor of the same race is the key to efforts to maintain racial preferences in medical education and hiring. However, the evidence does not support the alleged benefits of “racial concordance,” as it is called in the research literature. Systematic reviews of studies on the effects of racial concordance do not show better outcomes for patients treated by doctors of the same race. Even a highly touted study cited by a Supreme Court justice that purported to prove the benefits of racial concordance was later revealed to be marred by researcher misconduct and a Harvard economist was unable to replicate the results when adding an obvious statistical control. Advocates for race-based preferences in medical education and hiring have grown desperate in their efforts to maintain racial preferences in the face of court decisions and legislation that seek to eliminate them. A new study in The Review of Economic Studies offers to rescue them, but, as this analysis will show, it turns out to be no more credible than past claims. There continues to be no scientific basis for maintaining racial preferences in medical education and hiring. Continue reading the full report below.

A Critical Review of “Gender Mix and Team Performance: Evidence from Obstetrics”

A new study posted by the National Bureau of Economic Research (NBER) claims to find that all-female teams of obstetricians cause fewer maternal complications during delivery than all-male or mixed gender teams. The theory is that gender norms make female obstetricians better at communicating and collaborating, which improves health outcomes.

The authors, Ambar La Forgia and Manasvini Singh, draw their conclusion from an analysis of medical records for all births delivered in Florida hospitals between 2006 and 2018. They note that birth records contain the name of the “attending” physician as well as that of the “operating” physician. In 77% of the 2.5 million records they have, the same person is listed as the attending and operating physician. But 23% of the time different names are listed.

La Forgia and Singh focus their analysis on the subset of births with two different doctors listed so that they can compare the rates of maternal complications with different gender combinations of the two doctors. They claim that when both the lead (as they relabel “attending”) and assisting (as they relabel “operating”) are female, there is the lowest rate of maternal complications. Having both the lead and assisting obstetricians as male results in the highest rate of maternal complications. Having a male lead and a female assistant is a little better than having a female lead and a male assistant, but both produce more complications than all female-teams and fewer than all-male teams.

The authors claim that the gender composition of obstetric (OB) teams is “quasi-random,” so the results should be treated as if they were the causal findings of a randomized experiment. To buttress that claim, they observe that the recorded prior health conditions of patients do not differ across different gender compositions. They also point to the quirks of hospital scheduling and the gender of who happens to be on call as approximating a true experiment.

The study looks impressive, with over a half million observations and 74 pages filled with tables and figures to address anticipated objections. And at first blush their findings seem alarming, claiming that “severe maternal complications are 15.8% higher in male-only teams and 7.1 – 10.8% higher in mixed-gender teams compared to female-only teams.” Closer examination reveals that even if their analysis is completely correct, the benefit of all-female teams of obstetricians would only reduce the rate of having at least one maternal complication by 0.18%, from an average of 2.61% for all cases to 2.43% for those served by all-female teams. In addition, their model estimating the effect of the gender composition of obstetric teams controlling for hospital and year fixed effects only explains 0.7% of the variance in outcomes. Since some of that variance is explained by the year or hospital in which delivery took place, gender composition only accounts for a fraction of 0.7% of the variance in maternal complications. Thus, gender composition does not explain at least 99.3% of the variance. These are tiny and weak effects that are only statistically significant because the analysis examined over a half million observations.

But there are serious flaws in their analysis that should make us doubt even these tiny effects. The rest of this review presents concerns with the validity of their analysis and the credibility of their conclusions.

Biased Selection of Cases into the Analysis

The main difficulty with the La Forgia and Singh study revolves around the fact that they are only examining outcomes for the minority subset of cases in which two obstetricians are listed on the birth record. It is clear that whether a birth record has one or two doctors listed is not random. In fact, it is highly likely that whether a delivery has one or two doctors is correlated both with having a female doctor and with the rate of maternal complications, and not due to anything unique in the communication and collaboration among all female teams. Consequently, the authors introduce a significant bias into their analysis. More specifically, there is a subset of cases selected into the study involving two doctors that are more likely to be typical deliveries with low rates of maternal complications if one of those doctors is female.

Let’s walk through each step of this argument. Having two doctors on the birth record in Florida does not necessarily mean that two obstetricians were both present and working together to deliver the baby. The authors acknowledge in some cases two doctors actively collaborate to manage a complicated delivery but in other cases “the Lead physician monitored the patient up until delivery, but the Assisting physician performed the delivery because of a scheduling change or the Lead attending another birth.” While both scenarios involved communication and collaboration, the nature of communication is quite different – one entails a delivery that is likely foreseen to be complicated (and hence the “assist” with both doctor’s present for the delivery) while the other entails a handoff (with only the “assisting” physician present for the delivery).

It is almost certainly the case that female obstetricians are more likely to be involved in the cases with two doctors that result from a shift change as opposed to an anticipated complicated delivery with one doctor supporting another. This is likely because, on average, female obstetricians tend to work fewer hours than their male counterparts, which naturally leads to more patient handoffs to another doctor for delivery.

A survey of 3,698 OBGYN doctors conducted by the American College of Obstetricians and Gynecologists found that female respondents “worked 10% fewer hours, saw 9% fewer patients, and performed 21% fewer procedures.” Another survey of 541 obstetrician-gynecologists similarly found 22.1% of women reported working 60 or more hours per week compared to 31.5% of male OBGYN doctors. Even after adjusting for age differences, this study found that “women worked 4.1 fewer hours per week than men.”

We find confirmation of this pattern within the results that La Forgia and Singh report. In Table A.1 we see that female obstetricians in their data set have 21.6% fewer deliveries per year than do male obstetricians (200.49 versus 157.17). If female obstetricians work shorter hours, on average, then they will more frequently be involved in the hand-off of patients because of shift changes, which results in a higher percentage of cases with female OBs having two doctors listed on the birth record than for male doctors.

It is also important to note that deliveries that take longer, on average, result in a higher rate of maternal complications. As an analysis of over 50,000 deliveries in Scottland found, “As the duration of second stage of labour increased each hour, the risk of obstetric anal sphincter injuries, episiotomies and PPH [postpartum hemorrhage] increases significantly. Women were over 2 times more likely to have a forceps or caesarean birth.” This analysis confirms the common guidance found on medical web sites, like this warning from the Cleveland Clinic about the increased risks of complications: “Prolonged labor increases your chances of needing a different type of delivery. For example, your healthcare provider may need to use medical instruments, like a vacuum or forceps, to help deliver your baby. Prolonged labor also increases your chances of having a C-section.”

The association between length of labor and the rate of complications means that routine deliveries that take less time and result in fewer complications are more likely to be handled by a single doctor if that doctor is male. But those same routine deliveries without complications are more likely to have two doctors listed on the birth record if the obstetrician is female because they tend to work shorter hours and would be more likely to hand-off those deliveries to another doctor because of a shift change.

The net effect of this is more deliveries without complications involving male obstetricians are handled by a single doctor and excluded from La Forgia and Singh’s analysis. At the same time uncomplicated deliveries involving female obstetricians are more likely to be handled by two doctors and therefore included in the La Forgia and Singh analysis. Their data set ends up disproportionately stocked with routine deliveries involving female obstetricians — the exact kind of low-risk cases that are less likely to appear for male physicians because they work longer hours and have relatively fewer handoffs. Rather than finding that teams of obstetricians involving females reduce maternal complications because of their advantages in communication and collaboration, they are really just drawing conclusions from a biased data set.

Despite filling 74 pages with analyses to address anticipated concerns with their analysis, La Forgia and Singh never present the obvious one to address the issues raised here. They never show whether female obstetricians share the billing with another doctor on the birth record at a higher rate than do male obstetricians. That is, they never show whether the data set they examine over-represents routine deliveries for female doctors while under-representing them for male-doctors.

Using the limited information they do report, I am able to estimate the over-representation of deliveries involving female OBs in their data set focused on those involving two doctors. According to Table A.1, the 1,010 female OBs included in the study are involved in 369,818 deliveries or about 28.17 each per year over the 13 years included in the study. Table A.1 also discloses the average number of births per year for doctors involved in each gender combination. Taking a weighted average for female doctors, yields 157.17 deliveries per year. Dividing 28.17 by 157.17 we see that about 17.9% of the cases involving female OBs have two doctors listed on the birth record. Making the same calculations for male OBs reveals that about 16.2% of their cases have two doctors listed on the birth record. Female OBs are over-represented in the data set La Forgia and Singh examine by about 9.3 percent.

Even modest biases in the data set can yield the results they report. The advantage of all female teams they claim is quite small. According to the “main specification” in Table 3, having an all-female team reduces the rate at which there will be at least one maternal complication by 0.18%. This tiny effect is only statistically significant because they have over half a million cases in their analysis. It is also worth noting that the gender composition of the obstetric team explains less than 0.7% of the variance in outcomes. They have a weak finding easily distorted by the over-representation of routine deliveries involving female obstetricians that are included in their analysis.

Incomplete Medical Record When Controlling for Risk Factors

Another concern is whether all-female teams of obstetricians tend to have patients who are at lower risk of developing complications. It is true that they have medical records that could contain information on “23 patient risk factors that are predictive of maternal complications.” And when they use that information to model the expected rate of complications, they find no difference across the four possible gender compositions of the doctor-pairs that they examine (See Figure 1, Panel A).

But as they acknowledge, “variations in coding behavior by physicians… could influence our ability to accurately assess the patient’s clinical risk.” They attempt to address this data limitation by also controlling for “the number of diagnosis codes recorded in the patient’s medical record” but that doesn’t really resolve the problem. Controlling for the number of diagnosis codes might correct for the fact that doctors inconsistently code problems that are observed, but it does not correct for problems that are unobserved and therefore entirely absent from the medical record.

If patients have not been receiving regular medical care prior to delivery, the doctors may only record the risk factors that the patient discloses at the time of admission. For example, the 23 risk factors include things like “substance abuse or smoking,” “known or suspected fetal abnormalities,” “hypertension,” “spotting complicating pregnancy,” and “excessive weight gain during pregnancy.” These are the kinds of risk factors that may not be entered into the medical record if the patient has not had these issues identified during prenatal care or fails to disclose them when admitted for delivery. The absence of information would be treated the same as not having that risk factor.

According to the March of Dimes, 17.3% of expecting women only begin to receive prenatal care in the second trimester and 7.3% first receive it later or not at all. These women are less likely to have complete medical records that would fully capture their risk factors. These same women are less likely to be active choosers of the obstetric practice they would prefer for delivery.

There are also all-female OBGYN practices that recruit patients by extolling the benefits of having female doctors. For example, a fairly typical practice in Massachusetts emphasizes on its web site: “Women-led OB/GYN practices like Essex County OB/GYN prioritize patient-centered care, focusing on the unique needs, preferences, and lifestyles of each individual. This approach means your voice matters in every decision made, and your care is tailored to suit your specific circumstances.”

More advantaged women with more complete medical records are more likely to be drawn to all-female practices like this even if only because they are more likely to be active choosers of who they want to help with their delivery. Disadvantaged women who lack prenatal care and have incomplete medical records are more likely to be assigned to whoever is on call at the hospital, including more male obstetricians. The difference in the rate of maternal complications can be driven by unobserved differences in risk factors rather than superior communication and collaboration skills among female obstetricians.

Do All-Female Teams Have Intersectional Benefits?

Further evidence to support this suspicion can be found in La Forgia and Singh’s claim that “female-only teams not only achieve the lowest complication rates for Black women, but are also the only team type to have no racial disparity in maternal outcomes.” That is, La Forgia and Singh claim that gender combinations of teams other than all-female ones not only produce higher average rates of maternal complications, but they do even worse with their black patients than with their white ones.

They have no theory to explain why female teams of obstetricians would have particular advantages with respect to black patients. The theory they are advancing is that female teams are better at collaboration and communication, but it is unclear why this should be more important in preventing complications among black patients than among non-black patients other than as a result of being better positioned to manage the additional risk factors that black patients, on average, may bring to delivery.

The patient’s race, like the presence of hypertension, fetal abnormalities, or substance abuse, is among the pre-existing qualities that women bring to delivery. La Forgia and Singh claim to have controlled for all relevant pre-existing issues to make their claim that the communication and collaboration advantages of female teams during labor are what cause lower rates of complications.

Rather than demonstrating an additional benefit of all-female teams, their finding of differential outcomes for black patients suggests that their analysis suffers from incomplete information about risk factors that are more often found among black women and are also correlated with whether patients have all-female teams. Rather than being evidence of a theoretically unlikely intersectional benefit, this result is evidence of serious omitted variable bias.

What About the Baby?

Even if the La Forgia and Singh study were correct in finding a lower rate of maternal complications for all-female doctor teams, their analysis would still be incomplete. There are often trade-offs between risking maternal complications and preventing serious complications for the newborn baby. When delivery stalls and the baby shows signs of distress, doctors have to consider interventions including C-sections that increase the rate of complications for the mother but may save the baby from worse complications. Recognizing that C-section may be over-used does not erase the existence of these trade-offs nor does it mean that the optimal outcome for both mother and baby is the one that minimizes risks for the mother without considering those of the baby.

The Florida data set that La Forgia and Singh are using includes information about outcomes for the baby, but they do not report any analyses of those outcomes. Without considering how all-female teams may affect complications for the babies they deliver, it would be at least premature to conclude that they produce superior outcomes because of better communication and collaboration.

Conclusion

We have no evidence that La Forgia and Singh intended to produce a biased analysis or mislead their readers, but that is what they have in effect done. They have limited their analysis to a minority subset of cases in which two doctor names are listed on birth records without considering how those cases may be selected in a way that biases their result. They have controlled for observed medical conditions without considering how incomplete medical records likely bias their analysis. They have focused on maternal complications related to delivery without considering the possible tradeoffs between those complications and outcomes for the baby.

While these defects in their study may not be intentional, they do not occur at random. Researchers are tempted to overlook or downplay problems when they have strong ideological priors about the results they expect to find. The danger of bias resulting from ideological blinkers has historically been held in check by the enforcement of academic standards and the general adoption of skeptical norms. The process of peer review is further meant to hold ideological priors in check.

Unfortunately, NBER papers, despite being broadly respected and influential, do not go through any peer review process before being posted. And the check of peer review has weakened as more reviewers and editors share the same ideological preferences. The forum of “Researcher, Heal Thyself” is meant to address these shortcomings by offering critical reviews even if the researchers and their peers are unable to identify concerns on their own.

Debunking Frakes and Gruber’s New Study on Racial Concordance

The rise of the child transgender industry over the past decade has relied, in no small part, on the financial incentives for physicians and hospitals to perform sex-denying medical interventions. These procedures offer a potentially lucrative revenue source.

But insurance coverage has been variable for medical procedures performed for the purpose of so-called “gender-affirming care” for minors. Further, dozens of states now restrict Medicaid funding for sex-denying interventions for minors or restrict minors’ access to the procedures themselves.

This report aims to show the avenues through which healthcare providers may be able to skirt official coding guidelines to secure insurance reimbursement for so-called “gender-affirming care.” By misrepresenting the medical procedures they are performing, providers can pass off transgender medicalization as, for example, routine endocrine care unrelated to pediatric medical transition. These “loopholes” may enable providers to get paid for procedures which otherwise may not be funded. In some cases, such practices may even be outright fraudulent or a means of evading state-level restrictions on child sex change interventions.

The Ideological Capture of Continuing Medical Education

The rise of the child transgender industry over the past decade has relied, in no small part, on the financial incentives for physicians and hospitals to perform sex-denying medical interventions. These procedures offer a potentially lucrative revenue source.

But insurance coverage has been variable for medical procedures performed for the purpose of so-called “gender-affirming care” for minors. Further, dozens of states now restrict Medicaid funding for sex-denying interventions for minors or restrict minors’ access to the procedures themselves.

This report aims to show the avenues through which healthcare providers may be able to skirt official coding guidelines to secure insurance reimbursement for so-called “gender-affirming care.” By misrepresenting the medical procedures they are performing, providers can pass off transgender medicalization as, for example, routine endocrine care unrelated to pediatric medical transition. These “loopholes” may enable providers to get paid for procedures which otherwise may not be funded. In some cases, such practices may even be outright fraudulent or a means of evading state-level restrictions on child sex change interventions.

Activism, Not Advocacy: The Radical Transformation of the American Nurses Association

The rise of the child transgender industry over the past decade has relied, in no small part, on the financial incentives for physicians and hospitals to perform sex-denying medical interventions. These procedures offer a potentially lucrative revenue source.

But insurance coverage has been variable for medical procedures performed for the purpose of so-called “gender-affirming care” for minors. Further, dozens of states now restrict Medicaid funding for sex-denying interventions for minors or restrict minors’ access to the procedures themselves.

This report aims to show the avenues through which healthcare providers may be able to skirt official coding guidelines to secure insurance reimbursement for so-called “gender-affirming care.” By misrepresenting the medical procedures they are performing, providers can pass off transgender medicalization as, for example, routine endocrine care unrelated to pediatric medical transition. These “loopholes” may enable providers to get paid for procedures which otherwise may not be funded. In some cases, such practices may even be outright fraudulent or a means of evading state-level restrictions on child sex change interventions.

How the AAMC Fails to Read and Correctly Interpret the Research It Cites

The rise of the child transgender industry over the past decade has relied, in no small part, on the financial incentives for physicians and hospitals to perform sex-denying medical interventions. These procedures offer a potentially lucrative revenue source.

But insurance coverage has been variable for medical procedures performed for the purpose of so-called “gender-affirming care” for minors. Further, dozens of states now restrict Medicaid funding for sex-denying interventions for minors or restrict minors’ access to the procedures themselves.

This report aims to show the avenues through which healthcare providers may be able to skirt official coding guidelines to secure insurance reimbursement for so-called “gender-affirming care.” By misrepresenting the medical procedures they are performing, providers can pass off transgender medicalization as, for example, routine endocrine care unrelated to pediatric medical transition. These “loopholes” may enable providers to get paid for procedures which otherwise may not be funded. In some cases, such practices may even be outright fraudulent or a means of evading state-level restrictions on child sex change interventions.

Debunking the Utah Department of Health and Human Services’ Defense of Pediatric Medical Transition

The rise of the child transgender industry over the past decade has relied, in no small part, on the financial incentives for physicians and hospitals to perform sex-denying medical interventions. These procedures offer a potentially lucrative revenue source.

But insurance coverage has been variable for medical procedures performed for the purpose of so-called “gender-affirming care” for minors. Further, dozens of states now restrict Medicaid funding for sex-denying interventions for minors or restrict minors’ access to the procedures themselves.

This report aims to show the avenues through which healthcare providers may be able to skirt official coding guidelines to secure insurance reimbursement for so-called “gender-affirming care.” By misrepresenting the medical procedures they are performing, providers can pass off transgender medicalization as, for example, routine endocrine care unrelated to pediatric medical transition. These “loopholes” may enable providers to get paid for procedures which otherwise may not be funded. In some cases, such practices may even be outright fraudulent or a means of evading state-level restrictions on child sex change interventions.