Pointing to this recent article in Physical Review Physics Education Research, Michael Weissman flags this passage:
In our initial analysis of the historical data, we [the authors of the above-linked article] noticed there are cases where applicants have similar physics GRE scores and GPA, yet one applicant is accepted while the other is not. Given that cases such as these might add challenges to modeling the data, removing such applicants might allow us to better characterize the general trends in the data. We, therefore, consider an alternative approach that detects similar applicants with different admission outcomes and removes them from the database.
Whaaaaa?
The last 2 sentences of the abstract of that article are just amazing:
Our inability to model the second dataset despite being able to model the first combined with model comparison analyses suggests that rubric-based admission does change the underlying process. These results suggest that rubric-based holistic review is a method that could make the graduate admission process in physics more equitable.
I pointed this one out to Jamie Robins, who wrote:
Almost as equitable as a coin toss. With more grant money and research they may get there yet.
Jamie then elaborated:
Of course I was being overly pithy. If many more advantaged than disadvantaged apply, a coin toss will not meet their goals. So they will require a rubric based selection for which the ML program does better than chance (unless both the the covariates used by the rubric and their correlates are withheld from the features given to the ML program).
Miguel Hernan picked up on one other thing:
Also from the abstract:
Yet, no studies have examined whether rubric-based admissions methods represent a fundamental change in the admissions process or simply represent a new tool that achieves the same outcome.
But I couldn’t find a mention to MIT’s experience as summarized here by its Dean of Admissions:
Standardized tests also help us identify academically prepared, socioeconomically disadvantaged students who could not otherwise demonstrate readiness because they do not attend schools that offer advanced coursework, cannot afford expensive enrichment opportunities, cannot expect lengthy letters of recommendation from their overburdened teachers, or are otherwise hampered by educational inequalities. By using the tests as a tool in the service of our mission, we have helped improve the diversity of our undergraduate population while student academic outcomes at MIT have gotten better, too; our strategic and purposeful use of testing has been crucial to doing both simultaneously.
I also enjoyed (that is, was upset by) this bit:
While we were able to develop a sufficiently good model whose results we could trust for the data before the implementation of the rubric, we were unable to do so for the data collected after the implementation of the rubric, despite multiple modifications to the algorithms and data such as implementing Tomek Links.
“Despite multiple modifications to the algorithms and data,” huh? Maybe they just weren’t trying hard enough. I don’t know what’s worse, someone who keeps altering his algorithms and data until he finds success, or someone who tries to do it but can’t succeed.
We laugh because it’s too painful to always be crying.
P.S. I can’t speak to the substance of the subject under discussion. Maybe rubric-based holistic review is a great thing. Just cos a bad paper has been published promoting it, that doesn’t mean it’s a bad idea. Just one thing: I hate the use letters of recommendation for admissions and hiring.
+1 for hating letters of recommendation
In the “after” data, they have a classifier with AUC of 0.626 ±0.006 “which is less than the minimum of 0.7 for a reasonable model”. Perhaps a good example of questionable dichotomous use of AUC.
They have some nice plots in their appendix (Figures 5 to 7). I would have dropped the boxplot aspect from them, but the plots show the comparison between before and after the policy implementation quite nicely.
Admittedly, I have no idea whether it makes sense to compare the sub-groups they compare with each other. But if it does, I think someone literate in the field of admissions might benefit from these plots.
I wrote a short Comment on this paper. It’s here: https://arxiv.org/pdf/2306.09875.pdf.
The key points are:
“1. Although the paper only tentatively concludes: “Overall, the results of this initial investigation are suggestive that our admission process did change…”, simple statistical tests show conclusively that the rubric system specifically succeeded in substantially increasing the fraction of low-scoring applicants who were admitted.
2. Although the paper argues, based on prior literature, that this change will improve
graduate outcomes, a more careful reading of that literature, including papers they cite, suggests the opposite.
3. A data-editing technique that is advocated and used in some of the analyses is likely to give biased results and thus should generally be avoided.”
The journal has asked me to revise substantially it to be more complimentary to the original authors but so far I’ve been too preoccupied with other things.
entirely off-topic p.s. on preoccupation:
https://open.substack.com/pub/michaelweissman/p/an-inconvenient-probability-v30?r=2byn6&utm_campaign=post&utm_medium=email