Colin Leslie writes:
I am a relatively new reader of your blog (and somewhat new to stats), having come across your work randomly—which is fitting I suppose. I’m not an expert on stats, but I do have a sort of insider-outsider thing going on because of the quantitative requirements of my PhD (just defended a dissertation at USC in Public Policy).
Reading your latest blog entry on stereotype susceptibility reminded me of an idea I got a while ago, and I thought it ought to be floated amongst people who care. I came to public policy from the legal field, and in that field there are private services (Westlaw and Lexis being the biggest and most well known) that flag case law based on a series of dimensions. For instance, is the holding of the case still good law? Has it been overruled, or been overruled in part? Is it uncertain about whether this case can be used for such and such proposition? The system is obviously very useful for practitioners.
You can probably see where I am going, and someone has probably proposed this already. As I go through Google Scholar or lists of references I never quite have an idea of where the errors are and am left mostly on my own, which is fine but also noisy because I’m just one person (and somewhat a noob). Some errors are so critical that they destroy the paper and line of research but might be hiding somewhere non-obvious. Some are just a question of assumptions that no one has spelled out fully. It would be nice if these researcher degrees of freedom issues were flagged in a more obvious way. Maybe this sort of thing could be crowdsourced a la Wikipedia. Obviously, scientific research might have more dimensions than case law, but I think the model still definitely holds. Do you think this could work?
I saw the Pubpeer thing you blogged about and it seems like others have been moving a little in this direction. I haven’t done much searching around on this myself, just had passing thoughts and thought I’d put it out there. I suppose it’s a surfeit of knowledge issue, but it bothers me a little that there’s often no easy place to follow up on research that is unambiguously DOA or retracted, or for that matter no place to see when research is unambiguously a progression. I’m sure the middle ranges are more controversial and harder to apply in ways that make people happy, but it would be nice to have some warning flags at least for the clear cases.
I don’t know. One challenge here is that there are lots more science papers than there are laws and court decisions. Although I guess you could just focus on the material from major medical journals: NEJM, JAMA, etc. There’d be no real reason to have a “Westlaw” for Psychological Science because nobody trusts anything in that journal anyway, right?
I think Elicit (the GPT-powered literature review website) tries to list possible critiques of a paper, which as far as I can tell works by running a sentiment analysis on the how the paper is cited.
Worth a test?
Retraction Watch has a database of retracted publications. It is publicly-funded/croudsourced and incomplete but has a lot of information. https://retractionwatch.com/retraction-watch-database-user-guide/
I don’t see this as solving anything. We already have peer review (the problem with peer review is the peers), retraction watch, post-publication review, PubPeer, etc. Calling for yet another layer of evaluation begs the real problem which is that the current system does not work. If reviewers and editors don’t monitor quality sufficiently, then I don’t see how yet more layers of faulty evaluation will work better. For that matter, the idea of legal review evaluating how “good” a legal decision is, seems to imply that problems with the legal system (the problem with judgements is the judges?). That is, all of these evaluations are flawed, and I would say increasingly so (in the case of legal decisions, the increasing role of politics has made those decisions seem a lot like peer reviewed publication to me).
The problem with any layer(s) of review are the reviewers.
I don’t have a solution, but I maintain that until people’s reputations actually are impacted by doing shoddy (or worse) work, no system will sufficiently evaluate quality of statistical work, legal decisions, or anything else. And, the impact of shoddy work on reputations appears to be declining, as many of the cases highlighted on this blog show. It takes an awful lot of bad behavior before anything happens to the perpetrators (e.g., Cornell and Wansink).
The current systems are woefully inadequate, beyond what is possible. There is no point in saying that reviewers aren’t doing enough when existing systems don’t even adequately track or highlight retractions, so people go on blithely citing papers whose author has gone to jail for federal fraud over said papers.
This is just another example of how the problems are vicious because they are incentive problems. There is no actual incentive to avoid citing retracted papers which support your point; no one is going to call you on it, and if they somehow did get a letter to the editor published a decade later mentioning it, you simply go ‘oops, well, all the other studies support me, so nbd’. But if all authors lost $10,000 of funding within a few weeks of citing a retracted paper, I bet all the citation managers and journals would be keenly interested in better tracking of retractions. I’m reminded of a postmortem from a startup a year or two ago intending to do, essentially, ‘Westlaw for biological science’, where the upshot was ‘everyone thinks the tools are awful, everyone is right, but they don’t care, and won’t pay a single dime, and this is why every attempt is doomed to fail’. For perspective, the simplest and most limited possible Westlaw subscription would be something like $100/month/person, and can easily go to $1000+. What would convince you to pay $1000+/month for a souped-up Google Scholar? The answer turns out to be ‘nothing’.
Why does Westlaw for law work? (At least in the sense of rational lawyers all choosing to pay for such subscriptions.) Because incentives matter enormously there, and citing a ‘retracted’ opinion (which had been overruled) will lead to your opponent or the judge pouncing on you, and lose you your case or lawsuit in a relatively quick and obvious fashion, and cost you potentially a lot of money; and done more than a few times, will end your career. This is why Westlaw (and its rivals) is a sine qua non of law firms. (Similarly, Bloomberg terminals in finance.) It’s an awkward truth that in law, where everything is made up, being wrong has many large short-term consequences, and in science and research in general, where everything is real, being wrong has few small long-term consequences. As long as you can cite blithely retracted papers, never mind ones that have failed to replicate or are just sus etc, how can anyone justify $100-$1000/month?
You scientists /statisticians don’t know how good you have it. I have worked both as a lawyer and in a lab. No comparison. There is some sloppy work and some fraud in labs, sure, and I have seen it. But nothing like the law. In addition there is the structure of the law and its practice: Juries are required to know nothing relevant. Lawyers are people whose education basically stiopped at the undergrad level. State judges are elected without any knowledge of their qualifications on the part of the electorate. Once they are on the bench, their discretion is almost unlimited, and most lawyers are terrified of them. The grievance boards are a joke. It is true that there are these appellate cases that the commentators here focus on, but think about it for a moment: by the time you are in the appellate courts, the facts are so stylized that the original words and acts of the litigants have gone through the wash of a trial, an appeal, and perhaps another appeal, and bear little relationship to anything that might without hilarity be referred to as real.
It looks like there is a lot of case law being generated in the US (see this report from the federal government broken down by federal court, and that’s just the tip of the iceberg). Westlaw covers not only federal law, but also international laws (which are often duplicated in multiple languages). It looks like there are only a couple million academic papers published per year.
I assume that there’s not equal coverage of every case in Westlaw, nor would there have to be equal coverage for every academic paper. If we were to work by citations, we could trim the couple million number down to something less than half that size at the largest.
Yes, plus state courts and admin tribunals generate tons and tons of case law. With legal cases I like having the big red stop sign telling me not to cite a legal standard that got explicitly overruled or changed by statute. With scholarly cites I want it to show me at least the errata and fraudsters. I assume these are not huge pools. You could also start putting in more nuance for research that is contradicted or, say, highly doubtful, and those pools I assume are bigger and more complicated to generate.
The difference is not just quantity. Westlaw Nexis/Lexis have the huge advantage of a hierarchy of actual decisionmaking. A judge can find X and then be overturned by a higher court, at which point Not-X is objectively true in that particular jurisdiction until overturned by an equal or higher court. These compendia allow one to trace this, both intra- and inter-jurisdictionally.
Except for the relatively rare case of retractions, no such definitive refutation is ever possible in academia. While academics can continue to stamd behind the results of their paper against attacks that cause everyone but them to give up, you’d never have any way of knowing that. On the other hand, a lawyer can be convinced he got robbed at the Supreme Court, but literally no one cares until the Court overrules itself.