Our proposal for scheduled post-publication review: “Even if each review took twice the effort of the average pre-publication review, our system would add only 1 percent to the total reviewing effort, while providing important perspectives on papers representing more than one-quarter of the citations received by these influential journals.”

The current system of scholarly journal review is absolutely nuts. The vast majority of review effort goes to papers that nobody reads. We can do better via scheduled post-publication review, for example every time a paper reaches its 250th citation, the journal commissions an outside review–not with the goal of retracting the paper, but to provide a new perspective. If the paper has been cited 250 times, it’s worth getting that new take–and it’s rare enough that a published article reached that level of citation, so the total cost of these new reviews would be low, only a small fraction of the cost of existing pre-publication review.

Andy King and I present the idea:

Problems with the credibility of empirical research have been discussed for decades.

The most common prescriptions for improving credibility — better review, public critique, and replication — have merit, but they fail to direct scarce resources where they can do the greatest good: those few publications with the greatest impact.

We propose an alternative, using review resources more efficiently and effectively by borrowing the idea of “replay review” from professional sports. The current peer-review system would continue to judge research articles in real time when they are submitted, but the publications that go on to have an outsized impact would be evaluated again, and in more detail, to confirm or refine the initial assessment.

Here’s the key insight:

All proposals for strengthening primary review face an economic challenge: Most of the resources spent on strengthening it are wasted because, for most submissions, the existing review process is already strong enough. At top journals, 90 percent or more of the submissions are rejected for apparent flaws, and thus, a strengthened review process will not change the outcome. Of the 10 percent that are accepted and published, most are lightly read and cited and thus have little influence.

And this is what we recommend:

Once a publication receives a specified number of citations, it would receive an independent review. These reviews would then be published in full, along with author responses, so that readers have additional guidance on how to interpret the initial publication. . . .

We crunched some numbers:

To assess the practicality of our proposal, we evaluated the submission and citation history of articles published in the 2014 cohort of the peer-reviewed empirical journals selected by the Financial Times for determining the research rank of business schools.

Skewness in citation rates means that a large proportion of the citation impact can be checked at a relatively low cost. Replay review of just those articles receiving more than 250 citations would mean that publications accounting for 28 percent of all the citations would be checked through further review. Even if each review took twice the effort of the average pre-publication review, our system would add only 1 percent to the total reviewing effort, while providing important perspectives on papers representing more than one-quarter of the citations received by these influential journals.

This is an elaboration of my efficiency argument for post-publication review.

P.S. The Chronicle of Higher Education gave our article the title, “Social Science Is Broken. Here’s How to Fix It.” We didn’t choose that title! Our recommendation was, “Social Science Needs Replay Review.” But those of you involved in the news media know that the authors of an article are rarely are in charge of the title. We like our suggestion of post-publication review, but we’re under no illusion that it would “fix” social science. It’s just one small part of the picture.

P.P.S. We also thank David Wescott at the Chronicle for editing our article.

43 thoughts on “Our proposal for scheduled post-publication review: “Even if each review took twice the effort of the average pre-publication review, our system would add only 1 percent to the total reviewing effort, while providing important perspectives on papers representing more than one-quarter of the citations received by these influential journals.”

  1. I like this idea a lot. One potential issue might be how “citation” is defined. In a world of AI/LLMs, citations are ever easier to create, so it would require some mechanism for distinguishing between “genuine” and “invented” citations. Even if it is required that the citations appear in peer-reviewed publications, we have the issue of citation padding (either by the authors themselves or by their students or colleagues). I suppose there is a question about whether someone would prefer to have post-publication review or not: perhaps sketchy researchers would like to avoid that. But that would be an added benefit by incentivizing bad researchers to avoid too much recognition. In any case, I think there is an issue of strategic behavior of authors that might complicate this otherwise great idea.

  2. Great article and proposal! I suggest that you take into considerations other consideration of impact, e.g. if the author testifies in Congress, etc. I also suggest you focus on the most egregious offending and mendacious disciplines like sociology.

  3. Who is this intended for? At least in the biological/biomed sciences, once an article reached 250 citations (or 100 or 50 or whatever) its value is likely to have been embedded within the community of its particular scientific discipline, and apart from poor articles that receive high citation numbers because the paper is cited as an example of something poor, a re-review is unlikely to be of much interest for the particular community. So are these re-reviews intended for a non-expert general readership?

    And what happens if the paper takes 10-20 years to achieve its 250 citations? Is anyone going to care very much about a re-review? More interesting are the commentaries (Nature “News and Views”, or Science “Perspectives” type) that accompany potentially impactful papers that give the general reader a basic heads-up about the paper and its (potential) significance. Otherwise highly cited papers and their impacts are likely to be written about anyway whenever the particular field is reviewed. 250 citations seems an awfully large number – some very impactful papers in relatively restricted fields may only achieve many fewer citations.

    It’s also not clear how the re-review would be done without taking into account the 250 citations the paper has accrued. Is the re-review meant to highlight good points, flaws, suggestions for changes, clarifications and corrections as in a primary review? It’s difficult to see how a re-review of a highly cited paper would be anything other than a re-endorsement of the status already afforded by a large number of citations – a sort of Matthew effect whereby focus is directed at work that has already achieved “high status”.

    To me post-publication “peer review” of the sort exemplified by PubPeer is a pretty good way of identifying flaws in published work. So even if a highly cited paper has a significant flaw (e.g. faked data), this is unlikely to be identified in a post-250-citation re-review if it hadn’t already been identified e.g. via PubPeer posts.

    I guess I may be being misled by the use of the term “review. What is being suggested in the top article isn’t really a review in the sense of a critical peer review that assesses whether the work achieves some bar for publication, but would be more of a commentary on a highly cited paper and its impact – is that the idea? It would be adding a little more work for reviewers and (highly cited) authors over and above the normal publication peer review.

    • All good points. One quick datapoint: I got my Ph.D. in 2002. I have only *four* papers with > 250 citations, and these were published in 2001, 2003, 2006, and 2012. (Lest the reader think I’m inept: By most measures, I’m a successful physicist.) Except perhaps the 2012 paper (my favorite), I can’t imagine that anyone would care about a “re-review” of these articles. Describing them along with other papers and the state of the field in a review article: sure.

    • Chris:

      1. The idea is the journal that published the paper would commission and publish the review, so that when people look at the paper, they’d also see the post-publication review. If that review revealed important problems with the paper, readers who went to the journal website would see it right away. If the review does not find any problems, that’s fine–it’s good for people to see this too.

      2. As we say in our article, the appropriate threshold would depend on the field and on the journal. We’re not suggesting 250 as a general number; it’s just a starting point based on a quick analysis of some journals in the field of business management.

      3. You ask, “the re-review meant to highlight good points, flaws, suggestions for changes, clarifications and corrections as in a primary review.” I’d say, Yes, pretty much. As with pre-publication review, the reviewer would have discretion. As an author and reader of many many pre-publication reviews, I know there’s a lot of variation. That’s fine.

      4. You write, “apart from poor articles that receive high citation numbers because the paper is cited as an example of something poor, a re-review is unlikely to be of much interest for the particular community.” I disagree! I think a re-review can provide valuable perspective to the reader, in the same way that a pre-publication review is of value to the editor and author.

      To put it another way, post-publication reviews will not be perfect. Neither are pre-publication review. Almost all of pre-publication review is a waste of time, as almost all of the pre-publication reviews are for papers that just about nobody will read. I think it makes sense to devote much more of this effort to papers that are actually being used. In our article we talk about post-publication review expending 1% of total review effort. If it were up to me, post-publication review would take up about 50% of total review effort. I just thought that would be too radical a proposal. Our proposal described above is modest, and I think it goes in the right direction.

        • Claire:

          I don’t think our proposal would remove the problem of citation of problematic work, but I hope it would be a step in the right direction. It’s possible to make a difference even without solving a problem entirely or even mostly.

  4. I think it was Andrew who wrote something like “peer review is just looking at symbols and images on a page”. It was a few years ago, but I’d be glad to take credit.

    Point being, thats no replacement for even someone else reanalyzing the data. And especially not recollecting it.

    So why not just stop with all the half-hearted shortcuts and build independent replications into the funding process?

  5. All, I think we (I’m Andy King) meant the 250 cutoff as an example. Depending on the discipline or journal it could be lower or higher. There could be a statute of limitations.

    In terms of replications, each one can be a massive undertaking, and they are very difficult to publish. In my experience, replication is often beside the point, because the original method or data have flaws. So the replication becomes a scaflold for a re-review.

    Corrigendum might work, but journals are extremely reluctant to publish them. See my efforts to get a typo fixed in the most cited article (since 2014) in Management Science.

    Readers cannot adjudicate each publication, and so many are led astray by flawed work. We need a way to guide readers on how to interpret the published record — thus replay review.

    Or that is my thinking, anway. Andrew may differ.

    • One question is where these re-reviews would be published. It sounds like Management Science would not be overly keen on publishing a re-review of the highly-cited paper you were unable to correct! So would these re-reviews be compiled elsewhere? Perhaps they could share a platform with something like Pubpeer.

      I guess another issue I have is who would the reviewer(s) be – and who would invite the reviews? The receiving journal editor does this now but it sounds like your plan would involve an independent platform that wouldn’t necessarily involve journals?? Since you refer to the possibility of “flawed” work, that suggests the likelihood of a critical review which might be problematic not least because the authors are likely to be reluctant to engage with the re-review process in your scenario.

      There is very probably a field-specific aspect to this. For example, on replications, in the physical/biophysical/mol biol fields there is lots of replication but much of this is “indirect” in that replications are not generally done for their own sake. However, if work is considered important it is replicated down the line and so papers tend to “live or die” impact-wise (and accrue citations) by a sort of natural selection. For example, during covid structures of viral proteins were published followed by (or concurrently with) other structure determinations, and so if someone publishes the structure of the viral protease (say), that was followed by papers describing the structure in the presence of a drug, or with a mutation to assess mechanisms and so on. Lots of replication is of that sort.

    • replication is often beside the point, because the original method or data have flaws. So the replication becomes a scaflold for a re-review.

      If its worth funding once, its worth replicating.

      Cutting the original research in half, and spending it on replicating the most promising 50% sounds great.

      Seems like you believe less than half are even worth trying though.

  6. I’m curious: in social science, what’s the average age at which a paper that gets 250 citations reaches that milestone?

    More generally, it would be interesting to see a graph of age at which n citations is reached vs. n; call this A(n). Note that this is calculated only from the papers that actually reach n citations. Given that re-review is most useful within some cutoff age Y, what’s the largest n such that A(n) < Y? My guess is that for Y being 3 years, this n* will be about 10, at most. Then one can ask what fraction of overall papers have n* citations, and see what this implies for the re-review workload.

    I have other questions / concerns about all this, but this will do for now! (Interesting topic, by the way!)

    • Raghu:

      That would be worth looking at. I can do a quick check by going on Google scholar and looking at the subset of my own papers that have between 200 and 300 citations . . . their publication dates range from 1991 to 2024. About half were published since 2010, so that would suggest a median time until 250 citations of approximately 15 years. I think this analysis is subject to selection bias, though.

  7. When it comes to citations, I keep thinking about Einstein in 1905 and his famous (only!) four papers of that year which is now referred to as “Annus Mirabilis”

    https://en.wikipedia.org/wiki/Annus_mirabilis_papers

    Unfortunately, he kept on going, and as far as I can tell, he did tail off. In fact, shockingly so that his writings outside of physics in the early 1920s are embarrassing.

    https://www.bbc.com/news/science-environment-44472277

    • I think his regrettable diary entries of the 1920s are evidence that being brilliant at one thing (e.g. physics) doesn’t mean you can’t be terrible wrong in other areas (e.g. race). However, as the article you linked says, he later became strongly anti-racist (effectively recanting those old views), so I would argue it’s a *good* thing he kept going and grew as a human in the 30s and 40s.

    • In the May 15, 1935 issue of Physical Review Albert Einstein co-authored a paper with his two postdoctoral research associates at the Institute for Advanced Study, Boris Podolsky and Nathan Rosen. The article was entitled “Can Quantum Mechanical Description of Physical Reality Be Considered Complete?” (Einstein et al. 1935). Generally referred to as “EPR”, this paper quickly became a centerpiece in debates over the interpretation of quantum theory, debates that continue today. Ranked by impact, EPR is among the top ten of all papers ever published in Physical Review journals. Due to its role in the development of quantum information theory, it is also near the top in their list of currently “hot“ papers.

      https://plato.stanford.edu/entries/qt-epr/

      That doesn’t seem like “trailing off”.

  8. It seems some people may be missing the forest for the trees here. My reading of the proposal (at least what I can glean from this post, as I would prefer not to make an account to read the article) is:

    1. There is an enormous quantity of published research
    2. A significant proportion of published research lacks credibility/reliability (ranging from poor methodology to outright fraud)
    3. It is expensive (money, time, etc.) to effectively validate the claims of published research, e.g. through more thorough pre-publication review, reanalysis of data, or full independent replication
    4. A small fraction of published research actually attains the status of “influential” in some sense of the word
    5. Of the negative effects that poor research has on science/society, it is mostly at the confluence of bad + influential
    6. Points 1-3 imply that it is infeasible to apply these credibility-raising practices at scale to all research (and attempts to do so have not been terribly effective)
    7. Points 3-5 imply that the negative impact of poor research can be cost effectively mitigated by ensuring influential (or soon-to-be influential) research is carefully validated

    So the general structure of the proposal is to preferentially apply some credibility enhancing process, in this case “Replay Review”, to research that exceeds some influence threshold (as an example, citations >= 250 is suggested).

    To my eye, the overall proposal is sensible. Various fields could define their own measures for an ‘influence threshold’, so focusing on the 250 number seems beside the point. It can reasonably be debated which credibility enhancing process is appropriate, but replay review seems like a reasonable start.

    I would even suggest that a gradation of credibility enhancing processes (of both increasing cost and reliability) could be applied as research crosses successive thresholds. So, say at 250 citations (or whatever number) a replay review is triggered, but at 750 citations (made up number), some funding is automatically awarded to a group to perform an independent replication.

  9. i want an andrew read on those graphs! why don’t the axes/dotted lines line up? why doesn’t the x-axis start at zero? why isn’t the first graph on the log scale?

    • Bob:

      The graphs were re-drawn for publication. They didn’t want to include the graphs at all, and I put a lot of effort into insisting that the graphs were included. In doing so, I hadn’t noticed about them not lining up. I do think this opens up a good research project–it could be a QMSS thesis!–to gather similar data and make similar (and better) graphs for other fields.

  10. One of the things not mentioned here is the impact (and therefore incentive) to the reviewer. Being asked to review papers. most of which are destined for obscurity, is already a pretty thankless task… unless the thanks of journal editors is important to you. The incentives to review papers which are already deemed to have some actual importance ought to be much more appealing to reviewers…. I would expect that journal editors could lean on more careful or distinguished reviewers, and reviewers might in many circumstances have incentives to out themselves as reviewers. Increasing transparency here has obvious benefits.

    I question the resource savings, though, because I think prepublication review by someone will still be necessary. Reviewers will no longer stand athwart the publication process as an obstacle, but I think would still be very, very valuable as outside eyes looking at work and saying…. yo, this could really use a graph in section iv and the methodological section made no sense to me. This pseudo-review would not be aimed at *allowing* publication, but at an attempt to improve the publication. The author and/or editor would be free to ignore it, but I think an intelligent subject matter expert sending feedback before publication is still required…. or at least a really good idea.

    • Jonathan:

      You write, “I question the resource savings, though, because I think prepublication review by someone will still be necessary.”

      We’re not suggesting removal of prepublication review, and we don’t claim resource savings. What we say is that our proposal “would add only 1 percent to the total reviewing effort.” That said, I do recommend that the conventional review process be reduced from three reviews to two, which would save 1/3 of the effort!

  11. Agree with all the folks noting that 250 citations seems very high. Dozens of research grants could depend on such a paper by the time it gets 250 citations. In my world research can become influential at conferences from the talks that preceed publication, so that by the time the paper comes out there may already be funded proposals and / or papers in press based in part on that work. Of course it would not be possible to change proposals that are funded before publication or papers that come out before publication with any post publication review. But sooner review is still better. Perhaps papers above the 50th percentile in total citations after one year?

    IN the post you discuss the modest investment in additional labor in review required for your proposal. However, there is also potential for labor and direct cost savings by discovering and surfacing problems with a paper before it is used extensively to support further research. Raghu points out that all his papers with >250 citations are around two decades old. It seems like by that time the review has already been effectively done and any research misdirections or unexpected implications that may have come from it have already been tested and either played out or firmly established, so any further review is just a formal declaration of what is already known and has been extensively tested.

    Anyway it’s still great to see formal proposals coming forward for improving the research process. The process sounds good – I like the idea that the review becomes a citable publication – but I would definitely approve of a much lower bar for triggering reviews.

    • Anon:

      My original idea was 100. We picked 250 because it’s approximately 1% more effort, but, again, yeah, it would be fine to be 100 and just reduce total effort on the other end by generally requiring 2 reviews instead of 3.

      • 100 sounds better, but I don’t know many papers reach that level or how fast they get there. In my mind the primary utility of having post-pub review is the benefit it could have for future dependent research. Independent of other factors, what is an optimal time frame in which a review can be useful? The number I have in my mind is about three years. Any thoughts on that? I feel like research starts to age a bit after five years and approaches historical after maybe ten years, so at that point the utility of a review declines.

  12. A kind of variation on this interesting idea is for journals to have retrospective reviews written on significant “anniversaries” of influential papers. Here is one example: https://aiche.onlinelibrary.wiley.com/doi/abs/10.1002/aic.14878
    This example has also been highly cited (nearly 400 citations, which is far far above average for papers in this field).

    This also points to a possible confusion of language, in that this article was definitely a review article, but it was not exactly a review in the sense that Andrew pointed to in his post.

  13. Most of the papers I read come to me from citations by other people. Many are in the form of of PDFs on my hard drive or scans of articles which were digitalized as long ago as twenty years. So i wonder how any late post-publication review would reach readers of the famous article, given that research academics are swamped by things they could read already.

    We have all seen textbooks with some claims which were ten or twenty years out of date when they were published, and twenty to forty years out of date when they were set.

    • Sean:

      Agreed, our proposal is more forward- than backward-looking. Post-publication reviews won’t change printed copies or pdfs. What we’d like to see happen is the journals commissioning and publishing the post-publication reviews and linking to them at the same place where the published papers appear. So when readers find the paper on the journal’s website, they’d see the review. And even if you just have a pdf, you could google the paper title and then the review could come up.

      • The PubPeer browser extensions available mean that this is already achievable now (assuming that at least a link to the post-publication review has been posted on PubPeer)

  14. One way post-publication peer review does occur is through additional publications, which may (fail to) replicate the original or may cite it (un)favorably. We may already be at the point where a handy LLM could be created and trained to tell you whether a paper has any negative citations and the nature of the criticisms. Such a bot would work best on the papers that have a large number of citations.

    This would take full advantage of all existing post-publication peer reviews with no (incremental) additional effort.

    • There are already a few commercial services doing this (e.g. scite.ai), although in my experience the LLM’s emphasis on problems with findings in the literature is highly dependent on the wording of the user prompt, and of course how many critical relative to uncritical papers there are in the literature (although I assume that both these problems could be overcome with specific model architectures)

Leave a Reply

Your email address will not be published. Required fields are marked *