“On the poor statistical properties of the P-curve meta-analytic procedure”

Richard Morey writes:

My colleague Clint Davis-Stober and I have a new paper at JASA about Simonsohn et al’s “P curve” forensic meta-analytic tests, which are supposed to help identify “evidential value”, “lack of evidential value, and “left skew” in a set of test statistics. You’ve blogged about P-curve-related topics before so I thought you might be interested in our paper.

All the p curve tests are transformations of significant p values, summed to produce a known distribution if the p values are uniform. Their choice of transformations make for tests with hair-raising properties, including extreme sensitivity to values near the criterion, arbitrariness, inadmissibility, nonmonotonicity in the evidence, and general inconsistency of their “power” estimate.

Basically, they use transformations that approach infinity at the on the left at the significance criterion (like the probit) but on truncated one-tailed distributions (the chi-squared and F test statistics). This (and a few other choices) is disastrous.

We make the point that many of these techniques were never vetted by experts, and often are just “verified” by a few simulations. For tests, this is not good enough, but nevertheless these methods can get popular because (in my opinion) they tell people what they want to hear.

I asked Morey why their paper didn’t mention Hedges (1984), given that they were aware of that connection. Morey replied, “That was due to our focus away from the ES estimation techniques, and toward their tests, which are situated more in the nonparametric combination literature (or rather, should be). We figured the estimation angle had been explored by other authors, and since the tests hadn’t, we’d look into the tests (which are quite popular).”

Blake McShane summarizes the Morey and Davis-Stober article as follows:

Morey and Davis-Stober use fundamental mathematical statistics to show that the P-curve:

– Does not test what it claims to test (i.e., skewness or “evidential value,” which as they note “is not a well-defined statistical or scientific concept”)

– Has poor statistical properties including sensitivity, inadmissibility, and (horribile scriptu!) non-monotonicity.

In sum, this means that of the two different sets of methods (bewilderingly) both called the P-curve by the authors of the three papers listed at this website:
– The two sets of forensic meta-analytic tests introduced in JEP:G 2014 and JEP:G 2015 are poor tests as per Morey and Davis-Stober’s new paper.
– The meta-analytic effect size estimation method “introduced” in Perspectives 2014 is an alternative and inferior implementation of Larry Hedges famous 1984 paper/method as per prior work.

My take on this is slightly different. I agree with Morey, Davis-Stober, and McShane on statistical grounds, and I like the Hedges (1984) paper. In practice, though, I don’t see the different methods (again, I refer you to my post from 2018 for details) as being so different in practice, and my reason for saying this is that I see these methods more as rhetorical tools than anything else. I wouldn’t trust the estimates or hypothesis tests from any of these approaches!

To label a method as a “rhetorical tool” is not an insult. Rhetorical tools can be useful!

I offer a three well-known examples of statistical ideas arising in the field of science criticism, three methods whose main value is rhetorical:

1. The famous claim of John Ioannidis (2005) that “most published research findings are false.” As I wrote a few years ago , I don’t buy Ioannidis’s mathematical framing of the problem, in which published findings map to hypotheses that are “true” or “false,” but I’m still glad that Ioannidis wrote that paper, and I agree with his main point. I see his paper as rhetorical, as making the point in a simple way that most published research findings could be false, even though I would not want to use his method to actually estimate the rate of false findings, however that is defined. Again, this is not a slam on that influential 2005 paper, just my judgment on how it is useful.

2. The concepts of Type M and Type S errors, which I developed with Francis Tuerlinckx in 2000 and John Carlin in 2014. This has been an influential idea–ok, not as influential as Ioannidis’s paper!–and I like it a lot, but it doesn’t correspond to a method that I will typically use in practice. To me, the value of the concepts of Type M and Type S errors is they help us understand certain existing statistical procedures, such as selection on statistical significance, that have serious problems. There’s mathematical content here for sure, but I fundamentally think of these error calculations as having rhetorical value for the design of studies and interpretation of reported results.

3. Multiverse analysis (see this paper with Sara Steegen, Francis Tuerlinckx, and Wolf Vanpaemel) is another one of my popular ideas related to science reform. The method has caught on, and I’ve used it to help understand what can go wrong in science, but I don’t plan to use it in any of my own research! As with examples 1 and 2 above, I see multiverse analysis as more of a rhetorical tool, a way to see what can go wrong if you pick out just one analysis. It’s a kind of living version of researcher degrees of freedom and the garden of forking paths, just to mention two ideas which are more obviously rhetorical.

I don’t think all my methods related to science reform are primarily rhetorical; for example, I really do use multilevel modeling instead of multiple comparison corrections (here’s one of many examples), but there are times that a method can be helpful merely by serving a rhetorical purpose, as with the three items above.

And that’s how I see the p-curve method. No way I’d want to perform a formal hypothesis test or anything like that to detect the presence of questionable research practices or whatever. For reasons discussed by Morey and Davis-Stober, I think that would be a bad idea. But that’s never what I saw as the virtue of the p-curve approach. To me, the value of that method has always been in the concept that you could display an empirical distribution of p-values and compare to your general expectation of what would happen if everything were reported. Doing some math to make the formal comparison is fine in that it sharpens intuition, in the same way that the math in the Type M and S errors helps us understand the implications of these methods. But there are just too many assumptions floating around for me to want to do anything else here. For example, no way would I declare a problem just because a p-curve test yielded a p-value of less than 0.05, nor would I say everything’s ok just because such a test fails to reject.

I get that people do use the p-curve as a hypothesis test, so I respect the Morey and Davis-Stober paper: you have to meet people where they are.

Here’s what I wrote on the topic back in 2018:

I get concerned if people take these methods too literally. Consider the classic file-drawer-effect paper by Rosenthal, which I assume was written to demonstrate how serious this selection problem can be, but is sometimes twisted around to give the opposite meaning (by doing the calculation of how many papers would need to have been discarded to be consistent with a particular pattern of published results, and then claiming that since no such massive “file drawer” exists, the published claims should be accepted). I wouldn’t want researchers to take p-curve, or the Hedges approach, as evidence that a literature of uncontrolled p-values is approximately just fine.

As is often the case, I find myself more convinced by the demonstration of bias than by the attempted bias correction. In that sense, I see the Hedges procedure, or p-curve, or p-uniform, as being comparable to Type M and Type S errors (Gelman and Tuerlinckx, 2000) as a way of quantifying some effects of selection bias in statistical inference, but the desired solution is to go back to the original, unselected, data. All these methods can be useful in giving us a sense of the scale of bias arising in idealized situations.

I also recommend my 2013 paper with the late Keith O’Rourke, Difficulties in making inferences about scientific truth from distributions of published p-values.

Morey and Davis-Stober’s reactions

I sent the above post to Richard Morey and Clint Davis-Stober, who had the following comments:

Davis-Stober:

I’m not sure what it means for a person to “benefit from the method in practice” in this case. “Forensic” methods like p-curve are kind of strange conceptually. Usually, we’re building models to explain data–with the humility that our models are probably wrong but we’re trying to test out explanations. Analogous to a point made by Pek et al. (2025), forensic methods like p-curve don’t seem to work this way–the data themselves are suspect and it’s not clear what generating process created them. How do we know when p-curve works in practice? I think these methods are kind of a conceptual dead end.

Pek, J., Wu, H., Liu, Y., Dusenbery, K. J., & Wegener, D. T. (2025). What does a Z-curve analysis tell us? Manuscript submitted for review.

Morey:

I’d like to reiterate that these tests are really not well-connected to Hedges’ (1984) work–they’re nonparametric meta-analytic combinations of p values (e.g. Birnbaum 1954). This is relevant because there are really three things that are called the p curve, and this has caused a great deal of confusion. Obviously, the first “p curve” is just the distribution of p values under some hypothesis. The second is the p curve tests (tests based on sums of transformed p values, which we discuss) and the third is the meta-analytic methods for estimating effect sizes (it is these that are related to Hedges work). These three ideas are all only tenuously related. Some of the reaction to our work has been incredulous that we would not look at distributions of p values. Importantly, that’s not our argument. We don’t have a problem with looking at distributions of p values, but we would stress that one should be extremely humble in drawing concrete conclusions based on them.

I’m not against the rhetorical use of statistical ideas, and in fact I think all statistical methods are basically rhetorical: statistical methods are thought experiments, or kinds of games). If we accept that, then the question is what is the argument being made using the method? In the case of the p curve tests we discuss, the argument is clearly about specific bodies of work. The authors claimed to have “a key to the file drawer”, and these tests are supposed to be that key. The construction of any test can be regarded as the premises on which whatever rhetorical conclusion, and we argue that in this case the strength of any argument you’d construct on the basis of the tests specifically is weak, because they statistical properties are poor.

In the spirit of transparency and scholarship, I believe that the authors should have done a much more thorough job justifying their tests (that is, if you like, fleshing out precisely what the premises of the argument is that these methods allow you to make). If they had published in a statistical venue, they probably would have had to do much more work in fleshing out the properties of the test. As it is, they did not, and so we’ve had 10 years of people using this method without really understanding what it is. There have been 80 years of statistical literature on these types of methods (meta-analytic combinations of p values); I think they should have interfaced with that literature so that people (including them) could be properly informed. Instead, they cited Abelson’s introductory statistical pamphlet to justify their choice of transformation. I think this is poor scholarship.

Simonsohn’s reaction

I sent the above to Simonsohn, who replied:

Thanks Andrew for asking for my take on this.

I have written a Data Colada[129] blog post discussing the critique by Richard & Clintin. It’s titled “p-curve works well in practice, but, does it work well if you drop a piano on it?” I tried to make it short and fun, while engaging with the substance of the critique head on.

In a nutshell, none of the four criticisms raised are new. More importantly, they are unlikely to be consequential in real-life analyses (rather than in carefully chosen edge cases used to make rhetorical points).

Every statistical tool can perform poorly in edge cases. This is true even of the simple t-test (e.g., with outliers) and regression (under too many situations to list).

The question is not “is this tool perfect in all situations”. No tool is.

The question is “in realistic circumstances, are we better off with this tool than with the best alternative?”

The JASA paper doesn’t even try to answer that question.

Youall can read all the above and make your own decisions. I guess that much will depend on what sort of problems you’re working on and what you’re planning to do with these statistical summaries.

48 thoughts on ““On the poor statistical properties of the P-curve meta-analytic procedure”

  1. I wonder about the correlation between the rhetorical merit and the actual usefulness of science reform methods. People love a simple expose of the bad stuff, but that often sweeps a lot of nuance under the rug. At least in my own work the approaches that I’ve come to see as indispensable are not so obviously related to the reform ideas that sell.

    • I am very far from an expert, but I could imagine that part of the reason why certain science reform ideas have been so successful is that they promote ”business *almost* as usual”? ”We can stay within the paradigm if we just make these edits to it!”

      As a complete aside, I would love a blog post that lists ”the methods you’ve come to see as indispensable”.

      • +1. The science reform ideas do not go far enough. The paradigm needs to be completely rewritten. Focus on getting good data in experiments. Then focus on using simple methods of explaining the data and making precise and falsifiable predictions, and build from there. Like physics did centuries ago.

        • Anon:

          Sure, but in the meantime we want to be developing new drugs, estimating the effects of new treatments, etc. I agree 100% that better measurements and scientific models are important, but statistical challenges will remain with whatever data we have.

        • Correct, the problem is the null hypothesis. Make the null hypothesis == research hypothesis and 99% of the problems disappear like magic. The current method is bizarro to science, in science you want non-significant deviations from your hypothesis.

    • Came here to say much the same thing. The fine-line between rhetorically useful math and Bullshitting (in the technical sense) seems to be whether the authors draw insight or inferences. The Drake equation is *very clearly* a rhetorical tool intended to provide insight into the rarity required for various features of the universe for life to be just us. Rosenthal’s paper did a great job of not taking itself too seriously and just sort of pointing out you’d have to have a massive file-drawer to have the literature be entirely false positives.

      Ioannidis, however, drew an inference and made a specific claim that most of the literature is false. This data-free argument has evolved to his current claims that science is net negative for humanity. The paper doesn’t add too much to Rosenthal’s model, beyond some bells and whistles that make the bullshit more persuasive by adding parameters. The same goes for the p-curve. The authors have used it pretty routinely over the past decade to draw inferences about the value of bodies of research, scientists, etc., and even gone so far as to say it provides direct insight into (or adjusted for) things like p-hacking.

      Had the authors simply done some simulations to point out that p-values should be small more often than they’re large, under certain conditions, it would have felt more like insight and a rhetorically useful tool

      • J:

        Regarding your statement, “Rosenthal’s paper did a great job of not taking itself too seriously and just sort of pointing out you’d have to have a massive file-drawer to have the literature be entirely false positives”:

        As a paper, Rosenthal’s paper is fine, but I strongly disagree with its message, and I think it’s misled a lot of people. I say this because I think that forking paths (selection within a project) is a much bigger issue than file drawer (selection among projects). I have the impression that the catchy phrase, “the file drawer problem,” along with Rosenthal’s calculations, led researchers for many years to underestimate the distortions caused by selection. I suspect that in an alternative world in which the term “p-hacking” had come first and “file drawer” had come much later, we’d have avoided many of the Ted-talk excesses of the 2010-2015 period. Maybe Kahneman might never have said that memorably wrong “You have no choice but to accept…” thing!

        • I think that’s exactly right. Knowing about p hacking allows you to apply a useful heuristic. Personally, I think I’ve entered values into the p curve app exactly once, but p values close to .05 do put me on alert when reading. If we treat it as a heuristic, I think it’s a good one in many realistic cases, especially when reading the older literature. I don’t know of a formal investigation other than the fact that in the Reproducibility Project Psychology, high p values and low intuitiveness/prior belief ratings predicted nonreplication, as any good Bayesian would expect.
          As with any forensic tool, there is an arms race though, so I guess these discussions are academic in more than one sense, as researchers learn what to do so as not to arouse suspicion.

        • I don’t think rosenthal set out to argue (or did) that file drawers were the cause of false findings in the literature. He makes the case that file drawers can explain a few errant studies with small effect sizes, but get us nowhere for large literatures with bigger effect sizes.

          I’d have loved an alternate universe where it prompted us to be uninterested by small but significant effects. Reading the lit three days it seems more normal to just specify a massive N and make a whole story about some tiny effect on a scale that doesn’t translate into anything. We wound up focusing on making p values small by increasing n rather than d.

    • The Datacolada post reads to me like some combination of responses 5, 6, and 7 about which Andrew’s ladder post to which you link says “responses 5, 6, and 7 increasingly pollute the public discourse.”

      On the other hand, the thread that Richard Morey links to below seems to do a really nice job of trying to make sense of this all from a sociological perspective and provides a stunningly close analogy with his “W-test.”

      It boggles the mind that the P-curve authors have never proven a single thing about the P-curve and shift the burden of proof on critics of it–especially given that their P-curve tests don’t test what they claim they test (and that they test what they do test so poorly). I wish this new JASA piece had come out sooner and my heart breaks for all those whose careers have been harmed as a result of reckless application of P-curve tests.

        • Fair enough and reasonable people can disagree. I definitely did not mean to suggest an equally-weighted combination. I agree that minimizing (Response #5) gets the largest weight but I personally saw some patching / protecting (Response #6) and smearing (Response #7).

      • They show that they responded to all of the criticisms long ago or showed that they weren’t important (also long ago). Maybe there are shades of #5, but nothing else is wrong with the post. They even showed that the authors of the JASA piece used the same criticism that they themselves used in one of their previous blog posts (the one about heterogeneity), and claimed it is new. (That’s in Appendix 1 and footnote 1.)

        • I don’t want to get involved in arguments with Anonymous posters in blog comments these days, but for the record this is not true. Simonsohn was not aware of the inadmissibility (he admits he was not familiar with this literature). He was not aware of the nonmonotonicity (he admits “We did pay less if any attention to the issue specifically at the .025 cutoff, and the JASA paper focuses on that cutoff for some examples” – by “some” examples he must mean “basically the whole nonmonotonicity section”). He has not shown that they weren’t important. This is, in principle, probably impossible given that “important” is just a value judgement (we know *he* doesn’t think they’re important!). But he never showed anything like “Under these sufficiently broad conditions, the bias in this estimate will be be below X”). Instead they did a few simulations and said it was “fine”.

          Regarding the “same criticism,” this is, too, is untrue and it requires a misreading of the Morey and Davis-Stober paper (section 4.3) to say it. In this section (and the supplement), Morey and Davis-Stober prove that the estimate is consistent in the trivial case with one noncentrality parameter, but in general inconsistent unless a very specific condition holds (and given the variance will go to 0 asymptotically, it is asympotically biased). In the referenced blog post with supposedly the “same criticism” there is simply one example that Uli Schimmack proposed showing that in at least one case, the estimates seems way off.

          If you see these as the “same criticism” I honestly don’t know what to say. But Simonsohn didn’t help you much, because in his Appendix he 1) takes the example out of the argument context, 2) creates his own graphs to make them visually more similar, and 3) drops discussion of one of their TWO supposed requirements for apparent bias (a substantial number of “low powered” studies) because it is visually obvious that our example is different on the “low power” end.

          I’d say this counts as #6 on the ladder, and with the implication of a “lack of credit” in the Appendix (when we’re making a point about inconsistency in general, which cannot be shown/refuted by a few simulations, and we cited the blog post anyway) veers into #7 (but obviously I have a dog in this fight).

        • I’d be largely comfortable with this if “show” were replaced by “declare.” This is especially interesting in this context in that it is emblematic of a larger issue: Morey and Davis-Stober do in fact show (mathematically prove) whereas JEP:G / Datacolada illustrate (via a handful of simulations) and then declare that they have “shown” (in the sense of demonstrated or proved).

        • You seem to be jumping the gun here. Mathematical proofs are only as good as their assumptions, and even if they are true, they could be irrelevant. In fact, this is exactly what Simohnsohn et al argue when it comes to inadmissibility–they argue that being inadmissible is not always a problem (and cite a well-known paper proving this (in the sense of statisticians)).

        • This comment seems confused. First, the Datacolada blogpost does not “cite a well-known paper proving this (in the sense of statisticians)” [if it does, I missed it and apologize].

          Second, and regardless, there is no need for such a proof nor is it possible. (In)admissibility, a textbook mathematical statistics concept typically introduced at the advanced undergraduate level, is a property of a method and an “(un)desirable” at that. Like any other property, the degree to which it is (un)desirable is application- and user-specific and depends on the extent of the violation (e.g., just like unbiasedness is a “desirable” property ceteris paribus, but we typically face a bias-variance tradeoff and are willing to take on some bias in order to reduce variance and thus obtain greater accuracy).

          Third, stepping back, proofs are proofs and can be indeed be relevant or irrelevant in a specific application. We don’t need Datacolada to tell us that and they do not have some special insight into or jurisdiction over whether they are relevant in a particular application.

          Analogously, I can read the post and see that the simulations contained within it are not so relevant in the applications I work and thus I will take them cum grano salis but your mileage may vary!

        • Morey and Davis-Stober in their argument for inadmissibility cite Marden (1982) for the proof that the procedure is inadmissible. Data Colada, in their argument against inadmissibility, cite a paper that cites Marden (Rice (1990)). Rice (1990) argues that sometimes using the admissible test is not justified, and to use the same test that Data Colada are arguing for.
          This is the argument that inadmissibility is “not important”. In certain situations that they explain in the appendix (results driven by one very strong test), it is not a good criterion.

        • Richard, you claim that their argument was based on seven simulations. Data Colada show that there were nine, and the two you ignored correspond to some of your critiques.

        • I don’t know whether you are confused or arguing in bad faith, Anon, but I will try clear things up. You first wrote that Datacolada “argue that being inadmissible is not always a problem (and cite a well-known paper proving this (in the sense of statisticians))”.

          When asked for clarification about which paper, you mention Rice (1990) and write that “Rice (1990) argues that sometimes using the admissible test is not justified…This is the argument that inadmissibility is ‘not important’.”

          Rice (1990) does not discuss admissibility. Therefore, contra your initial claim, Rice did not prove being inadmissible is not always a problem (in the sense of statisticians), which in any event is not possible as I wrote above (for the “value judgment” reason Richard writes about above).

        • You’re nitpicking my wording, Replicate. Always a bad sign. Arrogance, pedantry, and confusion…it’s all here.
          Let me make this clear to you. Rice (1990) does not mention the word “admissibility”. It says that Fisher’s test is “inappropriate when asking whether a set of tests, on balance, supports or refutes a common null hypothesis”. (p.303). It has “greater sensitivity to data that refute a common [null] compared to data that support it,” i.e. “differentially sensitive [to such data]” (p.305). It mentions Stouffer’s test, which Data Colada use, as “not differentially sensitive to [the same]” (p.305). Fisher’s test has more power, which is what makes it admissible, but it is still not useful in this case. This is in the appendix and can be confirmed by looking at the original paper.
          My initial wording could have been improved, but my point stands.

        • You not only are being nasty but also are grossly misrepresenting matters.

          You write, “This is in the appendix and can be confirmed by looking at the original paper.”

          I did confirm by looking at the original paper, the appendix to which reads:

          APPENDIX
          An interactive computer program that carries out all of the tests described in the text can be obtains by sending a self-addressed envelope, with proper postage attached or provided, to the author. A disk copy of the program can be obtained by sending a formatted (IBM format) 5 1/4″ floppy disk along with a self-addressed disk-mailer, with proper postage attached. The program in standard Pascal. A compiled version of the program (Turbo Pascal, Borland International) is also available.

        • The appendix to the Data Colada blog post. I gave the page numbers of the article in the previous comment. I’m not sure if you’re trolling me at this point.

    • From the thread: “I can tell you that no number of simulations is going to show you what is wrong with the test. You already know that something must be wrong, but only because of your *theoretical* knowledge about Z tests.”

      Why? You must have gotten the idea that the Z test was most powerful from somewhere. If I simulated a bunch of different tests, I would eventually see that the optimal one was the Z-test. You even showed the curve for it. Why can’t I get the curve through simulation, or simple calculations?

      I think the “meta-science” viewpoint/style is better than the focus on so-called “rigorous proof” that statisticians have. Maybe Simohnsohn didn’t use the style correctly and his P-curve is wrong, I don’t really have an opinion on that. But I argue that the M-S style (as well as humility, not fooling yourself, etc) is better for doing actual science/statistics.

      Even Andrew says that “guarantees are just assumptions”.

      • Pretty ironic to talk about humility of the M-S approach there, given the sheer arrogance and self-absorption of Simonsohn’s posts, in both thinking that the onus should be on others to disprove, and even worse, being all snarky and uppity because they weren’t asked to look at or review Morey’s piece before publication! Frankly quite embarrassing, and I’m impressed by Morey and co’s restraint in the exchange

        • First of all, I was saying humility in addition to the M-S style. But I don’t read the posts as arrogant at all. I get the opposite impression. Saying that just because it is non-monotonic 2.8% of the time it’s obviously wrong and should be *immediately* memory-holed and laments about sociology and etc reads as extremely arrogant. Simonsohn also wrote published papers on the p-curve a decade ago, giving extensive arguments for it. Now, the onus really is on others to dispute those.

          Simonsohn also asked politely to see the JASA piece before it went public and the authors refused. There is no snark there. Also, Data Colada’s feedback policy is always to talk to the authors of a piece they do a blog on privately first. It is only fair to expect similar treatment.

          I’m impressed by Simonsohn’s restraint.

      • I initially thought that this was a facetious or trolling comment intended to undermine the M-S approach but apparently it was meant in earnest so let me do my best to try to explain.

        We know that the Z-test is most powerful because of a proof. Sure, any proof is only as good as its assumptions. However, because of the proof, we know that under the very wide range of conditions covered by its assumptions, the Z-test is most powerful.

        If on the other hand you simulated a bunch of different conditions as you suggest in your comment, you would only know that (i) it is the most powerful of the tests that you chose to evaluate in your simulations (ii) only in those very specific conditions that you chose to simulate, and then still (iii) you would only know this up to simulation error (which you can control by running more iterations of each particular simulation).

        Of course, whether and the degree to which the assumptions apply in any given application (as well as the harm caused by violations of them) requires careful thought and judgment and reasonable people can in some cases come to different conclusions about this.

        • “you would only know that (i) it is the most powerful of the tests that you chose to evaluate in your simulations”

          “under the very wide range of conditions covered by its assumptions, the Z-test is most powerful.”

          These are the same thing. You only know the Z test is the most powerful under the range of conditions that you chose to include in the proof. Why can’t I simulate or calculate the extreme cases and get to the same conclusion? I can vary all the parameters I like, and everything that applies to reality has error bars associated with it. There is no such thing as a “proof” in science.

          Mathematical proof is not some mystical thing that gives you knowledge that you cannot get any other way. A proof is simply a strong argument, and the intuitive feeling that the statement is true (a prerequisite for a proof) is usually given to the prover from some kind of heuristic evidence (which can include simulations). Try e.g. “Proofs and Refutations” by Lakatos, Andrew has talked about it in the past.

          https://sites.stat.columbia.edu/gelman/research/published/Lakatos_AMM.pdf

          I’ll also give an example from biophysics. There are things analogous to motors at the molecular level inside of cells. These molecular motors can behave in an almost optimal way, but this depends on the conditions they are in. Physicists don’t prove things, but they still know that there is an optimum (100% efficiency) and they can calculate the optimal behavior of the motors. The experimental data never reaches the curve corresponding to 100% efficiency, but it can get close. See figures 2 and 5 from this paper: https://doi.org/10.1039/c0cp02118k

        • They are *not* at all the same thing:

          A proof can apply to a continuum or infinity of conditions (as the proof regarding the Z-tests does).

          A simulation is necessarily limited to the specific discrete and finite number of conditions actually simulated (and then only holds up to simulation error).

        • I see that you ignored the main arguments that I made. Let me state them again, clearly and simply:
          You can get a good idea of the optimal behavior that you expect through simulations and calculations (I showed an example from biophysics). No proofs required to gain this knowledge. Proofs come later if you’re a statistician or mathematician.

          Second: you can simulate the extreme cases and in so doing, find the limits of the infinity of conditions. If something has a condition with a parameter that varies between 0 and 1 (an infinity of conditions) simulate it at 0 and at 1, or as close to those values as you can get.

          Give me a counterexample if you disagree.

          Why do statisticians have some type of religious obsession with proofs?

        • The second point, stated more clearly:
          You can approximate a continuous infinity of conditions with a discrete and finite set of them. An analogy: when you make a plot of a smooth curve, you approximate the uncountable infinity of points that make up the curve with a discrete and finite set of them that you see on the screen or paper.

        • I can’t speak for anyone’s religious beliefs, but with your clarification it seems we are in agreement:

          A proof can demonstrate something for the entire continuum between 0 to 1 (to use your example) while a simulation can only do so for the discrete and finite set of values between 0 and 1 actually simulated.

        • To be clear, this is not intended as a knock on simulation just a recognition of one of its limitations.

          With Morey and Davis-Stober, I agree that “Simulation is a powerful tool and can help build intuition.

          However, I also agree that “it is limited by the simulator’s imagination and confirmation bias” and “it is not a substitute for formal analysis. Simulation may provide hints of problems with a procedure, but only if the simulator’s formal knowledge helps guide the choice of simulations. A simulator might quit after running a few simulations that tell them what they think is true while problems remain uncovered” and that simulation.”

  2. This discussion reminds me of the extent to which we hedged our claims in the RIVETS preprint (10.31234/osf.io/ctu9z) regarding the statistical meaning of a series of “hits”. It felt like we were treading on eggshells at the time with our disclaimers, but on re-reading it, I’m not sure that we were conservative enough. As Davis-Stober says, “the data themselves are suspect and it’s not clear what generating process created them”, but we still managed to multiply four small numbers together and claim that the resulting number kind of worked like a p value.

  3. P-curve,
    Kicked to the curb
    Because just around the corner,
    There’s something new to replace the former
    There’s bound to be another trend
    Or something new around the bend
    Perhaps P-swerve might do the trick this time
    If nothing else, it completes this rhyme

  4. I read the footnotes of the blog post no. 129 by Simonsohn and footnote 5 is about the pros and cons of sharing critiques in advance. I wanted to note a few things concerning that:

    The first one is that it is my personal experience that one doesn’t always, or even regularly, or even normally, receive a reply from academics when you send them an e-mail. Especially perhaps when you send something critical of their work. I may not be the only one with that experience, and if I remember correctly there was a paper published around 2012-2014 on PLoS ONE that replicated some academic’s work. The back-and-forth in the comment section of that paper showed the original author that mentioned that the replicators should have contacted them earlier (or something like that) to which the replicators replied that they did so twice but received no reply.

    Regarding the possible pro of sharing critiques with the original authors (I presume before publishing) that there might be errors in the critique or that it may come across as snarky or inaccurate or whatever that the original authors may spot and point out seems technically correct, but practically pretty useless in my view. What one person sees as snarky or inaccurate is not the case for another, and to me it seems that the original authors do not necessarily have some perfect notion of when something is snarky or not, especially when they might be relatively more sensitive to that because they are being criticized. It’s also a bit strange to me, from a certain perspective, to think the original authors might be very good at spotting possible errors (because of their possible “expertise”) and should possibly be asked for their view beforehand when the critical scientists are essentially wondering and/or showing that the original authors may have not spotted errors or other problematic issues.

    The “without sharing in advance, you take control over my time (…)” comment can also be viewed the other way around. If the authors being critical have to first send in their critical comments to the original authors, they are dependent on the original authors to read, and reply. As mentioned earlier, it is my experience that I don’t usually receive a reply from academics, and (coincidentally) I will wait about two weeks before concluding that it doesn’t seem I will receive a reply. Can I then also state that “by sharing in advance, you take control over my time (…)” because I have to wait, at least two weeks in my case, before I can move on or think about other options, or I will have to wait until the original author replies within that two week period.

    The statement that “For the audience, the most important party, it’s always better to see argument and counter-argument jointly so they can make up their minds with all the information” might be correct, or not, or it may be correct in some cases and not in others. For instance, it can be possible that a critical piece being published without the original authors’ reply visible might cause a more thorough reading and understanding of that critical piece compared to the situation where the original authors reply or views cause certain people to not really read things properly, or to merely side with the original authors because of X, Y, or Z. Also, I don’t see what this all has to do with sharing critiques in advance. I mean, does sharing critiques in advance lead to the audience somehow seeing argument and counter-argument jointly, and if so how exactly?

    Concerning “It’s crazy we were not invited as reviewers of this paper, counter normative to how many journals operate, it’s a conciliatory act to make up for that” I would like to note that I had to look up the word “conciliatory”. If I understood that word correctly, I wonder how this relates to sharing critique in advance. Regardless of whether it’s strange for the original authors to not be asked as reviewers, how could the critical authors make up for that when I reason that sending their paper in starts the entire review thing which they then can’t roll back. I don’t understand then how they could have send their critique to the original authors as a possible “conciliatory act” when they, presumably, do not know whether the original authors will be (or are the) reviewers.

    And finally, the comment that “Not sharing it may become a distraction if that becomes part of the conversation instead of the substance.” can also be thought about differently. One could wonder who is responsible for this possible distraction. Or one might wonder who is also responsible for this to possibly happen. Or one could wonder what even is a distraction. Regardless, I enjoyed pondering these things written above as a result of the comment that not sharing it may become a distraction if that becomes part of the conversation. Perhaps a distraction can become a separate focus, which might be good for something.

    • > the comment that “Not sharing it may become a distraction if that becomes part of the conversation instead of the substance.” can also be thought about differently

      Nice paper you got there, pal. Be a shame if not sharing it became part of the conversation, innit?

      • Look as a fellow Anonymous I can’t really say that it might be sub-optimal to use Anonymous, but when it comes to replying to another Anonymous I just can’t help but saying that that might be sub-optimal in that case.

        Regardless, I don’t understand your comment. Like at all. What do you mean by “nice paper you got there, pal”? I just noted some things about the footnotes in the blog post…

        • The comment (not by me, no idea who Anonymous is) is a sarcastic variation on the kind of thing a UK gangster says in the movies when he’s running a protection racket… like “Nice business you have here, it’d be a shame if someone burned it down, wouldn’t it?” Apparently implying that people complaining about papers are like gangsters? Not really sure.

        • Ah okay, so it might be a humorous comment and not a comment more directly related to something I wrote. Thank you for the reply and the information about the movies (I did not pick up on that possible reference)!

        • On the other hand (and consistent with all the anonymity here), it could be a playful way of intimating things like fear of bullying or retaliation.

        • Daniel’s right, it was just a different way to think about that comment.

          It may not be the intention of the footnote to defend a “you should share of else” approach to the issue but there is some nice irony in introducing the lack of sharing in the conversation while pointing to the risk of such a thing happening.

        • @Just Another Anon: Yes, that’s in line with the comment by Daniel Lakeland in my view. It was likely a humurous comment I did not pick up on. Thank you for your reply as well.

          Concerning the Anonymous-thing I would like to note that I just use it because I started using it when I first commented on this blog over 10 years ago if I remember correctly. I am not an academic. I also, for instance, regularly post a title of a manuscript I have written to make a point or refer to some more information about a certain topic which nullifies the anonimity. I will do that again now, here, by pointing out that my comments above about the blog post no. 129 and the footnote no. 5 might be another reason why I think there should be much more attention for logic, reasoning, and argumentation in psychological science.

          If my points above make, at least some, sense perhaps it points to something worthwhile to think about some more. To quote something from one of my manuscripts titled “Pyschological Science replicates just fine, thanks” which can be found on SSRN:

          As Katzko (2006) states: “(…) sound reasoning is as much a part of a scientist’s methodological toolkit as are procedures for data collection and analysis.” (p. 210)

          I can’t help but feel the footnote no. 5 referred to above in the comment I wrote down might be (depicted) sub-optimal, but perhaps I am making mistakes in reasoning and/or interpreting things incorrectly or sub-optimally. Whatever may be the case, perhaps it all points (again) to the importance of reasoning and argumentation (if those are the appropriate terms to use in this all). I have tried to get some more attention concerning this all in the best way I can via comments on this blog here and via manuscripts I have written. I hope this is all something worthwhile to note, mention, and try to get some more attention for.

  5. Morey said “in fact I think all statistical methods are basically rhetorical: statistical methods are thought experiments, or kinds of games). If we accept that, then the question is what is the argument being made using the method?”

    If everyone using stats methods absorbed this advice then the literature would be infinitely better. It’s amazing how many papers don’t even mention the core assumptions on which some conclusions are based, and these are usually the most important (e.g. sample is externally valid; adjustment set removes all bias; zero risk-of-bias etc.)

Leave a Reply

Your email address will not be published. Required fields are marked *