We want to go beyond intent-to-treat analysis here, but we can’t. Why? Because of this: “Will the data collected for your study be made available to others?” “No”; “Would you like to offer context for your decision?” “–“. Millions of taxpayer dollars spent, and we don’t get to see the data.

Dale Lehman writes:

Let me be the first (or not) to ask you to blog about this just released NEJM study. Here are the study, supplementary appendix, and data sharing statement, and I’ve also included the editorial statement. The study is receiving wide media attention and is the continuation of a long-term trial that was reported on at a 10 year median follow-up. The current publication is for a 15 year median follow-up.

The overall picture is consistent with many other studies – prostate cancer is generally slow to develop and kills very few men. Intervention can have serious side effects and there is little evidence that it improves long-term survival, except (perhaps) in particular subgroups. Treatment and diagnosis has undergone considerable change in the past decade. The issue is of considerable interest to me – for statistical reasons as well as personal (since I have a prostate cancer diagnosis). Here are my concerns in brief:

This study once again brings up the issue of intention-to-treat vs actual treatment. The groups were randomized between active management (545 men), prostatectomy (533 men), and radiotherapy (545 men). The analysis was based on these groups, with deaths in the 3 groups of 17, 12, and 16 respectively. Figure 1 in the paper reveals that within the first year, 628 men were actually in the active surveillance group, and 488 in each of the other 2 groups: this is not surprising since many people resist the invasive treatment and possible side effects. I would consider those that chose different groups than the random assignment within the first year as the true effective group sizes. However, the paper does not provide data on the actual deaths for the people that switched between the random assignment and actual treatment within the first year. So, it is not possible to determine the actual death rates in the 3 groups.

The paper reports death rates of 3.1%, 2.2%, and 2.9% in the 3 groups. If we just change the denominators to the actual size of the 3 groups in the first year, the 3 death rates are 2.7%, 2.5%, and 3.3%, making intervention look even worse. If we assume that half of the deaths in the random prostatectomy radiotherapy groups were among those that refused the initial treatment and opted for active surveillance, then the 3 death rates would be 4.9%, 1.2%, and 1.6% respectively, making active surveillance look rather risky. Of course, I think allocating half of the deaths in those groups in this manner is a fairly extreme assumption. Given the small numbers of deaths involved, the deviations from random assignment to actual treatment could matter.

The authors have the data to conduct both an intention to treat and actual treatment received comparison, but did not report this (and did not indicate that they did such a study). If they had reported details on the 45 total deaths, I could do that analysis myself, but they don’t provide that data. In fact, the data sharing statement (attached) is quite remarkable – will the data be provided? “No.” That really irks me. I don’t see that there is really any concern about privacy. Withholding the data serves to bolster the careers of the researchers and the prestige of the journal, but it doesn’t have to be that way. If the journal released the data publicly and it was carefully documented, both the authors and the journal could receive widespread recognition for their work. Instead, they (and much of the establishment) choose to rely on their analysis to bolster their reputations. But these days the analysis is the easy part, it is the data curation and quality that is hard. Once again, the incentives and rewards are at odds with what makes sense.

Another question that is not analyzed but could be if the data was provided, is whether the time of randomization matters. The article (and the editorial) cites the improved monitoring as MRI images are increasingly used along with biopsies. Given this evolution, the relative performance of the 3 groups might be changing over time – but no analysis is provided based on the year upon which a person entered the study.

One other thing that you’ve blogged about often. For me, the most interesting figure is Figure S1 that actually shows the 45 deaths for the 3 groups. Looking at it, I see a tendency for the deaths to occur earlier with active surveillance than either surgery or radiation. Of course, the p values suggest that this might just be random noise. Indeed it might be. But, as we often say, absence of evidence is not evidence of absence. The paper appears to overstate the findings, as does all the media reporting. Statements such as “Radical treatment resulted in a lower risk of disease progression than active monitoring but did not lower prostate cancer mortality” (page 10 of the article) amounts to a finding of now effect rather than a failure to find a significant effect. Null hypothesis significance testing strikes again.

Yeah, they should share the goddam data, which was collected using tons of taxpayer dollars:

Regarding the intent-to-treat thing: Yeah, this has come up before, and I’m not sure what to do; I just have the impression that our current standard approaches here have serious problems.

My short answer is that some modeling should be done. Yes, the resulting inferences will depend on the model, but that’s just the way things are; it’s the actual state of our knowledge. But that’s just cheap talk from me. I don’t have a model on offer here, I just think that’s the way to go: construct a probabilistic model for the joint distribution of the all the variables (which treatment the patient chooses, along with the health outcome) conditional on patient characteristics, and go from there.

I agree with Lehman that the intent-to-treat analysis is not the main goal here. It’s fine to do that analysis but it’s not good to stop there, and it’s really not good to hide information that could be used to go further.

As Lehman puts it:

Intent-to-treat analysis makes sense from a public health point of view if it closely reflects the actual medical practice. But from a patient point of view of making a decision regarding treatment, the actual treatment is more meaningful than intent-to-treat. So, when the two estimates differ considerably, it seems to me that they should both be reported – or, at least, the data should be provided that would allow both analyses to be done.

Also, the topic is relevant to me cos all of a sudden I need to go to the bathroom all the time. My doctor says my PSA is ok so I shouldn’t worry about cancer, but it’s annoying!

I told this to Lehman, who responded:

Unfortunately, the study in question makes PSA testing even less worthwhile than previously thought (I get mine checked regularly and that is my only current monitoring, but it is not looking like that is worth much, or should I say there is no statistically significant (p>.05) evidence that it means anything?

Damn.

50 thoughts on “We want to go beyond intent-to-treat analysis here, but we can’t. Why? Because of this: “Will the data collected for your study be made available to others?” “No”; “Would you like to offer context for your decision?” “–“. Millions of taxpayer dollars spent, and we don’t get to see the data.

  1. A review of the data availability statement shows that it goes beyond a simple no.

    Specifically, it states, “Requests for data extracts will be considered by the Chief Investigator” and gives a link for further information.

    The link does not copy and paste due to an imbedded space (%20). But editing it gives:
    https://www.bristol.ac.uk/population-health-sciences/projects/protect/collaboration/

    The first two paragraphs at that web page are:
    We welcome collaboration from potential investigators provided that any new studies relate to the field of urological research.

    If you would like to use ProtecT Study data for a new research study or a publication, please complete the Access to data and/or samples from the ProtecT study Proforma (Office document, 29kB) as the means to seek approval from the ProtecT study PIs.

    • I think the more appropriate link is to the data sharing agreement from the publication (linked above). It does go beyond a simple “no” but is highly qualified – data extracts will be considered, subject to approval, and after some time period. So, they aren’t simply denying access to any data, but it is surely not an open invitation. I’d put the odds of me getting the data I want at 10% or less (based on my previous experiences requesting data which have been successful 0% of the time).

  2. Dale wrote: “However, the paper does not provide data on the actual deaths for the people that switched between the random assignment and actual treatment within the first year. So, it is not possible to determine the actual death rates in the 3 groups.”

    I think that is true. However, Figure 3 appears to show that 85% of those assigned to radiotherapy had some radical intervention within the first six months of the follow-up period. For prostatectomy, the corresponding number is about 80% in the first year of the follow-up period.

    It seems to me that you should be able to make some inferences from the data in that chart about switching.

    Note also that at about seven years of follow-up, 50% of the active monitoring group had undergone a radical intervention. That rose to about 60% at 15 years.

  3. There is so much confusion about prostate cancer, so before this day’s blog gets going, I suggest everyone read this 2013 BMJ article first

    https://www.bmj.com/content/346/bmj.f548

    which is written by Gerd Gigerenzer and features the former and still pincushion, Rudy Giuliani who, in 2007, manages to get everything wrong, a foretaste of things to come.
    Rudy’s major mistake is that he confuses survival rate with mortality rate. PSA testing in the U.S. is done at an earlier age than in other countries so that a five-year survival rate is higher in the U.S.; but the mortality rate is, more or less, the same in this country as it is in other developed countries.

    “To Giuliani this meant that he was lucky to be living in New York and not in York, because his chances of surviving prostate cancer seemed to be twice as high in New York. Yet despite this impressive difference in the five year survival rate, the mortality rate was about the same in the US and the UK. Why is an increase in survival from 44% to 82% not evidence that screening saves lives? For two reasons. The first is lead time bias. Earlier detection implies that the time of diagnosis is earlier; this alone leads … ”

    After reading Gigerenzer’s article, then go on to read
    https://www.medscape.com/viewarticle/828854

    where you will find a history of the misuse of the PSA test as seen by the discoverer of the PSA test.

    “It was much to Dr. Ablin’s dismay that more than 2 decades later, in the mid-1990s, the US Food and Drug Administration (FDA) approved the use of PSA not only to test for recurrence of cancer, but also as a possible predictor of cancer. Since then, Dr. Ablin maintains, the United States spends billions each year administering a preventive prostate cancer screening test to men, using PSA, that produces false positives in the majority of cases. In his interview with Dr. Topol, Dr. Ablin explains why physicians and patients should proceed with caution when using PSA as a marker for preventive screening.”

    • Interesting pieces. Thanks for posting those links.

      The first linked paper contains a sentence that ties back to a major theme of this blog—statistics education. That statement is:

      For decades medical schools have failed to teach students statistical thinking (biostatistics does not seem to help much).

  4. “cos all of a sudden I need to go to the bathroom all the time. ”

    Have you taken to drinking more carbonated beverages lately? If so, stop. (Sometimes the carbonation can be irritating. For some folks. For others, the alcohol or caffeine can do it. Just.Drink.Water. In sensible amounts.)

    I’ve been reading arguments on breast and colon-cancer screenings and wondering if they were going to go the way of the PSA test, but those cancers kill enough more people than prostrate cancer that I’m probably wrong on that.

  5. You can have an enlarged prostate without having prostate cancer. It’s the enlargement that impinges on the urethra and prevents the bladder from emptying when you pee. There are relatively low risk prostate surgeries (short of total removal) that can address this. I write from experience.

  6. Over 30 years ago, I was one of the first to have a PSA test done and thus, before numbers were agreed upon for what was considered to be normal vs. concerning/indicative of cancer. Although my PSA stuff was eventually–norms were not yet quantified–deemed very high and mounting, I was consoled by the fact that my Gleason score (a measure of how normal/abnormal the cells appeared) was really low, “1+1.” Because of the topic of this blog today, I looked at the current situation with Gleason scores,

    https://en.wikipedia.org/wiki/Gleason_grading_system

    To my surprise, this grading system changed in 2005 and changed again in 2014 so that comparisons with the existing system back in the last century is tangled to the point that confusion is likely. For example, no longer n1+n2 as it was back then; currently, a 5-point Gleason grade is in effect.

  7. The data is the data and should be presented clearly, fully, and accessibly. Intention to treat is the convention; there are reasons for this. We should be leery of manipulating the data. I think that consciously or unconsciously it is commonplace to favor treatment/no treatment, and it is hard to alter our beliefs. I found that patients had deeply held views that were hard to impact. Prostate cancer patients were more resistant to physician input than say lung cancer patients. This is just a personal observation, and I welcome the opinion of others. I tried to present the current state of the art in a neutral fashion, but I acknowledge that I may have let views sneak in. The biggest problem I had was that the situation is complex, and no one wants to hear a seventeen minute lecture. People, patients and doctors, want a two minute answer. Not always possible.
    I have had a procedure on my prostate based on an emotional reaction and not on thinking like a physician/scientist.

  8. Less on the data sharing, more on intent-to-treat etc.: It seems like often people want to do as-treated or per-protocol analyses, but these have pretty severe problems. Is there some reason we’d trust those here (or developments in these approaches I don’t know about)? My sense is that basically you want to do either a ITT or IV approach, where the latter could include not just a point estimate for compliers but also things like partial identification.

    • An observation on defining as-treated.

      The paper states that, after 10 years, 54.8% of the men in the active-monitoring group had received radical treatment (radiotherapy or prostatectomy). Figure 3 shows that about 10% of the active-monitoring group had a radical intervention in the first four months of the study.

      It seems to me that classifying someone in the active-monitoring group who had a radical intervention in the first week of the study as being in an intervention group would completely obscure the proper comparison between the treatments. Active monitoring is quite different from giving a placebo. In comments on an earlier post on this blog, I offered criticism of ITT analysis. However, on reflection, it seems that ITT analysis is quite reasonable in this case.

      Certainly, if one wished to do an as-treated analysis, one would have to figure out how to classify those in the active-monitoring group (61.1% after 15 years) who eventually received radical intervention.
      Bob76

      • I am not sure how to interpret the people who were under active monitoring but underwent a radical treatment within 4 months. The practical reality is that to receive the radical intervention within 4 months means that the patient must have elected treatment almost immediately upon entering the study. The “active monitoring” wasn’t likely to have played any part in their decision. I’m speculating here, as I am not a medical professional although I have some personal experience with these choices. So, my guess is that these are people that decided they were better off getting treated despite being assigned to the active monitoring group. So, I’m not sure how they should be treated in the analysis – though it seems like a good idea to explore what difference it makes and/or apply a more complete model as Andrew suggests.

        Again, the appropriate treatment of these people might depend on how the model is to be used. From a patient perspective, it might be different than from a public health perspective. If I, as patient, were just diagnosed as having suspicious lesions and given the choice of active monitoring (which will probably mean a biopsy before long) or radical intervention, the random assignment results for people that crossed over might not be of much interest. It just isn’t clear to me that ITT is right, though the as-treated might also not be appropriate.

      • > Active monitoring is quite different from giving a placebo.

        There was a similar confusion in some reactions to the colonoscopy screening study discussed recently. Not screening doesn’t mean that the colonoscopy cannot be used later for diagnosis in the “standard of care” group.

        https://statmodeling.stat.columbia.edu/2023/05/16/colonoscopy-corner-misleading-reporting-of-intent-to-treat-analysis/

        In this case there was already an initial diagnosis but people may decide to do something different from the treatment assigned at different points in time and for different reasons. (The article mentions that “Decisions to change the management approach in the early years were often made without evidence of progression, which probably reflected anxiety on the part of either the patients or their physicians.”)

  9. I admit that I don’t do medical treatment analysis on any kind of regular basis, but it seems to me that ITT analysis is the kind of thing one does when one has “Frequentist estimators” rather than generative models. The right model is enroll person -> randomize to treatment group -> person makes decision about actual treatment -> actual treatment occurs -> outcomes observed

    In order to analyze this you need to have a model for how people make decisions. You would hope that people actually record sufficient data on the person that you could estimate such a model, but my impression is people waste a lot of money and time doing these studies without the slightest notion that they should be estimating a generative models and hence don’t collect that data. Examples would be surveys of patient symptoms and their severity, having patients rate their quality of life, having patients rate their anxiety about having a procedure done or about NOT having a procedure done, etc.

    So we spend a lot of money and time, generate a bullshit dataset based on zero hypothesis or model and feathering the nest of people who run the studies, and eventually have nothing that qualifies as usable and anything we do have is pseudo proprietary so that the study authors can continue to collect rent on it by reanalyzing and following up and preventing any “competition” or critical reanalysis.

    • I tend to agree. Adherence is under-studied in the medical statistics field, in my opinion. I think the people that like ITT and highlight the selection problems with as-treated analysis implicitly assume that the patient’s decision to follow or not follow the randomization reveals important information about their condition – such as they cross over to a radical intervention because they really don’t feel well. On the other hand, a preference to ignore the initial randomization and focus on the actual treatment chosen seems to implicitly assume that a person chooses a course of treatment based on personal preferences unrelated to their actual medical condition (such as fear of surgery or the opposite – for example, “just get rid of my prostate so I don’t have to worry about cancer”).

      In both cases, what is missing is a model that treats the individual choice whether or not to follow the randomization as something that needs to be analyzed. Without that, I don’t think we can decide which approach is most appropriate. There may be enough data on the individual patients to make that decision part of the analysis. If not, then I think the analysis should be done both ways (perhaps even more than two ways – in the prostate cancer example I think crossovers that happen shortly after randomization might be different than those that happen after a year or two of active surveillance). And, of course, I’d prefer the data to be made available so these alternatives can be explored by those that care.

      • If you at least collect some survey information about the patients experience and state of mind you’d have some chance of building that model. My impression is most of the time this data isn’t even considered for collection much less collected and categorized etc.

        >people that like ITT and highlight the selection problems with as-treated analysis implicitly assume that the patient’s decision to follow or not follow the randomization reveals important information about their condition

        And so… they throw that information away! wth?

        I’m with you 100%, a patient crossing over is information, throwing it away is not helpful. If only because it lets us highlight a source of uncertainty which should then increase the uncertainty in the final analysis but would not in an ITT or any other frequentist analysis.

    • If you had a model of how people make decisions maybe your “right model” wouldn’t need to include the “randomize to treatment group” part at all.

      • Maybe, but then again it depends on the model. One component of the model might be people trying to discourage them from altering from the randomized treatment course because of the problems it causes in the study for example. Your model might suggest that in general people are much more likely to refuse to stay with the “monitoring” group because people feel like “doing something” is better.

    • From a health services policy/business perspective – the ITT analysis seems to make some sense. If you have a great (as treated) but very expensive treatment that no one manages to follow-through with you might make one policy. If you have a modestly effective treatment that people succesfully take, you might make a different policy. That’s a different target audience than the person making the (seemingly simple) decision to go through with a given treatment or not.

      I think you may underestimate the difficulty in the term “model for how people make decisions”. As someone who self-admits to “[not doing] medical treatment analysis on any kind of regular basis” you may be falling into a common trap of thinking everyone else’s job is easy (I think there’s a name for this). In my experience, the idea that such a model is remotely plausible based on available data is utter fantasy.

      • I agree with the difficulty of actually modeling whether or not people adhere to the protocol. But I don’t agree with your first paragraph. In the case of the colonoscopy study, the effectiveness for those that went through with it (took up the “invitation” to be screened) was strong enough that the real health services policy question is how to get more people to get screened. If you simply assume they won’t go through with it, then you get the misleading impression that it is not very effective, and as a result, you don’t put resources into the more manageable question of how to encourage people to follow through.

        In some other cases, you are absolutely correct that lack of follow-through is a major impediment to treatment effectiveness and will not be feasible to overcome. I think this particularly applies to drugs with severe side effects. For example, many anti-depressants have bad side effects that will make adherence low, and it isn’t obvious that much can be done about that. As a result, the effectiveness should take into account the actual behavior of those prescribed the drugs.

        The point should be that such details should affect the analysis – that one size does not fit all. Simply saying ITT is the right way to analyze all such situations misses the point that adherence is something that needs to be taken seriously, perhaps even modeled.

      • The point isn’t that modeling adherence is easy, it’s that if you aren’t modeling adherence you aren’t doing the job. It’s like saying sculpture is hard so when you hired us to make a sculpture for your garden we just delivered a block of granite and drove off…

        Even at the level of building a simple linear logistic regression of adherence vs random assignment group and pretreatment survey results plus uncertainty you would be doing vastly better. But if you don’t do the pretreatment survey you haven’t even started to address the issue. The psychology and symptomology and such are part of the process, so they need to be part of the model.

        It being hard is admittedly a difficulty for career advancement but not for science, because science cares about what we learn from an experiment. If the answer is “not much” that is the correct answer, not to pretend that you’ve now got the answer.

        ITT is just a wrong answer. Noone disagrees with the causal chain I listed above. We all know people are recruited, they get a random assignment, they decide to adhere or not, they get the treatment they decide to do, and then they get some outcome. Their random assignment is relevant to their choice of treatment but not determinative.

        Suppose I did a mouse experiment and randomized mice to different surgeries and then told the surgeons to do the surgery they were randomized to unless they see conditions A,B,C in which case do a different surgery, and then analyzed via intention to treat. That would just be disingenuous. The fact that we don’t know a precise rule for adherence doesn’t mean adherence isn’t part of the question. So build a probabilistic rule for adherence!

        • The whole “it’s too hard” gets me every time. Usually we have the following:

          1) no one ever even tried
          2) usually no one has any bayesian background at all, so when they say “it’s hard” they mean “I can’t look in a statistical manual and find the commands to put into SAS/Stata/R”
          3) There is no one from psychology involved in the research project at all
          4) No one even thought to ask psychologists or people who study decision making like marketing people
          5) They never even thought to collect any information relevant to the task
          6) No-one understands the purpose of doing this stuff… it’s not because you know how to get the right answer so why didn’t you just program it in? it’s because if you don’t do the adherence model you don’t get anything like a reasonable **quantification of uncertainty** and you therefore get a VASTLY overinflated sense of certainty about the outcomes/effectiveness
          7) The vastly overinflated outcome certainty is the point, and a politically desired outcome. no one wants to say “we did a study and found that it’s nearly impossible to tell anything about how well we treat cancer” even if that is 100% the correct answer.

          Something as simple as

          adherence ~ Binomial(inv_logit(k * symptom_severity * assigned_monitoring + random_individual_factor))

          with random_individual_factor having a group level prior distribution would be a vast improvement over what is done and you’d get an estimate of the probability of adherence, also you can fit probability of good outcome based on adherence with symptom_severity informing, telling you something about whether outcomes are due to selection or severity.

          This is just as a pointer in the direction of what proper research would look like. Obviously fitting a variety of models and thinking hard about the problem would be needed. But fatalistic “this is impossible” is just laughably bad.

        • Bob76 thanks for the reference, but this paper appears to get it wrong again. The ITT gives the answer to “what is the *average* causal effect of *randomizing* someone to the treatment when adherence follows the same rules as seen in this study”.

          There’s nothing “unbiased” about it because adherence isn’t a random number generator, and furthermore the answer it gives leaves us with no mechanistic estimate of the effect of the actual treatment itself.

          The correct model is unambiguous, we all agree on it right?

          1) Someone runs an RNG and randomly assigns patients
          2) patients discover their assignment and make a decision with the assignment being one factor
          3) patient gets treatment patient decides on
          4) outcome is observed

          The model has 4 groups:

          1) group assigned to treatment, takes treatment
          2) group assigned to control, takes treatment
          3) group assigned to treatment, takes control treatment
          4) group assigned to control, takes control treatment

          We observe the adherence fraction **exactly once** so you’re extrapolating to adherence in other contexts on the basis of **one** data point and claiming to have an estimate of something useful at a population level. That’s based on a strong and implicit prior that adherence in the future will be tightly constrained to be similar to adherence in this trial and offers **no** individual level decision value at all. A person who individually wants to decide what to do gets **zero** benefit from the analysis if you use an ITT because it’s impossible to distinguish between rates and causes of adherence and effectiveness of the treatment.

          The medical industry thinks their only alternative is “As Treated”. But the SAME thing is true in a different way for As Treated. Because we treat a mixture of those assigned randomly and adhered, and those assigned to control who didn’t adhere. We don’t know why someone would choose to adhere vs not adhere, and we don’t have covariates that could help an individual decide whether their reasons for wanting the treatment they prefer are similar to the reasons that people in the study had.

          The problem is that the medical literature sees two possible methods of analysis: average effect based on ITT, vs average effect based on As Treated.

          This dichotomy is just WRONG, both of them get the wrong causal generative model, neither analysis is correct! The correct model models the 4 groups and the conditional probability to be in the 4 groups as a function of pre-treatment covariates, including both psychological covariates, cultural covariates, as well as symptomatic and disease progression knowledge.

          A generative model with survey of the patients prior beliefs about the treatments, about the severity of the disease, pressures from family members, subjective experience of the symptoms, prior belief about the effectiveness of the treatments, fear of surgery and anaesthesia, fear of side effects, fear of pain… those are absolutely necessary to get us an individual level predictor.

          And individual level predictors are the **only** ones that offer us a hope of extrapolating to the general population accurately because we have *zero* guarantee that in other populations adherence and its reasons are similar.

        • > 7) The vastly overinflated outcome certainty is the point, and a politically desired outcome. no one wants to say “we did a study and found that it’s nearly impossible to tell anything about how well we treat cancer” even if that is 100% the correct answer.

          It’s funny because the problem of most people with the reporting of ITT results in medical trials is that they don’t like the “we did a study and found that it’s nearly impossible to tell anything about how well we treat cancer” and they’d rather be told “we’re doing great” even if the certainty about that outcome may be vastly overinflated.

        • > The medical industry thinks their only alternative is “As Treated”. […] The problem is that the medical literature sees two possible methods of analysis: average effect based on ITT, vs average effect based on As Treated. […] This dichotomy is just WRONG

          This also seems funny – but maybe I’m misunderstanding what you wrote…

          In any case, there are _three_ main methods of analysis (as the article linked mentions): intention-to-treat, as-treated and per-protocol.

        • But Carlos, the reason they don’t know what’s going on in ITT is precisely because they didn’t make a model of what’s going on, not because they couldn’t have figured it out from an appropriate analysis.

          As I say elsewhere, the generative causal model is pretty clear, there are 4 groups, there are various different reasons for crossover, and if you don’t model any of that you just pretend that treatment assignment is real randomization, you get a wrong answer that’s automatically confounded with a known confounder (namely, the patient’s decision). No wonder you can’t figure it out! And “As Treated” just assumes strongly “patient decision making didn’t matter”. So they’re both wrong.

        • your discussion of per-protocol did not appear when I was writing my reply.

          Suppose we have the following RCT for seizure treatment.

          1) Half of the people will be randomized to stab themselves in the eye twice a day with a pencil
          2) Half of the people will be randomized to take an anti seizure medication.

          Of the people randomized to (1) 100% are non-compliant and take the anti-seizure medication

          In the final outcome 90% of both groups experience some reductions in the frequency and severity of seizures.

          ITT analysis: There was no statistically significant difference between stabbing yourself in the eye with a pencil vs taking anti seizure medications p = 0.84

          As Treated analysis: Seizure medication reduces seizures in the period after taking it when compared with the prior period intra-patient p=0.048 larger sample sizes are recommended, particularly of eye-stabbing.

          Per-Protocol analysis: No data was collected on the control group, among the half of people who were assigned to the medication there was no statistically significant reduction in seizure rates p = 0.055

          Bayesian analysis: Eye stabbing was strongly resisted by patients due to fear of pain and permanent injury, no amount of officious annoying MD coaxing could change the patients unreasonable prior bias. All patients chose anti-seizure medications, the medications dramatically reduced incidence of seizure among patients except those whose seizures were caused by brain injury from previously taking high doses of methamphetamine. Bayesian analyses are to be taken as mere opinions given by the one guy on our team who doesn’t have a PhD, the analysis was originally withheld due to its highly improper subjective nature however we were forced to reveal it by reviewer number 2. No data will be provided to anyone who wishes to reanalyze the data in any way other than our favorite analysis which is ITT.

          Title of paper: Eye stabbing an effective an cheap intervention in seizure control, a gold standard Randomized Controlled Trial

          /sarcasm

        • Daniel, almost nothing is true in your haughty, condescending, angry rant number 1,010,313,131. The rant I am referring to is copy/pasted below. Your tiring shtick of demeaning people without a strong Bayesian background, over-simplification of a problem, and aggressive attitude are very tiring. I do agree with your point #7, though.

          The fact is, there are plenty of excellent Bayesian statisticians who do excellent work in health outcome modeling (especially cancer).

          There is a lot of information to consider that you ignored in your proposed solution. Initial participation in a screening strategy is important, and your proposed formulation (in a different post) , or something like it, could be used for a start, but the majority of an individual’s screens will occur longitudinally. For example, there are 26 opportunities to take an annual FOBT test for CRC screening for screening from age 50-75.

          There isn’t good information to model this stuff at a population level. Behaviorally, it’s believed there are folks who will adhere perfectly, folks who will adhere imperfectly, and folks who will never adhere. Modelers do not know what this distribution looks like. Additionally, at an individual level, it’s been observed that adherence changes as a function of a prior test result–for example, adherence over time declines after multiple, sequential, negative screen results. This has been referred to as ‘screen fatigue.’ It’s also been observed that initial participation rates increase as a function of age, and there is usually a large bump once an individual reaches Medicare age. There are also different adherence levels to diagnostic follow-up, conditional on a positive screen result. There are many other types of relationships that exist.

          At a population and decision-making level, it gets even more complicated. What’s the correct analytical framework to compare different, competing tests, taking adherence into account? If 40% of the population is adherent to colonoscopy and 50% to FOBT, is it fair to compare them head-to-head by plugging in those numbers into a computer simulation? Maybe not, because lot of FOBT takers may never want to be screened with a colonoscopy! In other words, the 50% FOBT group may consist of 80% who would never want to be screened with a colonoscopy. These are different, non-overlapping groups of people!
          _______________________________________________________
          “The whole “it’s too hard” gets me every time. Usually we have the following:

          1) no one ever even tried
          2) usually no one has any bayesian background at all, so when they say “it’s hard” they mean “I can’t look in a statistical manual and find the commands to put into SAS/Stata/R”
          3) There is no one from psychology involved in the research project at all
          4) No one even thought to ask psychologists or people who study decision making like marketing people
          5) They never even thought to collect any information relevant to the task
          6) No-one understands the purpose of doing this stuff… it’s not because you know how to get the right answer so why didn’t you just program it in? it’s because if you don’t do the adherence model you don’t get anything like a reasonable **quantification of uncertainty** and you therefore get a VASTLY overinflated sense of certainty about the outcomes/effectiveness
          7) The vastly overinflated outcome certainty is the point, and a politically desired outcome. no one wants to say “we did a study and found that it’s nearly impossible to tell anything about how well we treat cancer” even if that is 100% the correct answer.”

        • Unanon, anyone who’s making an effort to build a model based on a description of a generative model is not included in any way in my annoyance. Yeah, I get it there’s good people making efforts to do stuff. But it seems like this is far from “normal”.

          I don’t read cancer biology daily, but I **have** participated in cancer biology research. The very first thing I discovered when participating in cancer biology research was that the cancer biologists and their colleagues who I was working with used Kaplan-Meier curves to compare different groups of tumors in terms of the aggressiveness and survival differences. So we started building a Bayesian model of survival as a function of some genotype scores and got REALLY different predictions from K-M curves, like so different it made no sense. K-M curves for cancer X showed years or decades to median death, and Bayes showed months or something like that. Which of course we assumed meant that we had bugs in our code… we worked at that for weeks… until we looked carefully at what was going on and sure enough we could show that the Bayesian model was correct and the incredibly different results we saw was due to just terribly inappropriate assumptions built into the K-M curves. We even built simulations to show how it occurred to convince ourselves.

          I’ve read RCT analyses because of personal interest in health issues, such as cancer or treatment of allergies or treatment of other things that have affected either myself or someone I know. I have NEVER seen anyone model compliance, I have almost never seen a Bayesian model involving a generative model of anything. Does that mean no-one does it? Of course not, there are millions of people writing papers any given year. I can say though that if you select maybe 50 of them from among things someone like me would be interested in, the chance of anything being done other than plugging and chugging into canned analyses is very low. I’ve seen non-inferiority analysis of accupuncture for allergies, I’ve seen use of surfactants for sinusitis, I’ve seen cancer, heart disease, vitamin D, various treatments for COVID, etc. The number of Bayesian models I’ve seen in medicine other than my own is pretty much exactly 1, and that was the Pfizer COVID trial and was Bayes pretty much in name only if I remember, it pretty much just added a prior to what was otherwise a frequentist style model if I remember correctly (though honestly I haven’t looked at it in a while, I just remember thinking that it had very little of real interest).

          I said something the other day in another forum, it was basically that there’s a huge disconnect between what I say and what people hear. When someone like me says “the problems with these analyses would be solved by a Bayesian model” what other people hear is “take the analysis just like this and add a prior to it”. But what advocates like me advocate isn’t anything like that, its **write down a theoretical generative model for how the data comes about, write down hypothesized forms for the functional forms, treat probability **as a credibility measure not a frequency** and then reanalyze entirely from scratch** It’s a completely different paradigm.

          I’d be happy to have you show me 4 recent papers in Cancer biology where they lay out a generative model for their data and model probability as credibility not frequency. Please, send the links!

        • ” Please, send the links!”

          I don’t read Cancer Biology, and I’m not sure why I need to curate publications for you, given your intelligence and ample time to post on message boards, but sure, here are some that I’ve read:

          Colorectal cancer:
          https://pubmed.ncbi.nlm.nih.gov/20076767/
          https://pubmed.ncbi.nlm.nih.gov/21127321/

          Breast cancer bayesian simulation model (MDACC): https://resources.cisnet.cancer.gov/registry/packages/bayes-mdacc/

          Lung cancer:
          https://pubmed.ncbi.nlm.nih.gov/30518460/

          Regarding your Bayesian survival models–this idea has been thought of before:

          Read
          Gustafson “Flexible Bayesian Modeling for Survival Data”
          Abrams “A Bayesian Approach to Weibull Survival Models”
          Osnes “Spatial Smoothing of Cancer Survival: A Bayesian Approach”

          Cancer incidence:
          Mezzetti “A Hierarchical Bayesian Approach to Age-Specific Back-Calculation of Cancer Incidence Rates”

          And so on.

        • I checked the first two CRC papers. Neither was quite relevant as they fit to observational data while the discussion was about RCTs.

          But anyway the first paper bugged me because they don’t even show a graph comparing predicted curve to data. Like age-specific incidence or mortality. That is kind of a bare minimum requirement for me, so I moved on.

          The second one did include this but was kind of strange. They use *a lot* of space explaining the details of metropolis-hastings for some reason. The model had decent fit, but it also had 20+ parameters and was more empirical than generative.

          I still can’t tell if figure 8 shows actual predictions or if the model was trained on that data. This is weird since like half the paper is spent explaining metropolis-hastings, maybe they tell us that important detail somewhere in those pages.

          I don’t see why it is so rare to simply plot the data along with your model fit. Having a model capable of fitting the data is only the first step though. Then you have to check whether it can predict new and other types of data. If it can, you may have something interesting/useful.

          Then it is time to compare it to other models that can pass those tests. Ie, which are simplest, have best predictive skill, and are derived from the most plausible assumptions?

        • Unanon, thanks for sending the links. The first article was an honest attempt to bring bayesian ideas to patient-level process models of cancer progression. I think it’s very worthwhile effort. I do agree with Anoneuoid that they really could have used some basic graphs. I didn’t have access to the second one.

          I certainly don’t think bayesian survival models were something I invented! I just don’t see them as something anyone does at any scale. It’s not surprising though because when I get involved in research with new people they very rarely have the slightest idea what Bayesian stuff even means. I can’t blame them, education doesn’t include any discussion of Bayesian methods usually. When it does it’s often just a couple textbook examples about diagnosing a rare disease, or adding a prior to an essentially classical framework.

          The most important part about using Bayes isn’t the Bayes at all, it’s the **process model** that describes what you think is going on. That’s where the *science* is. Bayes is just the way you inform that model with data. The first article you sent with the micro-simulation was an attempt to do that! kudos to them.

          I was sent the following paper this morning, which discusses how RCTs should define what it is they are trying to estimate.

          https://bmcmedicine.biomedcentral.com/articles/10.1186/s12916-023-02969-6

          Running ITT or per-protocol or as-treated or whatever is basically just a way to turn data into publications, its not an honest attempt to learn from data.

          When I say it’s not an honest attempt I don’t mean that people are being dishonest, like deceitful, I just think they are usually doing “what is done” and are unaware of the problems, because they haven’t thought through the point of their analysis (ie. in the paper above “the estimand”) the analysis is more or less meaningless as we don’t really know what it tells us.

          Also people really want randomization to be magical causal identification. In the context of an RCT where patients have choices, the randomization is really just a form of suggestion “hey the randomizer suggests you do A” and offers no real causal identification power at all. Ignoring that is hugely problematic.

          I’ve encountered this in the context of engineering. For example in discussing which locations to sample for evidence of deterioration or concrete cracking or whatever, sometimes people want me to create a randomized sample plan, and then they want the ability to deviate from that plan when they are in the field and see problems. If the randomization is just a suggestion and the field people can do whatever they like after hearing that suggestion, the data means something completely different from what it would mean if you did in fact randomize. For some reason no-one ever wants the ability to deviate from the random sampling plan to in detail document exactly how there were no problems at all at site X. But they do want that ability when they see a crumbling pile of garbage on their way to the randomly chosen site.

          I tell them when they collect extra data we’ll use it to conditionally describe the kinds of damage that can be found in the worst cases, and of course it’s a finite population, so it goes to estimate the conditions at that one location, which can be a big deal when you’re talking about 10 or 20 buildings, but it won’t have the same informational value as a random location.

          If people are designing RCTs where they *know* the patients will often deviate from the randomized treatment, then they **need** to model that process, or they don’t have much of anything at all in terms of information. People find that easier to understand in the context of the engineering problem. “oh yeah, the inspector wandering around looking for problems will tend to find them and spend all their efforts there” they find it much less satisfying when it comes to medicine, because they think “we can’t force the patient to do what we want, so what should we do just throw our hands in the air and not do medical research?” If your whole analysis schema is built on taking sample averages and “extrapolating to a population” you have no basis on which to do anything about the non-randomness problem.

          A lot of medical research is just wrong. That’s not even a controversial opinion from a jerk like me on the internet, that’s like the subject of a bunch of articles! https://link.springer.com/article/10.1007/s00192-017-3389-1 and citations there for example.

          I think that’s really a problem. We spend hundreds of billions of dollars annually on stuff that’s *wrong*, in some cases worse than if we’d done nothing, and actually harms patients.

          I don’t criticize this stuff vehemently on the internet because I’m a jerk, I do it because I care deeply about the content of society’s collective knowledge. When so many times after you’re asked to look at something you find a major methodological flaw making the publication either meaningless or mean exactly the opposite of what the authors claim it can be extremely disconcerting.

        • > the reason they don’t know what’s going on in ITT is precisely because they didn’t make a model of what’s going on, not because they couldn’t have figured it out from an appropriate analysis.

          That’s the point. It’s not the desire for a vastly overinflated outcome certainty that push people for ITT analysis – if that’s what they wanted they would be better served asking you to figure out an appropriate analysis (or getting the overinflated outcome certainty on the own doing the as-treated or per-protocol analysis).

          > As I say elsewhere, the generative causal model is pretty clear, there are 4 groups, there are various different reasons for crossover, and if you don’t model any of that you just pretend that treatment assignment is real randomization, you get a wrong answer that’s automatically confounded with a known confounder (namely, the patient’s decision).

          Maybe I’m misunderstanding you again… Treatment assignment is real randomisation. That’s the reason for doing an analysis based on treatment assignment – instead of actual treatment – despite the obvious reasons against doing so.

          However in your “four groups” model there will be confounding because patient’s decisions are not independent of the potential outcomes with and without treatment. There is no straight-forward way to compare these groups.

          You listed a few things that are absolutely necessary to get us an individual level predictor. I’m not sure if it was implicit that you also need some not-so-pretty-clear assumptions about the correlations between the (partial) self-assignment of individuals into those four groups and their individual expected outcome with and without treatment.

          (Aside: In general intention-to-treat is not strictly the same as treatment even if compliance is total. Knowing that there is an intention to treat can have an effect on outcomes distinct of the effect of the actual treatment. That’s one of the reasons to have blinding – trying to simplify things as much as possible so they don’t get “too hard”.)

        • Carlos, wow it’s hard to reply to stuff on my phone. I hope this makes it to the right place.

          What I intended to say was one reason people do something other than a bespoke analysis is because they don’t like the idea of having to say “here’s some stuff we hardly have any idea about, we have to put that in our model but we haven’t got a clue about it” and then have it lead to a crazy large uncertainty interval

          You see this a lot. elsewhere in this thread one of the commenters says something about how people haven’t got any idea about the distribution of various individual characteristics etc. They don’t know, therefore rather than making a model that explicitly says “we don’t know this” they stick with “standard” or “approved” types of analysis that don’t have those issues even included in the model.

          It’s also confounded by people thinking that probability distributions are physical quantities in the world. So they feel there’s literally nothing they can do, since they don’t know the formula for the pdf… Throw up your hands…

          When it comes to per protocol, the randomization doesn’t really get you much if compliance was not good. You discover the conditional distribution of outcomes among people who are like the ones who comply, but you don’t discover the causal effect because the people who comply with the alternative treatment B are not necessarily similar to the ones who comply with the treatment A

          If you want to estimate the treatment effect of doing A instead of B you need a group of people doing A who are similar in all relevant ways to the ones doing B. Randomization followed by noncompliance does not get you that.

          If you model the individual decision it leads to a bunch of uncertainty unless you have a decent model, but it gives a straightforward way to estimate the causal effect by difference between people who are relevantly similar.

          Note that I’ve known a bunch of people who have had prostate cancer diagnoses, friends of family etc. Most of them made a decision on pretty straightforward decision criteria, things like asking their oncologist what they would do if it were them, and asking their friends what they did facing similar issues and whatever. I think you’d have a reasonable chance of making an ok model with a decent survey and discussion with patients.

        • > When it comes to per protocol, the randomization doesn’t really get you much if compliance was not good. You discover the conditional distribution of outcomes among people who are like the ones who comply, but you don’t discover the causal effect because the people who comply with the alternative treatment B are not necessarily similar to the ones who comply with the treatment A

          Sure. When your analysis is not based on the randomized groups the randomization doesn’t get you much in per-protocol or as-treated analysis including the one you propose. When you deviate from the randomization you can’t discover the causal effect without an additional assumption of exchangeability. This is an observational study now: finding the right covariate adjustments becomes an absolute necessity.

          > If you want to estimate the treatment effect of doing A instead of B you need a group of people doing A who are similar in all relevant ways to the ones doing B. Randomization followed by noncompliance does not get you that.

          Sure. But on what basis do you presume that your model gets you that? If two subjects are similar why would one do A while the other does B? Something causes that difference and supposing that this something is irrelevant for the outcomes is a strong assumption. (Not to mention the practical difficulty of defining similarity when you have more than a handful of covariates and no two subjects look the same.)

          > If you model the individual decision it leads to a bunch of uncertainty unless you have a decent model, but it gives a straightforward way to estimate the causal effect by difference between people who are relevantly similar.

          Sure. If you have a correct model it’s straightforward to estimate causal effects from observational studies. Why would anyone waste time and resources with experiments then? Having the correct model that gets you that is not as straightforward and pretty clear as you seem to imply though.

        • Carlos, I don’t mean to imply that getting a good model is easy, only that it is the only way to get the answer. A huge amount of science is using statistics to launder the fact that they have no hope of addressing the question. Utilizing “named” procedures with hundreds of references that have been utilized by others before is part of the laundry process.

          In these cases, the right answer is to build the individual decision model, then it’s got a bunch of parameters that you have broad and largely useless priors over, and the collected information lacks informativeness… So the final answer is that we have gotten hardly any closer to an answer than we were at the beginning, despite spending $14M or whatever. It’s not an answer people LIKE but it’s the intellectually honest truth. Utilizing the “well known” ITT analysis or per-protocol or whatever is just a way to say “don’t attack us we are doing the accepted stuff”

          You WOULD have a better chance at information if you did some surveys and collected info relevant to the question though. That’s pretty rare among biomed articles I’ve seen. But I don’t read them all day.

          The fact is, the social construct of the science industry has built self protection schemes.

          Similarly, if I ask aunt Mabel who would win the last presidential election she could have told me it would be 50/50 trump or Biden +- 5 percentage points. Taking a large number of telephone polls with 10% response rates etc got us essentially no closer to the truth than aunt Mabel. A good model would have had some kind of time series willingness to respond bias model and that would have been hard to identify, so it would lead to smearing out the prediction so in the end after hundreds of polls we know barely more than 50/50+-5

          I think sometimes people hear me say we should do such and such type of analysis and they think I’m saying it would be easy to extract information where no information exists. Au contraire, it would be easy to see from the analysis that in fact no information is there to extract. That would be the right answer!

    • I don’t think it’s feasible to measure all patient variables which would be needed to model medical decision making. Even if we somehow could measure all relevant variables, the resulting adjustment set might be very high-dimensional, so we would need an extremely large number of observations to get good estimates. And all this is assuming that you are able to specify a good model in the first place, which is not at all obvious.

      The whole point of randomisation and intention-to-treat analysis is to avoid these very difficult challenges.

      • I think you overlook the biases that exist in training of medical professionals. The “science” makes it difficult to take adherence seriously to study, yet all practitioners know that adherence is a real issue. But it requires understanding the individual patient and more than just their medical condition. Admittedly, it is a difficult thing to predict, but clinicians do think about it and take it seriously. “If I prescribe drug X, then how likely is this patient to follow the prescription guidelines?” But it is difficult to try to answer that question – for one thing, the time involved is probably not reimbursed by the insurer. For another thing, it wasn’t covered in medical/nursing school. For another, it involves collecting information far beyond what usually appears in the medical record. All of which present real practical difficulties. But as Daniel keeps saying, ignoring the difficult is a poor scientific practice.

        (I do believe there are circumstances in which ignoring modeling adherence can be justified – it would be under circumstances where the model would be particularly difficult to collect relevant data for and/or where understanding adherence probably doesn’t matter much. The reasons I focused on the colonoscopy and prostate cancer cases is because adherence is so clearly relevant and something somewhat tractable).

        • “This thing is hard so we’ll do something wrong” is not a good answer ever.

          “This thing is hard, and we have strong reasons to believe it doesn’t matter much so we’ll do something approximate” is OK, but as you say, that’s not what’s going on in colonoscopy or prostate cancer.

          People also seem to naively believe that they have to somehow model “what’s going on in the head of the patient” as opposed to “how what the patient says in a survey should alter our probability assigned to whether or not they comply, and whether or not they have an underlying condition that would strongly affect their outcome”.

          A probability model for patient decision making is hard, relative to pushing a button on SAS, but it’s hardly quantum mechanics of their brain.

          For the prostate cancer model, to get started, let’s imagine we ask the following questions:

          1) What is your age
          2) What is your biological sex
          3) What race do you identify as (set of choices)
          4) What religious affiliation do you affiliate most with … (set of choices)
          5) On the following scale rate the discomfort that your condition causes… 0,1,2,3,4,5,6 with some text descriptors from “none” to “severe discomfort”
          6) What is your education level (from some grade school to PhD, and a check mark for whether it’s a biology/medicine related degree)
          7) Check all that apply: what are the primary motivations for entering the study (things like “to get free treatment” and “to improve the state of knowledge about the disease” and various other things)
          8) What is your income level?
          9) Do you have insurance that would cover this treatment outside the study? What is the level of coverage?
          10) What is your level of concern about having surgery? (from zero to “I am very nervous about having surgery”)
          11) What is your level of concern about potential side effects (similar scale)
          12) Which side effects are you concerned about (check boxes)
          13) What are your existing thoughts about watchful waiting (0 I believe it is not right for me, up to 6 watchful waiting is my currently preferred treatment)
          14) What is your current belief about the severity of your condition: similar
          15) What is your current belief about the aggressiveness of your condition: similar
          16) What is your current belief about metastatic tumors: 0… I have no reason to believe… 6 … I have medical biopsy proving metastatic

          Maybe about 2-3 other questions.

          Now, hypothesize at least the *direction* that each of these affects the probability of compliance with surgery assignment, and with watchful waiting assignment.

          Place a prior over coefficients of each one with bias towards the direction hypothesized in a logistic regression with nonlinear response on any variables that may seem appropriate.

          Now, hypothesize the direction with which some subset of these predicts actual aggressiveness of the underlying condition, for example using the patients own beliefs, using the reported level of symptoms, using info about metastatic condition, etc.

          Place prior over coefficients in a hidden variable model for severity… again with nonlinear response if necessary.

          Collect data…

          Posterior distribution of parameters…

  10. The first article was an honest attempt to bring bayesian ideas to patient-level process models of cancer progression. I think it’s very worthwhile effort. I do agree with Anoneuoid that they really could have used some basic graphs. I didn’t have access to the second one.

    I looked closer at it. Shouldn’t the model include parameters that can be checked some other way like number of cells, division rate, and genetic error (eg, mutation) rate?

    Say the human body is ~10^13 cells and (guesstimating from organ mass) the colon is about 1% of that (~10^11 cells). Then those cells need to all be replaced daily for 100 years. That would require a minimum of log2(365*100*10^11) ~ 50 cell divisions (ie, times the DNA has been copied since the egg was fertilized). That is consistent with the Hayflick limit of ~ 60 divisions before fetal cells stop dividing in culture. So let’s assume that, in one way or another, the body comes close to attaining this minimum.

    On average the DNA is being copied at a rate R of ~0.5 times per year (the rate is probably much higher in childhood and lower in the elderly, but keep it simple for now). Then assume most genetic errors occur during the copying process, and there is redundancy so a cell needs to accumulate n errors before it becomes cancerous. That is essentially the Armitage-Doll model.

    We don’t know exactly why there is a copying error during some divisions but not others, so at this point we go to probabilistic model.

    If there is an approximately constant probability p of an error each division (independence), and no possibility of correcting the error once it is established in the daughter cell, then after d divisions the probability of one mutation happening will follow the geometric distribution. Ie, 1 – (1 – p)^d. Then the probability n mutations accumulate in the same cell will be probability of the first *and* second *and* third, etc. Ie, if p is constant you multiply each. So raise it to the power of n: (1 – (1-p)^d)^n

    From earlier we got that the DNA is copied anew about once every two years. That means we should see cumulative cancer incidence follow a curve of (1 – (1-p)^(r*t))^n, where r ~ 0.5 and t is years of life. Then the age-specific incidence will be the first derivative, or finite difference, of the cumulative function. This would give the actual incidence, but not necessarily the observed incidence.

    I compared to the data in Unanon’s 2nd paper. They reported a sensitivity for the screening of ~0.4 and divided the incidence up into four categories (which you can roughly combined mentally). After fitting by eye and was able to get near the same performance with *only two* free parameters. In R:

    p = 0.05
    r = 0.5
    t = 1:100
    n = 10
    p_detect = 0.4

    p_cancer = (1 – (1-p)^(r*t))^n

    plot(t[-1], 10^5*diff(p_detect*p_cancer), panel.first = grid(),
    xlab = “Age”, ylab = “Incidence (per 100k pop)”)

    https://i.postimg.cc/MGfYPWxm/crc.png

    There are a lot of questionable assumptions being made (independence of mutations, constant cell division rate, negligible lag between carcinogenesis and detection, etc), so I do not take these parameters seriously. But in principle all the parameters have a meaning that can be checked by some other method, and it captures the main patterns in the data. Ie, very low incidence until 30-40 years old, inflection point around 60 years old, plateau at around 80 years old, and peak rate of ~400 per 100k pop.

    So, I don’t believe we need these models with 20+ free parameters (that can’t even be measured via some other method) to explain this data.

    • Also, here is the paper that model was partially based on:

      The mature differentiated cells of the colonic crypt are ultimately the product of a single stem cell situated at the base of the crypt [13–16]. The simplest model envisages a stem cell that divides daily or every other day. The stem cell produces two daughter cells. One daughter becomes the new stem cell. The other daughter undergoes successive mitotic divisions through approximately eight generations to produce 256 mature cells that populate the crypt for one or two days. The cycle then repeats; the new stem cell divides to produce daughter cells, one is the new stem cell and the other divides to produce 256 mature cells that displace the previous set. At age 80 years the stem cell at the base of the crypt will be 30,000 mitotic divisions from the primordial stem cell which in turn is at least 30 or 40 mitotic divisions from the zygote.

      […]

      The hierarchical model is based on the assumption that the number of mitotic divisions that separate the zygote from stem cells and mature differentiated cells in the adult is reduced to the absolute minimum [17,18]. This will reduce the mutant frequency in stem cells and greatly reduce the risk of malignancy. The minimum number of cell generations is easy to calculate, it is simply log2(N), where N is the total number of cells produced in a human lifetime. This number, N, is also equal to the total number of mitotic divisions that occur in a human lifetime. But the key question is how many times has the DNA in a particular stem cell been copied? Because it is in the process of copying that errors occur. Approximately 7 x 10^15 mature cells are produced in a human lifetime and these could be produced in 53 cell generations (2^53 = 9 x 10^15 ). In 60 cell generations a total of 10^18 cells would be produced, enough for over 1000 years of human life. Thus it is possible that, even in extreme old age, the mature cells of the body are fewer than 60 generations from the zygote.

      https://pubmed.ncbi.nlm.nih.gov/25459141/

    • It’s a really interesting model, pretty straightforward. I’d probably modify it by having a time varying cell division rate and a few things like that but really it would be nice to just see it fit to all the different cancer types and cancer diagnosis rates by age.

      • I’d probably modify it by having a time varying cell division rate and a few things like that

        Sure, I did this a few years ago and basically could fit all the curves provided by SEER with 7 free parameters. My conclusion was we need more data to constrain the parameters.

        But it isn’t just more of the same type of data. There are a few big issues with how cancer incidence is reported:

        1) The most commonly detected cancers by far, skin cancers, are so common the main ones aren’t even reported to cancer registries like SEER. Melanoma is reported, but the much more frequent (by 10-100x) basal and (to lesser extent) squamous cell cancers are not. These are the cancers that are easiest to detect, so should suffer from the least bias due to diagnosis rates.

        2) Binning incidence/mortality by ten-year intervals and then using 85+ as a catchall even though that is when we often see the interesting drop in cancer rates.

        3) Lack of data on the number of undiagnosed cancers. Until the 1970s, autopsies were performed on ~25% of patients who died in the hospital. Now it is much lower, closer to 5%. And that still doesn’t count all the people who did not die in a hospital or who had cancers that would also be missed upon autopsy.

        4) Constraints on the number of cells and turnover rate in various tissues as a function of age.

        5) Constraints on the rate of diagnosing various cancers as a function of age, eg doctors may not bother with a tumor in someone who is 90 because they assume heart disease or whatever is a more pressing issue.

        Then obviously not every cancer cell generated proliferates to the point it would be detected, and only does so with some lag. The immune system could clear it, no one is looking, and so on. So really the model predicts an upper bound on detected cancers, and ignores the proliferation phase.

      • Regarding the proliferative phase, you would do something like assume a tumor must grow to 1 cm^3 to be detected, which requires 10^7 – 10^9 cells:

        A tumor reaching the size of 1 cm 3 (approximately 1 g wet weight) is commonly assumed to contain 1 x 10 9 cells. This paper comments on the probable origin of this “magic” number and on some possible reasons why it has remained in use until now. However, mostly in epithelial tumors (85% of all human tumors) a cell number one order of magnitude smaller would be more realistic.

        […]

        Assuming that “all cells remain cycling and no cells are lost”, the increase of cell side from 10 μm to 30 μm theoretically requires 25 instead of 30 population doublings to form a tumor clinically detectable.

        https://pubmed.ncbi.nlm.nih.gov/19176997/

        That is about log2(10^8) ~25 divisions, ie doublings. Then doubling time for many cancers is something like 2-10 months (also see fig 5):

        The mean simulated doubling times for pancreatic cancer, melanoma, hepatocellular carcinoma (HCC), renal cell carcinoma, triple negative breast cancer, non-small cell lung cancer, hormone receptor positive (HR+) breast cancer, human epidermal growth factor receptor-2 positive (HER-2+) breast cancer, gastric cancer, glioblastoma multiforme, colorectal cancer, and prostate cancer were 5.06, 3.78, 3.06, 2.67, 2.38, 2.40, 4.31, 4.12, and 3.84 months, respectively. For all cancers, clinically reported doubling times were within the estimated ranges.

        https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8383152/

        So there should be a lag of ~5-20 years from carcinogenesis to detection. Eg, for doubling every six months it would be 25*6/12 = 12.5 years. NB: this is the 0.5 divisions/year rate estimated above.

        I don’t understand why people are coming up with these strange models like unanon cited instead of following something like the line of reasoning outlined here.

Leave a Reply

Your email address will not be published. Required fields are marked *