Shameless plug alert: Win prizes by forecasting real healthcare data to help UK’s health service save lives

This post is by Lizzie, on behalf of my colleague Will Pearse in the UK about a cool forecasting competition. It’s different in some ways than the cherry competition, but the same in other ways, predict blooms, predict beds… Only predict! (Only connect!)

Every four hours of delay in admitting patients from an emergency department adds up to roughly 25 potentially avoidable deaths per month (Howlett et al. 2026). We’re running a contest where you can forecast those risks 1–10 days ahead so UK hospitals can take action.

The data in the contest are all real healthcare data. The data come from the Bristol NHS system, with 220 variables ranging from daily counts to 15-minute feeds (e.g., bed occupancy, ambulance waiting times). You build a model in R or Python, submit by June 5, and forecasts are judged by mean squared error over short- and medium-term horizons.

The winning model(s) will be implemented in Bristol’s live system to flag emerging system pressure and so this is, literally, a chance for you to save lives with your stats know-how.

If you like this sort of thing (…and of course you do, you’re reading this blog…) then please sign up for the SPHERE-PPL mailing list. We (SPHERE-PPL) are organising a series of forecasting challenges just like this one, and we can tell you more about them.

Press release: https://www.imperial.ac.uk/news/articles/natural-sciences/life-sciences/2026/researchers-launch-forecasting-challenge-to-help-predict-severe-patient-harm-in-nhs-hospitals/

GitHub repo to enter: https://github.com/SPHERE-PPL/NHS-EAD-forecast

Black and white, gray and in between: What color is the media?

This post is by Lizzie

Recent reports of the toxic workplace culture of the Uber-fancy Copenhagen restaurant Noma confused me, initially. I thought we already knew this place and its lead chef, Redzepi, was toxic and it had closed. I eventually figured it out–these were new reports and involved more than reported before—and was surprised to find that Noma in Copenhagen closed but Redzepi came back and was ushering in folks in black SUVs past inflatable mushrooms to dine in LA (and other pop-ups).

This got me thinking of some of the parallels with academia. I think/hope we don’t have quite the level of bullying/abuse described in the recent NY Times investigation but we do seem to have:

  • A culture of entree via unpaid internships (volunteering) where you get the worst job.
  • A tension between what level of bullying is `part of the process’ versus too much.
  • Some media-savvy enough folks who may try to control the narrative.

I somehow have never done a post on how lame I find our ‘volunteer’ culture so I should do that. For now I will just say it’s a great way to reduce equity and has repeatedly been called out in conservation  (one of the worst offenders).

For the last one, I turn to the case of Tom Crowther (which also touches on the second point) that has had me wondering for some time: how much media outreach is too much for a lab? As you may or may not know, Tom Crowther was a assistant professor at ETH Zurich who published a series of high profile papers (including perhaps the most famous one suggesting we could nip anthropogenic climate change in the bud by just planting a lot more trees, including in lots of places that don’t have trees for good reason, like the savannas of Southern Africa where the carbon-rich grasslands are actually an excellent carbon sink, let alone a bit of a biodiversity hotspot) and was covered breathlessly in news and related bits of the journals of Science and Nature for his work (ETH went full Lance Armstrong effect on him).

What was less covered, was that recently he did not receive tenure from ETH after a series of bullying and harassment complaints were filed against him. I was fascinated by how little media there was on this and still wondering why. Given how much he was covered in the news, given the movies he was making, this seemed a story worth covering. I have a couple hypotheses for why.

Hypothesis one I am calling `isn’t it all always so confusing?’ It seemed messy (but isn’t that always the case? Hence my hypothesis name). The reports came some time after they reportedly happened, they appeared via a news story, not through formal ETH processes, and a number of lab members said they were very happy in the lab and urged anyone discussing this to stop.

Another (perhaps related) hypothesis, is that Swiss laws seemed to squash part of the story. The first I heard of the reports were from a story ran in Tages Anzeiger about a rising star professor at ETH who had bullied members of his lab. The newspaper had reports from multiple lab members they planned to publish, but were blocked because — apparently (as best I understand) — in Switzerland if you hear something will be published about you, you can try to block it (like a preemptive libel law, I gather). Why the accusers did not just go to a US, UK or some other nationality media outlet I never understood (that would get around the law, no?). Anyway, this all led to the story dripping out, to where everyone seemed to know one of the accusations included Crowther unzipping at a party and accosting one of his lab members with his privates, and it’s all filmed so we don’t have to wonder if this happened (later this was all detailed in the pro-Crowther version of the story in NZZ).

[Can I just take a short moment here to say that I am still confused by the number of colleagues who have said to me any of the following: Oh, but it was probably all in good fun! Can’t people take a prank these days? Maybe ETH should have told him not to do stuff like that; he needed more guidance. But should you really lose your job for this [at a government funded university where you have massive control, support and responsibility]?]

And then my last hypothesis (definitely perhaps related) is that the Crowther lab’s media branch semi-successfully controlled the story. The lab had an entire media branch and ETH found in its report that funds were used “not in accordance with ETH regulations. Specifically, funds were allegedly used for services such as crisis communication, legal advice and marketing.” It others words, he used ETH funds to combat the media reports surrounding the allegations against him. He admits this in the report and “explains that this was part of routine research communications before the recent media crisis caused by the media coverage at the end of August 2024. When the recent media crisis arose, he regarded the crisis communication services as a standard aspect of public relations.”

This last hypothesis has really made wonder how much media is too much media for an academic lab? As a climate change scientist, I have been raised on the dogma that we need to communicate our research. We need to help counter disinformation, get the word out. Fossil fuel companies have been pouring money into the disinformation campaign for as long as they have been accurately forecasting rising temperatures due to their fossil fuels (or pushing us to work on our own personal carbon footprints). Isn’t the massive branch of your lab doing media just taking this to the appropriate level? Or, is the outcome then that you control the media and you control the world? You become a mini-Rupert Murdoch and believe you deserve all that control?

And what happens then? Perhaps there is one more parallel between this story and Redzepi at Noma. Redzepi stepped down from Noma and many funders pulled out. But already wonderings abound about whether he is just lying low, waiting for the `squall’ to break. How long do you have to lay low? Crowther has resurfaced at KAUST University, which has hired him as a professor and has launched a new institute.

There’s an opportunity with post-Redzepi Noma to create something new. There’s also a chance he just comes back and nothing changes. The funders have a role in that outcome, but so do the diners. Do they come back? Or do they perhaps also want something new?

Missing forecasts

This post is by Lizzie.

Greetings from sunny Vancouver! I saw a magnolia in bloom today, which feels super early to me. But less early in that cherry and other plum trees have been blooming around the city. We have not had the coldest winter and had a recent warm-ish spell with lots of sun … so maybe that’s the trigger? The usual ‘two bucket’ model for spring flowers is that plants need to fill a bucket of winter cool (cool, but not cold, what this bucket is doing is covering some sort of latent variable related to ‘chilling,’ a possibly mystical biological concept*) and then a bucket of spring warmth (this makes more sense — warmth leads to cell growth and division). A little sun would likely help that second bucket.

Figuring this out of course is the goal of the International Cherry Blossom Prediction competition, which Jonathan Auerbach, David Kepplinger and I run.

Predictions just came in** and they tell me to hold my breath a while on Vancouver but that blossoms are coming this month probably. And I should wait a month for the east coast too — I hear it has been a cold winter there. We’re not far off the Washington Post’s predictions (which just came out) and of course we’re hoping to beat them.

We’re not sure how we compare to the National Park Service’s forecasts. They have had one each of the four years before when we ran the competition but seem quiet this year. I could imagine what may have happened but if anyone actually knows, let me know in the comments.

BTW, the bloom video is of Hibiscus, not cherries (if you want to visualize how far apart they are on the tree of life, go to https://www.onezoom.org/ and then type in ‘Hibiscus’ in the search, pick the first one, then type ‘Prunus’ and pick your favorite from the list, it will zoom you over). And big thanks to Tim Savas for the amazing video.

* I am hopeful we’re getting closer on ‘chilling’ — the best idea so far is that it’s capturing some enzyme or such that breaks down the callose (which I call a sugar, but someone told me is better defined as a polysaccharide) that blocks the plasmodesmata (the little wholes in cells that let them communicate with each other). A new study has new info and suggests this whole thing could connect to bet-hedging in plants … I am dubious about that last part.
**If you have comments on the figures, see also here.

The fifth annual Cherry Blossom Prediction Competition is open!

We are once again running the annual Cherry Blossom Prediction Competition, open to all! As part of the competition this year, the Washington Statistical Society is hosting the hackathon “Petals and Probabilities,” where students will work together to predict when the cherry trees will bloom this spring, followed by networking. It is not necessary to participate in the hackathon to enter the competition, but if you’re interested in coming to the hackathon, read on!

The hackathon will be an in-person event hosted at Georgetown University (111 Massachusetts Avenue) on February 21st, 2026 from 9 AM to 4 PM. The featured lunch speaker is Scott Olesen, Lead Data Scientist at the Center for Forecasting and Outbreak Analytics in the Centers for Disease Control and Prevention. Free food will be provided, and monetary prizes will be awarded.

Interested students from undergraduate through PhD should bring a laptop and register before the event (at this link). Reach out to hackathon lead organizers Ujjayini Das (ujstat at umd.edu) and Patrick Roney (pcr44 at georgetown.edu) with any questions about the hackathon. Sponsors include the Washington Statistical Society, the American Statistical Association, and the Massive Data Institute at Georgetown University. Competition organizers are Jonathan Auerbach, David Kepplinger and myself.

Protecting data from the public and ourselves

This post is by Lizzie.

Many years ago I received an email from some ecologists I had never met who wanted to use data from my PhD (remember my PhD field work that I told you about the other day)? Oh glorious day! Me and my shrubs were out in the sunlight of other researchers’ interests. And, better than that, all my data was already public so when they asked for the data I just pointed them to a link and wished them well. After some back and forth they asked me to read the resulting paper to make sure they had not mis-used or mis-understood the data from my PhD. I thought this unnecessary but read it anyway. I found nothing to complain about.

All I recall is that I didn’t love that they called the field site I worked at ‘a moonscape’ so I suggested they change it to ‘shrubscape.’ This was partly because I think Mojave desert when I think moonscape and I worked in a coastal habitat with more vegetation (see photo), but also because I was on a total kick of trying to get more folks using shrub as an adjective (I was thrilled when I published about ‘shruboreal arthropods’). This brings me to point 1 I wanted to make about data sharing today.

The idea that we should all contact the folks who produced data to make sure we use it correctly is misguided.

It’s a nice idea and I don’t want anyone to think I don’t enjoy meeting the people who produced data that I use (or really anyone uses, I watch movies to see the CERN people in their hard hats producing data and I would enjoy talking with those folks too) or that I don’t realize a lot of things can make data less straightforward (when I was a student we called these ‘demonic intrusions’). It’s that the idea leads to small (and more often less amazing) science. I think this is especially true in ecology where we spent decades sharing stories with each other of what we each found in our particular system (side shout out to Ryan Thum who will forever live on in my memory of him talking to me during the first year of my PhD and saying: if you truly want to advance ecology, then you should get two wheels, one labeled with different systems and one labeled with different questions. When each PhD student shows up they spin each wheel and that’s what they do their PhD on).

Think about how big the science you can do is if you require everyone who wants to use data produced by someone else to contact them. I definitely am not analyzing tree rings as a topic! (But shout out also to all the researchers who have all replied to our queries about their public data, including showing us photos of their western redcedar tree cores when we were incredulous that they could have 2 cm of growth in one year). No more global analyses of advancing plant leafout for the IPCC! There’s the obvious logistical problems of the effort to do this (including all the dead people who produced data that is now public), but also the issue of what we would ask them exactly. As a researcher who does large-scale analyses using other people’s data it’s on me to make sure I validate the data, understand the data and read all the metadata. Is this a way to skip those steps and just get an emailed approval check-mark? Or is there so often missing information—that I would somehow not find in the process of doing science (visualizations, analyses, etc.) and the researcher who shared the data would have forgotten to mention—that would derail the whole finding? I doubt it and that’s because of:

1) The person most likely to benefit from you posting your data with metadata is you. Future you. Because you will not remember all the details of how it was collected and which day someone turned off the automatic watering on your experiment and dried out your plants in a year, or two or ten.

So what are the odds you will be so good at helping others use your data in the future?

2) I don’t know so many cases where people screw it up. And by that I mean: I don’t know any cases. If you know of a good case tell me. And if you’re someone who is being told to worry about this, ask for the good examples of this problem. There must be a name—beyond the boogeyman—for this tendency by those opposed to data sharing to toss out the risks of data sharing that they themselves have never observed. It’s just out there in the ether, being so scary, like a camp story someone told them, that someone else told them about two campers … who were waiting for a friend … at night … in a dark parking lot…. Boo!

My other point (point 2) sort of relates to this boogeyman problem, which is the belief that data should be protected from ‘the public.’ This was a new one I did not know so many people were into until recently and so I am not (yet?) sure how to combat it (I am open to ideas). It seems pretty obvious to me that data generally thrives in the sunlight.

Open data is a major tenet of the open government movement that many democracies have signed onto (and, must I say this?, open data is not a major policy of autocracies) based on the reality that open data makes discriminatory policies obvious (e.g., redlining). There are realities where public government data has been misused to harm minority groups, but most often the reverse is true. Closed data makes it easier for governments to unequally distribute resources and hide other nefarious actions. So I am nervous about how many people seem quick to invoke minority rights and concerns to back up government policies to keep data closed—with zero evidence to support their supposition (is there a term for this? There should be). Shouldn’t our first reaction be that data should be open without good evidence to the contrary?

The Desperation of Causal Inference in Ecology

This post is by Lizzie. The Figure is take from Frank 2024.

I was in a meeting a little over a year ago in which I asked a student to define causal inference. The definition he gave me focused on complex approaches often used to try drag causality out of observational data. So I asked if causal inference involved experiments at all? “No,” came the reply. I double-checked. “No.” The student was certain. Someone else following up later did not change their mind.

Experiments cannot help with causal inference.

I knew we had a problem then, but how did it happen? I’ll tell you my version of what happened and some of what I can put together for how this happened, but I am open to other theories and ideas. And if perhaps the new ‘causal inference’ movement in ecology really has — finally — struck on a way for us to figure out ecology, then time will obviously prove me wrong, and you’re welcome to beat time to it in the comments section.

I could start back with Sewall Wright and the demes of cows I was once told he used to visit and Fisher and his fields of corn (or some agreeable consistent crop just waiting for its split plot design), but I will just start in the 1990s with path analysis in ecology. Path analysis (what I would call structural equation modeling with standardized coefficients) was hot in the 1990s in ecology. There was a chapter on it by Mitchell in the book ‘Design and Analysis of Ecological Experiments’ in 2001 (perhaps around its peak). It had this on the first page:

plant traits → visitation → pollination → reproduction

Isn’t that great? I could link plant traits to plant reproduction via those traits’ effects on (insect) visitation (to flowers) and how all that racey visiting led to pollination and then — reproduction (and then I might even make a run at … plant fitness!). I mean it is great. I like the idea. I liked the chapter. I did a path analysis. But I didn’t call it path analysis, I called it structural equation modeling because, by the time I was publishing, path analysis had hit some bumps.

Namely, everyone had done path analysis and many of those people had done it poorly in one way or another and suddenly all those little paths looked like a lot of made up stories with lots of little asterisks representing lots of significant p-values that didn’t really hold up to scrutiny. Shocker! (No, not shocker.) So, we all stopped doing path analysis and (within a few years it seems to me) we started doing structural equation modeling, sometimes with standardized coefficients. But we never called it path analysis again.

We couldn’t let go of path analysis because the dream was still alive. We wanted causality. We wanted to link things to explain how the world works. And manipulating plant traits is hard (have you ever tried to paint flowers different colors in a field? Or paste on tiny hairs (which we call trichomes)?), but measuring them is comparatively less hard. We wanted causality from observational data. That was the dream.

And, the dream is still alive. After all this time.

And the dreamers seem to have just discovered some of the basics of causal inference for observational data from the social sciences and econometrics literatures. With this, they have discovered that diversity (more species) in grasslands leads to lower productivity, not higher (Dee et al. 2023) and linked white nose syndrome in bats to increased infant mortality across the eastern US (Frank 2024). This latter paper is the one that rattled me because I attended a discussion group with colleagues and found out how many of my colleagues are excited by these ‘new techniques’ and how they have learned from them the amazing power of fixed effects for finding causality and the dangers of random effects to lead us astray.

Huh? Fixed effects to save the day and random effects of doom?

I tracked some of this down to me thinking of the common ecology definition of fixed versus random (I think some closer to definition #2, page 245 of Gelman and Hill: “2. Effects are fixed if they are interesting in themselves or random if there is interest in the underlying population. Searle, Casella, and McCulloch (1992, section 1.4) explore this distinction in depth.”) whereas the ‘new’ methods in ecology are using (I believe) definition # 5 (“5. Fixed effects are estimated using least squares (or, more generally, maximum likelihood) and random effects are estimated with shrinkage (“linear unbiased prediction” in the terminology of Robinson, 1991). This definition is standard in the multilevel modeling literature (see, for example, Snijders and Bosker, 1999, section 4.2) and in econometrics….”).

This explains some of the interesting lines I found in these papers, including:

Random effects account for clustering in data via the error structure of the model (Bolker et al. 2009; Gelman and Hill 2006), rather than estimating cluster means as part of the data generating process of a model (i.e., via fixed effect for each cluster’s mean, using the terminology of the mixed models literature). (Byrnes & Dee 2025)

The time-varying site attributes (μ_{st}) are also modeled in a fully flexible way that allows a year- specific effect for each site (in the estimation, an indicator for each year is interacted with an indicator for each site). (Dee et al. 2023)

I think the authors of this new Tower of Babel for ecology have also defined random and mixed effects to mean only ever linear models with lmer-style partial pooling on intercepts (never slopes I presume?) with fixed effects on slopes (back to definition #2). They even go so far as to refer to this as the “Common Design in Ecology” (they also capitalize Ecology and Ecologists in Byrnes & Dee 2025, which I find odd — is this high German? Personally, as an ecologist, I don’t think I need an capital letter) and explain:

Without more variable transformations, the multi-level modeling approach does not easily lend itself to controlling for as many unobservable sources of confounding as can be done in our linear, additive, fixed-effects panel data estimator. (Dee et al. 2023, in supp)

I thought about calling this post ‘The Tower of Babeling Causality’ or ‘The Problem with Statistical Terminology,’ but the real problem is not how lost in the weeds of words we get in with terminology. It’s partly how easily ecologists do want ‘new’ terms and approaches that will solve everything. The authors who have come armed with econometrics panel data approaches and instrument variable analysis (when most ecologists don’t know what is an instrument in their experiments) had the ground laid for them by all the ecologists who are enthralled by ‘random effects.’ I agree we have too many people trained to believe that chucking enough categorical covariates (site, plot, year …) on the intercept of a simple linear model will save the day. It’s a problem how much we sway from this being correct statistics to that being correct statistics. And how quickly think a new approach will change everything in ecology. We seem to quickly learn — and re-learn — that bad stats can easily lead you astray, but never take on that good statistics alone will not save you.

To be clear, I don’t have a giant problem with these methods. I have a problem with how they are presented as saviors (and somewhat how they are presented as new, but perhaps we need the ‘new’ and ‘savior’ angle to follow Grace) but I have a bigger problem in how rapidly they are being taken up. I fear the next 10 years I will live in a sea of piranhas where lots of ecological problems explain 5-10% of infant mortality and plant productivity.

And that’s the other problem — the bigger one: how much people want this causality. They want to believe that we have the data and methods to show that a disease that wipes out bats leads to an 8% increase in infant mortality. Of course we should want causality, we’re scientists, but the drive for causality seems to jettison a lot of the stuff we also need as scientists, especially estimates of uncertainty and the ability to leave room for uncertainty so that we search out better methods and better answers. I don’t know if bat decline has increased infant mortality 8% (though I highly doubt that number given the language of the author and how ‘outrageous’ he thinks it is that he is expected to share all his data for people to believe his claims). I just know we have managed to do science before and make progress and it wasn’t because we got better statistical methods or memorized a glossary of one particular set of people’s definitions of DAGs and fixed effects.

I am cited in one of these papers for old work I did where I compared shifts in the timing of flowering and leafout with warming over time (due to anthropogenic climate change and natural variation) and experimental warming (due to infrared heaters or teeny tiny plastic greenhouses — also, hello instruments in ecological experiments!). Estimates from experimental and observational data were different — the effect of warming in observational data was bigger. I did lots of different statistical analyses to figure this out, I even did something probably close to the ‘Common Design in Ecology’ (although the authors don’t seem to ding me for this) and the effect never went away. With Ailene Ettinger and other colleagues, I eventually got all new data and found the same thing using slightly different statistics. But that wasn’t why we got all the new data, we got it to test hypotheses about what drove the difference. And we found out that it appeared to be two things: warming experiments dry out soils which delays leafout and flowering and warming experiments over-report their warming (so their per degree estimates look smaller than they should).

I did all of this without ever invoking the term ‘causal inference.’ And that’s what really worries me for trainees today; that ‘causal inference’ will now mean a narrow branch of amazing ‘fully-flexible’ completely un-confounded statistics. We’re ecologists; we actually can manipulate some stuff. And somehow we’re going so gaga for econometrics statistics to give us causality through time-invariant fixed effects (or whatever) that we have students who don’t know how experiments could relate to causal inference.

What’s the solution? If you ask me, be less gaga over any statistical method (and I do love my own statistical methods so I could practice a little more of what I preach) and teach everyone basic mathematical notation and basic biological models. Teach them that generative modeling doesn’t belong to any one part of statistics or to only fixed or random effects. Teach them to be able to write out a simple biological model and simulate data from it and then fit their statistical model to it. ‘Only connect!’ Connect the models you learn for ecological theory with those you learn in stats. (And maybe teach them about the long debate in conservation biology about Cassandra’s curse, but that is a topic for another post.)

The mantra and mania of data sharing

This post is by Lizzie. The photo is from a photos folder I found from my PhD called ‘favorite stake photos.’

I wrote this post before seeing Andrew’s post for today.

When I was a grad student I spent a remarkable amount of time wandering around Sweetwater National Wildlife Refuge visiting 56 shrubs. Over and over again. I visited the shrubs almost every day. It wasn’t wandering, it was a structured, efficient route into the site, hitting each `replicate,’ then down the hill, up the next. Some days I was opening or collecting pitfall traps under the shrubs, some days I was vacuuming the shrubs (both of these tasks were to collect arthropods), other days I measured soil respiration, collected soil samples under the shrubs, I took clippings of the shrubs. There was also climate data to collect under the shrubs, little litter bags I was variously adding and removing under the shrubs. Vegetation sampling! I have forgotten much of it but I recorded a daily log so it can come flooding back. It felt like a lot of work.

In the end I published four papers about those 56 shrubs (which became 54 after a fire). Stuff about invasive grasses and carbon cycling, and bugs and stuff. Solid work, science maybe inched forward? Maybe it stepped right, but because of all the data I think it inched forward.

After grad school I joined a sort of think-tank that had been funded by NSF to promote data synthesis in ecology. Ecology needed it. It was (is?) a bit of a stuck field with a cacophony of individual studies in different places with different shrubs (or in lakes, forests etc.) — ‘boots and bucket’ ecology I heard it called.

You put on your boots, grabbed your bucket and — voila! Ecological science. The think tank was a renegade endeavor, trying to make sense of all the individual studies, by looking across them for patterns and maybe even testing some theory now and then. There was a tension between ‘boots and bucket’ ecologists at the time and ‘synthesis’ ecologists. According to the ‘boots and bucket’ tribe the `synthesis’ ecologists were stealing all their data for flashy papers. They were upsetting the order of things. Some said they were getting it all wrong because if you didn’t collect the data, you didn’t know the system enough, you could never figure anything out (yes, let’s all take a minute and think about where a field with this idea would be headed). Others said data would stop being collected and everyone would just do ‘synthetic ecology’ and never have much data.

This world was swirling far above and away from me and my 56 shrubs at Sweetwater, but at the think-tank they gathered all the new postdocs and told them about the power of data sharing. Science advances if we share data! Think of the questions we could answer if all the data were shared and organized! It just takes a hour or so to post your data. Go ahead, post your PhD data!

I was totally in. I posted all my PhD data and I dove in on the power of data sharing and wrote a paper about it for climate change biologists. I found papers showing that the massive improvements in pediatric oncology (for leukemia it went from 4 to 94% survival) could be attributed in part to data sharing. I read up on GenBank and drooled at a field so close to mine in topic but so far away in data sharing — and also trounces ecology in finding important interesting science IMHO. I felt like scientists should take an oath to advance science, and if they took that oath, then clearly they would see that they have to share “their” data. We’re trying to mitigate climate change people! Share your data!

Fast forward 20 years and all the scare tactics of the anti-data sharing folks have not come true. There’s no drop in data. (Though I got this argument recently from a marine biology postdoc, who then retreated a little from the premise when I asked for data on the declining data — given, and she did manage to agree, that this had been happening for 20 years at least so shouldn’t we see the pattern? — she then said data is only be propped up by PhD students who have to collect it for their PIs, so I guess suggesting a radical shift in how data are collected? And some verifiable decline in other data types? Through I didn’t try to steer the argument anymore.) Journals require data sharing. Granting agencies do. I think the synthetic ecologists won.

But a bunch of folks — beyond that one marine biologist postdoc — missed the message. I have been running into major governmental and non-profit data-collecting agencies that will not share data over the last two years.

For today, I will tell you about just one of them.

It’s the Canadian Forest Service (CFS). My lab recently contacted them for a big tree growth responses to climate across western North America analysis we’re doing. We have a lot of data, because these data in the US are generally public. They’re either on the ITRDB or they were uploaded with papers (there are certainly some that are not shared, but I like think they all will be, as the USFS and related US agencies do usually have a mandate to share data), but I happened to know that the CFS usually does not share data without co-authorship. They don’t share plot level data, they don’t share tree ring data, they don’t share data unless you sign an agreement with them and guarantee them co-authorship (and some other weird stuff that sort of sounds like they control whether you can publish what you find or not, but I think they have had to back off on that, so now there’s just related smushy language I suspect).

We asked anyway. I told my lab it’s important to ask and not just assume (even if everyone has told you that you will not get the data without co-authorship) and here’s the reply I got:

We are particularly interested in collaborating with researchers who bring expertise in Bayesian methods to help advance our analyses and explore future growth projections. With that in mind, we would like to explore the possibility of a scholarly collaboration with you.

Entering into collaboration would help streamline access to tree-ring data across Canada by removing certain barriers. Some of the data you requested are under restricted use, with licenses granted only to CFS researchers. Others require external requestors to obtain authorization from the original data owners—a process that can be time-consuming. Additionally, some datasets (highlighted in red below) have not yet been published, and we are actively encouraging collaborative projects that incorporate these data.

To provide further context, it is common practice for National Forest Inventory (NFI) data to have access restrictions, particularly for raw or highly detailed data. These restrictions are in place for several important reasons. NFI plots are located on both public and private lands. Disclosing precise locations could compromise data integrity or infringe on landowner privacy. Some datasets also contain sensitive ecological or proprietary information. Also, NFI data are designed to provide an unbiased representation of forest resources. Controlled access helps prevent misuse or misinterpretation, especially given the complexity of the data and the ongoing updates and revisions. NFI data support national and international reporting obligations, policy development, and collaborative efforts across jurisdictions. Ensuring consistent and validated use is essential. Finally, while publicly funded, the collection and processing of NFI data represent a significant investment (CFS & NSERC). Responsible dissemination protects this investment and ensures proper attribution and use.

I am working on a reply to this and open to all ideas/suggestions. I’ll give you what I have so far, vaguely in order of the arguments they have given (which is probably not the best order).

  • I appreciate their reply and understand their perspective, but requiring collaboration for access to data slows scientific progress, reduces equity and diversity in access to data, and has never been shown to be helpful or beneficial to science to publishing robust results, and thus is something my lab has a policy against (we do).
  • If you have data you have had for a while and not published (7 of the 30 datasets we asked about), but would like the data analyzed, then publish the data. This is the best way to get data analyzed and then it will likely be analyzed by different teams of researchers so CFS would get maximum insights from the data.
  • Sensitive data can be fuzzed, jittered or otherwise changed enough to meet privacy standards but still allow others to use the data. Certainly for our purposes, given the grid-size of the climate data we’re using it is hard to imagine this would not be possible.
  • The best way to get data cleaned, corrected and properly interpreted is to share it widely. The more eyes on the data, the quicker these issues can be spotted and fixed. Further, lack of access suggests there is something to hide, which is extremely concerning.
  • It is precisely because these data are used for national and international policy that making them public seems critical. (Can anyone help me here? This seems so obvious that I am not sure how else to say it.)
  • Data is far more widely used (and cited) when publicly shared. More papers and research seems like a better return on the Canadian taxpayers’ investment, no?
  • If you really want to charge for the data, then charge for it — but make access of those data available to all who can pay.
  • Fundamentally, there is a large number of researchers — Canadian and otherwise — who would use the CFS data and don’t because of this policy. These are excellent researchers who simply either do not have time for the efforts of collaboration with one team that requires collaboration in exchange for data when all other teams make the data public or do not want to support this process because it slows progress in science and the more researchers who sign onto it, the more it is tacitly condoned. Lots of good scientists I know — myself perhaps soon to be one — will not use CFS data because of the current access policies. Or, if they use it once, they won’t use it again.
  • Taxpayers paid for these data to be collected, they should get to see it and use it how they please. And with the way things are going, I would add that — if the commitment is to data quality — the more people who can access and download it now, the better. Political regimes of the worst kind often remove and restrict data.

I don’t think CFS researchers have anything to hide. They run a really nice database of their data (I know, because you get to search around it to request the data) and they are helpful and sharp when I meet with them or we correspond over email. I also don’t think widely incorrect papers or policies have been prevented by this restricted access. But I think that some researchers have been fed a steady diet to make them fear these possibilities and I am not sure how to disabuse them of this version of the world.

I hang out with a lot of people who share their data — climatologists often share it, and folks related to them (e.g., dendrochronologists in the US), phenology people usually are better (though I have had recent issues) — and I hang out with folks who don’t.

The people who share data are happier. They don’t spend time telling me all the horrible things that will happen if they share data. They don’t spend their time worrying about it. They just share their data and move on.

It’s like people who spend all their time talking about work-life balance. I find them much less happy than the people working until 11pm some days — those folks are often also the ones tango dancing until 1am the next night, or leaving on multi-day kayak trips or getting in a Truck Surf hotel to tool around Morocco surfing. The ones talking about work life balance seem to set on searching for something I think they would find if they stopped searching for it so desperately.

 

I am no longer chairing defenses or joining committees where students use generative AI for their writing

This post is by Lizzie. The photo is from Mount Rainier. 

I decided this week on some new rules for myself relating to graduate student training.

  • I will only chair defenses where the the student states they did not use generative AI at all in the writing of their thesis.
  • I will only join committees for graduate students when the students agree not to use generative AI at all or in limited (pre-defined) situations for their writing.

Why am I doing this? Because I have limited hours in my days, weeks and life and thus limited hours to dedicate to graduate student training and I want that time to be used most effectively in training folks. Time I spend reading AI-generated text — and possibly editing it for students — is not currently a good use of this time.

Why am I doing this now? Because it’s been bubbling up for a while but very recently it was like someone threw a small grenade in the pot and that got my attention. Meetings with an old friend and colleague pushed me and, as though the universe wanted to make sure that I got the message, I walked out of his house and arrived at the airport where I checked my email to prepare for a PhD thesis defense that I was chairing the next day and found the thesis was written in part with generative AI.

I personally never would have noticed this as this reality was tucked into one sentence in the preface, but the outside examiner’s report flagged that entire paragraphs appeared written by generative AI (chatGPT or similar). As chair a good bit of my job is overseeing how the report by the outside examiner is treated and considered (for those of you not familiar with the role of chair, you’re in the same boat as me when I started a faculty job here in Canada, you can see what the chair is supposed to do here) and so I scrambled to figure out the official UBC rules. They’re here.

I had already agreed to chair this defense, which was starting in about 16 hours so I felt that I could not back out, but I did want to get a sense of whether these rules were followed. I wanted to know how the student used generative AI and whether they knew when to use quotation marks around text from it versus just take it and run (this is what I understand of UBC guidelines, and it honestly makes sense to me, but is a debate), and when and how they discussed this with their supervisor and supervisory committee. Had they asked their supervisor or supervisory committee to read and edit AI-generated and AI-edited text without telling them it was AI-generated and/or AI-edited text? I suddenly realized that didn’t really seem okay to me.

My take-away from the whole thing that followed? We are in a deep mess that an alarming number of my colleagues would be happy to pretend is not happening and we have left students adrift at sea (which, I realized cycling home from the defense, rhymes with chatGPT). I feel fairly sure many of my colleagues would have been happier if I had not asked these questions and I was pretty horrified by some of the discussions there and in days since with colleagues. My small survey through conversations suggests to me that most people are fine to have their students use generative AI to edit their text, maybe help them write it etc.. One said, ‘I am so relieved I don’t have to edit my students’ writing so much anymore.’ I guess this means most people think it’s fine to ask their colleagues to review AI-generated text produced by their students? There was also a lot of throwing their hands up and saying, ‘oh, but there are no good guidelines, so what can we do?’

But reading the UBC policies and thinking about this a little made me easily disagree — I can come up with some guidelines, just like I can come up for guidelines for old-school plagiarism. And just like old-school plagiarism (before the era of Turnitin and such) I can never know if students follow those guidelines. But I can make it clear that not following them is academic misconduct to me and let students decide if they want to do that.

So here’s how it works if you’re a student and you want me to join your supervisory committee:

  • I’d prefer you not use generative AI to edit your writing at all (and I would rather read text with some grammatical errors).
  • I understand that you may want to use it to edit your English grammar. If you want to do this using generative AI you must promise me you will not use generative AI beyond this and show me that you know the difference. Thus, you need to:
    1) make up a short (say, 1-3 pages) document showing examples of both how you use AI to edit your writing *and* examples where using it without quotation marks would be considered plagiarism. Maybe also include examples where it is editing your language and not your grammar, and confirm you will not do this.
    2) You also need to find a way to share or record all of the changes made so you can document them.
  • These guidelines apply to your writing of text, not your writing of code. (Update from 19 July 2025: I see coding with generative AI as a different ball of wax. I am not asked to review your code line by line generally and it can be tested in ways your writing cannot, but this does not mean I am totally fine with generative AI for coding and not for writing).
  • Update (19 July 2025): These guidelines are not necessarily set in stone for the rest of my career. They are the ones I see as best for now.

What does this mean for my own students and trainees? I don’t want them to use generative AI for text I will help them with. This is their chance to have me edit their grammar and flow and everything and I think they should use it fully to learn it as well as possible. If they want to write someone else an email using Grammerly or leave the lab and do whatever, fine — but while I am here to help, I want the help I give not to be editing generative AI text.

I would hope that other supervisors can offer this also and thus, when I do read text that is not edited by generative AI, it won’t have so many grammatical errors. I don’t think that, as supervisors, we should all feel so relieved to stop editing our students papers and other writing. We do realize that all that time we’re spending will now be reading AI-generated and/or edited text, right? (Not to mention, all that text we’re working on is thus being rapidly fed into the models owned by private companies.) That said, I’ll be chatting with my lab about this in the coming weeks and am open to being swayed but the argument has got to be good.

On that note, I’ll address a few points I have heard more than once.

  • This is unfair to non-native speakers of English. The scientific publishing world (and conferences etc.) have elevated English and that is unfair in many ways. But I don’t see this as good fix to it at all. Depriving non-native speakers of the opportunity to have me help them with their writing in English does not seem a great outcome. No one in my lab currently is a native English speaker and I am fine editing their text — even if it takes slightly longer (though honestly I don’t think it does, because editing grammar is so quick compared to really teaching someone to write well), but I will be looking to hear what they think.
  • Asking people to quote from chatGPT or similar is insane, no one will do that. Yes, got it. Now ask yourself why you’re okay hiding that you’re doing it by not quoting it. If it’s not your writing, why did you give it to me in such a way that I cannot tell that it’s not your writing?
  • The world is changing, get with it. Sure, but this is about me spending the time I have to train and help others by reading/editing AI-text and that’s not a good change. If we want AI-generated text to be acceptable then I think we need a lot of other changes too, such as that we write A LOT LESS. I know this has been discussed before and if we come to where papers are just methods/results and published (totally open) data with some AI-generated/edited text then maybe I will start reading the bits of text again, but asking me to do it now across a 100-200 pages thesis is not a good use of my time and that means poorer training for students.

These rules flow to co-authorship and I expect I have already agreed to co-author work with other people’s students where they are writing with AI assistance, so I will have to figure out how to deal with that. I feel especially bad it has taken me over two years to come up with these rules and guidelines so students know what I want and know why I am asking for it.

I think we have really failed students in not sorting this out for them and I would be very interested to hear if you and/or your school have guidelines for thesis students (MSc and PhD). I have found pretty scant info or real useful rules.

It seems we’re all just going along and I am pretty worried about what I have heard from colleagues while discussing this. One of them, seeming exasperated at me when I said that lifting a sentence or most of a sentence from generative AI needed to be in quotation marks, ‘This is silly. How is this different than an advisor just editing text for a student?!’ To which I said, ‘because a supervisor can be a coauthor on a paper and, according to UBC guidelines, chatGPT cannot be.’ This left a long silence from which we never really recovered.

PhD position at UBC in Temporal Ecology Lab

This post is by Lizzie

My lab has an open position for a PhD student to join the lab. We’re looking for someone bright, motivated and collaborative to study how seed and seedling pathogens influence forest regeneration and diversity. This project would be part of a broader PhD with room to develop your own projects. This project is in close collaboration with the Plant Ecology Group at ETH Zürich, which is led by Professor Janneke Hille Ris Lambers.

If you’re interested, please find more information (including how to apply) on this page.

We’re open to folks from diverse training backgrounds. If your background is more computational and/or mathematical but you’d be excited to spend a few weeks outside and do a little lab work then you could be an ideal candidate; if you’re really interested, please apply no matter how well you think you line up with the ad.

Application review begins 1 July 2025 so apply soon for full consideration.

Ecologists’ endless quest for automatic inference

This post is by Lizzie.

At the end of a recent course I taught on Bayesian approaches (which reminds me I should blog an update on that) a student asked ‘so when do we divide up our data into test and training?’ This stopped me a little as the whole course was on a workflow approach to science and stats that I hoped hammered home how to gain mechanistic insights from simulated data, preparing you for more insights using retrodictive checks on a model fit to your empirical data, etc.. I was on the spot suddenly realizing some gaps and failures in my course content. I also should not have surprised, as ecologists are going in big time on machine learning (are there other uses for test/training data? Yes, but that’s the dominant place this language is in use in my field now, IMHO), and we (I) don’t step back and teach the different approaches.

In discussing this with a stats colleague recently he mentioned the endless search for automatic inference. `Feed in data, pull crank, get scientific inference.’ It’s the opposite of the workflow to me. I also think it’s not going to work well, but it’s clearly the dream, and an alarming percent of ecology is devoted to it, without even knowing it.

Machine learning is the new best hope of automatic inference for ecology (and a lot of other fields) without anyone seeming to notice what they’re not getting. It’s amazing to me how many students seem blithely unaware of what machine learning is going to give you — (good) predictions for out-of-sample data, but a difficult time finding interpretable parameters and all the science that can go with them. (And, yes, I know some of the machine learning approaches are working on changing this.) So they see it as the inference approach.

The previous best hope of automatic inference was model comparison (LOO is the new magic, AIC was a big — BIG — hit, before that was stepwise regression with an alarming number of ecologists never learning any potential for problems with stepwise regression, but I digress) and it’s still running strong in some circles. Fit 6 or 600 or so models and compare them to see which is best. In my area, the models balloon since we have no idea what climatic driver to include. For example, I think water matters to trees growing outside, so for a precipitation variable, should I use total precipitation? Our maybe just during the growing season? Or, wait, maybe divide up growing and non-growing season. But then for the non-growing season, should I use snow depth? Snow water equivalent (SWE)? This is so hard, and there’s no clear answer.

Automatic inference to the rescue! You can put them all in with model comparison, including a suite of possible interactions, and see which ones really matter. Yay!

Did this work? Not at all if you ask me. I recently saw a tree ring talk that did this but you can tell the best fitting model actually made no biological sense after they thought about it more, so they presented the ‘second best model.’ And I am quite sure the second and third best model were pretty similar in any comparison metric you wanted to throw at them and they might have had really different answers to how the world works. (Ecologists have tried one way around this — model averaging, which I don’t think offers much either.) I am not sure why everyone is doing this other than that (1) we have all tacitly agreed it’s okay and (2) the other option seems harder, more uncertain and maybe we have not all tacitly agreed it’s okay.

What have we never gotten out of this as best I can tell:
(a) We start to see new patterns in what matters in these model comparisons and say, ‘hey — all this work together really shows we should focus on SWE in this context. Thank goodness we did model comparison as there is no other way we would have figured this out.’
(b) We use something we learned in model comparison to design an experiment that teaches us something new. Like, ‘wow, I never thought extreme heat in August would be so important, I will now set up an experiment to test the role of extreme heat in August. I am so glad I put that predictor — and extreme heat in every other month and in 3-month windows — in my model so I could find this out.’
(c) The feeling of joy at saying, ‘look at my minimum adequate model! This is great and so helpful.’
We never get these things because the results are almost always a mess. We all know this as best I can tell so we don’t even look closely at them as reviewers any more.

What’s the other option?

The other option to me is that you pick your few best-guess damn variables — the ones you can make predictions about and describe the functional relationship of them to your response variable(s) and you put those in your model. Maybe you fit a few models, but not endless models. In my experience, the first step in this process alone (picking those variables) gains me way more insights than any model comparison ever has. Why? Because it’s the opposite of automatic inference. It requires me to think.

What’s the downside of this other option? One would be that we pick the wrong predictors and never see that amazing predictor we would have just tossed in on model comparison. But given where 20+ years of model comparison has gotten us I am discounting this possibility. The other — and this is what students in my classes are really worried about — is that we don’t all tacitly agree this is okay. Many students I suggest this to don’t think it’s okay. They see how widespread model comparison and its ilk are and worry they cannot get published without it. They aren’t even trained in how to pick those variables.

We’re so over the top on automatic inference we don’t even train our students to be prepared for anything else. And worse yet, we tell them they’re doing (good) science.

With machine learning* we’re slipping even further away from science and our training is getting even worse as best I can tell. Students at UBC in data science learn to ‘tidy’ data as though there is no domain expertise in this process. ‘Tidy’ means removing outliers, gap filling and other things that horrify me to see students learn in their first term. How on earth do they know what an outlier is when they don’t even know what the data are? After this they learn random forests and some simple neural nets. Science done.

What’s the solution? I desperately hope people smarter than me are working on this question. One answer is obviously raising our standards and discounting work that doesn’t really give us much from whatever model comparison they used. Another is better training — I think we all need to admit that training has got to change with machine learning on the rise. A lot of students I work with now only take data science — they learn only machine learning and don’t know what a regression is or think it is anything they use. They need to see how interconnected all the inference methods are and what aims each one works well on for now (and not) and be prepared that that might change. This seems tractable. What seems less tractable is better training in science — training students to know there’s no automatic inference for science and getting useful insights is actually messier, harder, and involves more uncertainty than most people tell you (but, if you ask me, it’s also a lot more fun).

*We’re somehow also now calling most of machine learning ‘AI’ in ecology. Are other fields doing this? Why (I mean, other than wanting to sound like you’re doing the absolute coolest, most cutting edge thing)?

Who could have imagined it? The IPCC working group 1 folks.

See https://www.ipcc.ch/report/ar6/wg1/figures/summary-for-policymakers/figure-spm-4/

Figure from https://www.ipcc.ch/report/ar6/wg1/figures/summary-for-policymakers/figure-spm-4/

This post is by Lizzie.

The Intergovernmental Panel on Climate Change (IPCC) is a United Nations group that helps assemble the best possible information on climate change and its impacts, and then tries to communicate it to the world. In 2017 they updated their former models for emissions scenarios (representative concentration pathways, RCPs) to include socioeconomic narratives and called them Shared Socioeconomic Pathways (SSPs, oy! The acronyms).

One has been ringing in my head for a while now. And I have been meaning to track it down, especially since some ‘who could have imagined this?’ conversations that I had months ago. I am finally making the time now as I wrap up teaching climate change ecology.

Here’s the description for SSP3 — Regional Rivalry – A Rocky Road:

A resurgent nationalism, concerns about competitiveness and security, and regional conflicts push countries to increasingly focus on domestic or, at most, regional issues. Policies shift over time to become increasingly oriented toward national and regional security issues. Countries focus on achieving energy and food security goals within their own regions at the expense of broader-based development. Investments in education and technological development decline. Economic development is slow, consumption is material-intensive, and inequalities persist or worsen over time. Population growth is low in industrialized and high in developing countries. A low international priority for addressing environmental concerns leads to strong environmental degradation in some regions.

Reference here and see also this.

There’s only five narratives. Missing from them is the chance to change which one we’re on; see also the Climate Action Tracker.

Let’s analyze how we analyze!

This post is by Lizzie. I thought of calling this post also: Lynx, hares and the utility of many-analyst studies, but I thought that schmeared stuff more than I mean to. I also took the photo from Paris last fall. Why a photo of Paris? Why not.

At an evening discussion event recently I made a passing comment about how I wish ecology knew more assuredly what causes the lynx-hare population dynamics (including their magical Lotka-Volterra cycles). Almost immediately several folks swept in that how could I think this when we *do* know. And then they proceeded to each share a different mechanistic (if you will) model, including:

– Maternal effects such that scared bunnies (aka hares) keep hiding out from lynx well after lynx population numbers have plummeted.
– Something to do with willow (the trophic level below the hares).
– Disease.
– Someone mentioned the moon (but then someone emailed me later about sunspots, maybe the moon was the sunspots misremembered?).
– And (I think) more hypotheses that I cannot now recall.

I was not surprised by this. One reason for this is that I have tried it before and got a similar set of assured and wildly diverging answers. The other is that ecology seems to me an inchoate field where we’re still struggling for general theories and how to sort out what is going on. I like to think we’re making it progress, but I am rather sure that it is currently slow.

This all brings me vaguely around to a recent paper that a reader pointed Andrew to and he then pointed me to it, which is a new many-analyst study by Gould and colleagues (lots of colleagues!) ‘Same data, different analysts: variation in effect sizes due to analytical decisions in ecology and evolutionary biology.’

As the title suggests the paper gives the same data and research questions to a set of teams (who signed up to do this, and be co-authors for their work) and then sees how different the ‘answers’ are. I say ‘answers’ because obviously a pain-point in this sort of study is how to decide what to extract from any analysis and deem an answer. The authors tried to think about this in advance, as they pre-registered their study, but I was amazed at how painful I found both the presence of the pre-registration — or, to be more specific: the continual side bars on ‘deviations from pre-registration’ — and sifting out what the authors themselves were trying to tell me. Before I get to the latter though, here’s my favorite deviation:

Some analysts had difficulty implementing our instructions to derive the out-of-sample predictions, and in some cases (especially for the Eucalyptus data), they submitted predictions with implausibly extreme values. We believed these values were incorrect and thus made the conservative decision to exclude out-of-sample predictions where the estimates were > 3 standard deviations from the mean value from the full dataset provided to teams for analysis.

I only skimmed the paper, but I think the abstract captures some of what confused me:

For both datasets, we found substantial variation in the variable selection and random effects structures among analyses, as well as in the ratings of the analytical methods by peer reviewers, but we found no strong relationship between any of these and deviation from the meta-analytic mean. In other words, analyses with results that were far from the mean were no more or less likely to have dissimilar variable sets, use random effects in their models, or receive poor peer reviews than those analyses that found results that were close to the mean.

This led me down a brief path of skimming some other many-analyst studies or viewpoints (and the religion one cited within). At the end of the path I realized that these studies are not simply trying to point out a potentially concerning level of variation in the answers obtained from different teams using the same data for the same question, but something more.

Some are suggesting this should be a new way to do science — as if doing this will give us more confidence in the answers, while others (including Gould et al.) were doing something else: trying to find out which types of analyses give ‘better’ answers. The authors don’t clearly say they’re doing this (at least not a quick skim) but why else would so much of the paper include peer reviews of the analyses or dissections of whether the presence of ‘random effects’ tend to give more similar answers (I felt there was someone who either thinks hierarchical models are ‘better’ or has heard that and thinks otherwise behind this particular analysis).

Either one of these aims makes me want much more than any of these studies are currently offering. For the former (‘let’s add this to how we do science’) I wanted more information on how exactly the authors think science advances, effectively if you’re telling me this will improve science I think you first owe me a good model of how science works so I can better assess your claim. In the latter (‘which way is better?’) I obviously wanted simulated data, where we could find out which methods got closer to the truth, because we would actually know the truth.

This got me to musing about what my colleagues the other night would think of the challenge of simulating data for a many-analyst ecology study. I wonder if some would bristle that we don’t know enough to simulate such data, but if that’s true, I think we have a real problem. And the challenge to ask ecologists to simulate data to then give out for a many-analyst study seems to me perhaps a better place to focus our efforts if we want to improve ecology than to ask more people to do many-analyst studies on different datasets given different questions (further, I’d be more interested in a many-analyst study on simulated data).

I think the value of many-analyst studies lies elsewhere. First, to show the variation (in which case, we do not need so many of them) and then, perhaps to get better models for specific applications. This is where it occurred to me the cherry blossom competition I run with Jonathan Auerbach and David Kepplinger is a many-analyst study of sorts, but in a very different spirit. We’re looking for more predictive models! We’re not disturbed by the variation in the way that I think I was supposed to be by some of the many-analyst studies I read (skimmed).

Indeed the most interesting part to me of Gould et al. was the discussion where issues were raised about whether the research questions given to the teams were too vague or whether readers should even be surprised by this variation. They write:

We recognize that some researchers have long maintained a healthy level of skepticism of individual studies as part of sound and practical scientific practice, and it is possible that those researchers will be neither surprised nor concerned by our results. However, we doubt that many researchers are sufficiently aware of the potential problems of analytical flexibility to be appropriately skeptical. I hope that our work leads to conversations in ecology, evolutionary biology, and other disciplines about how best to contend with heterogeneity in results that is attributable to analytical decisions.

I see their point, but then I wonder about how well they added to the conversation when I have no idea why they did so many tests or what their question(s) exactly was. I also think any such conversation should be framed with a solid grounding in both how science works (who knows, but there are theories and ideas, none of which were really mentioned) and how statistics work. Both of these areas should give all scientists a good dose of skepticism, so do we really need many-analyst studies for that? I would hope they’re offering something more.

Three side notes:
1) This work reminds me of the debate over whether the bird Parus major was declining due to mistiming with its caterpillar food resource with anthropogenic warming (the peak of caterpillar abundance was advancing faster than bird reproductive timing, these birds use caterpillars to feed nestlings). The Dutch team said it was happening in their birds, the British team said it was not happening in their woods, but whenever I worked on them together in a hierarchical model they looked the same to me (see interactions 221 for the Dutch and 180 for the UK in Fig 1B here). And then this paper came out after the Dutch team had more data.
2) Another big take-home to me from the discussion of Gould et al. was a search for a simple answer of how to fix this mess, as opposed to acknowledging there’s no one or easy thing that would fix this, as often mentioned on this blog.
3) I was also sort of disturbed how these papers seemed to think of model averaging and model comparison, but I will save that for another post.

Contestants and AI competitor predict cherry blossoms this month!

This post is by Lizzie.

This year’s International Cherry Blossom Prediction Competition has closed and we can now wait to see how contestants’ and the bot’s predictions fare. Results are up on the website, with average predictions around late this month (DC, Vancouver, Liestel) to early April (Kyoto and New York City). The bot predicts DC and Vancouver to be earlier than the average of the contestants’ entries and later for New York. For Kyoto and Liestel (which, incidentally have the most historical data) the bot and average of the contestants’ entries are quite close.

Thanks to Yu-lin Hsu for being our “AI-handler” (using ChatGPT o3-mini-high), co-organizers Jonathan Auerbach and David Kepplinger — who do most of the work compared to me — and all our great sponsors, supporters and contestants.

We’ll find out the winners once the trees blossom!

(Incidentally, I saw the plum trees starting to bloom last week on the road I cycle on to get to work. It felt early! I have spent the last couple years thinking they would be early so perhaps the year I am distracted is the one year they might sneak up on me.)

Entries due this Friday! Predict cherry blossom dates in five cities

This post is by Lizzie

As February draws to an end, so does your chance to enter the Cherry Blossom Prediction Competition! We challenge you to predict the bloom date of cherry trees at five locations throughout the world for cash and swag prizes. Entries are due Friday. 

This year, contestants will not only compete against each other for the top prizes–but against artificial intelligence. We will include one or more submissions from the most popular large language models. Our AI handlers will prompt the AI with the contest rules and the entries from previous competitions. The handlers will then execute any code written by the AI as part of their entry. Judges will review all entries without knowing which were submitted by humans and which were written by AI.

Any human that beats the AI will receive commemorative memorabilia indicating they “beat the bot in the 2025 International Cherry Blossom Prediction Competition.”

Beat the bot in this year’s Cherry Blossom Prediction Competition

This post is by Lizzie.

It’s back with a twist!

Once again, the Cherry Blossom Prediction Competition will run throughout February 2025. We challenge you to predict the bloom date of cherry trees at five locations throughout the world and win prizes.

However, this year, contestants will not only compete against each other for the top prizes—but against artificial intelligence.

We will include one or more submissions from the most popular large language models. Our AI handlers will prompt the AI with the contest rules and the entries from previous competitions. The handlers will then execute any code written by the AI as part of their entry. Judges will review all entries without knowing which were submitted by humans and which were written by AI.

Any human that beats the AI will receive commemorative memorabilia indicating they “beat the bot in the 2025 International Cherry Blossom Prediction Competition.”

Help teaching short-course that has a healthy dose of data simulation

catsinNH

This post is by Lizzie. I hope you like the cats photo from this summer. I do.

I am looking for help. I decided to change my term course (12-14 weeks-long) on `introduction to Bayesian modeling with some hierarchical modeling’ (no, that’s the not the official title, but that is the gist) to a three-week intensive. I have been thinking about doing this for a couple years, but finally decided to do it now. My feelings in teaching the term-length versus short course is that: (1) a lot of students discover during term that Bayesian approaches are a lot more work than what they can get away with in frequentist methods and lose interest, but are still stuck in class for weeks — this way, when they can find this out, the course will soon be over! (2) As a corollary to (1), taking a Bayesian course (to me) mainly means finding out if you want to dive in and do it more (as you will never learn enough in a term course to be off and running fully), so with a short course, more students will take it and find out if they want to learn more. (3) More students will take a short-course and provide more support for those who want to continue on (building a useful community). (4) More students will take the course and learn how to simulate data, and I want more people to learn this. There’s lots of other reasons, but those are my big ones.

The downsides are that students will likely arrive unprepared and I won’t have time to go through the basics the way I do in 12+ weeks and that generally, there will be less content. Also, no more analyzing your own data as a term project that I support. People will miss that.

The students will be mainly MSc and PhD students in ecology or evolution (but not in molecular evolution, think more: people studying populations of birds and how their phenotypes might evolve) who will hopefully not be taking many other courses. Most will work in R and have some background in statistics, but not have been too challenged to fully understand the statistics they’re using.

I have two days of 3.5 hours each per week for three weeks (about 21 hours) and plan to cover:
Week 1: Data simulation for linear regression; what are priors and some ways to check them
Week 2: Fitting a model in rstanarm to simulated data; diagnostics
Week 3: Introduction to hierarchical models and posterior predictive checks

What I could use help on:
– Suggested classroom activities and problem sets
– Good example datasets or vignettes and generally any other good resource that could help me teach this material
– Advice on how to structure or approach a short course to make it work well
– Recommended background info to point students to

I have some ideas of examples I might use, like simulating data to show how much bigger a sample size you need to estimate an interaction in week 1 perhaps, and the Olympic figure skater example (judges, skaters from Regression and Other Stories) as homework for week 3, but I am really not sure so all advice and ideas would be most welcome.*. I am especially hoping to emphasize data simulation throughout, so ideas there are extra welcome. Please let me know in the comments any ideas/thoughts you have (or you can email me if you much prefer). Thanks in advance.

* One thing that I am NOT looking for is a big debate on the values of trying to teach all the way to hierarchical when people may not really have the basics.

Scientific publishers busily thwarting science (again)

This post is by Lizzie.

I am working with some colleagues on how statistical methods may affect citation counts. For this, we needed to find some published papers. So this colleague started downloading some. And their university quickly showed up with the following:

Yesterday we received three separate systematic downloading warnings from publishers Taylor & Francis, Wiley and UChicago associating the activity with [… your] office desktop computer’s address. As allowed in our licenses with those publishers, they have already blocked access from that IP address and have asked us to investigate.

Unfortunately, all of those publishers specifically prohibit systematic downloading for any purpose, including legitimate bibliometric or citation analysis.

Isn’t that great? I review for all of these companies for free, in rare cases I pay them to publish my papers, and then they use all that money to do this? Oh, and the university library signed a contract so now they pay someone to send these emails… that’s just great. I know we all know this is a depressing cabal, but this one surprised me.

In other news, this photo is from my (other) colleague’s office, where I am visiting for a couple days.

I’ve been mistaken for a chatbot

… Or not, according to what language is allowed.

At the start of the year I mentioned that I am on a bad roll with AI just now, and the start of that roll began in late November when I received reviews back on a paper. One reviewer sent in a 150 word review saying it was written by chatGPT. The editor echoed, “One reviewer asserted that the work was created with ChatGPT. I don’t know if this is the case, but I did find the writing style unusual ….” What exactly was unusual was not explained.

That was November 20th. By November 22nd my computer shows a file created named ‘tryingtoproveIamnotchatbot,’ which is just a txt where I pasted in the GitHub commits showing progress on the paper. I figured maybe this would prove to the editors that I did not submit any work by chatGPT.

I didn’t. There are many reasons for this. One is I don’t think that I should. Further, I suspect chatGPT is not so good at this (rather specific) subject and between me and my author team, I actually thought we were pretty good at this subject. And I had met with each of the authors to build the paper, its treatise, data and figures. We had a cool new meta-analysis of rootstock x scion experiments and a number of interesting points. Some of the points I might even call exciting, though I am biased. But, no matter, the paper was the product of lots of work and I was initially embarrassed, then gutted, about the reviews.

Once I was less embarrassed I started talking timidly about it. I called Andrew. I told folks in my lab. I got some fun replies. Undergrads in my lab (and others later) thought the review itself may have been written by chatGPT. Someone suggested I rewrite the paper with chatGPT and resubmit. Another that I just write back one line: I’m Bing.

What I took away from this was myriad, but I came up with a couple next steps. I decided this was not a great peer review process that I should reach out to the editor (and, as one co-author suggested, cc the editorial board). And another was to not be so mortified as to not talk about this.

What I took away from these steps were two things:

1) chatGPT could now control my language.

I connected with a senior editor on the journal. No one is a good position here, and the editor and reviewers are volunteering their time in a rapidly changing situation. I feel for them and for me and my co-authors. The editor and I tried to bridge our perspectives. It seems he could not have imagined that I or my co-authors would be so offended. And I could not have imagined that the journal already had a policy of allowing manuscripts to use chatGPT, as long as it was clearly stated.

I was also given some language changes to consider, so I might sound less like chatGPT to reviewers. These included some phrases I wrote in the manuscript (e.g. `the tyranny of terroir’). Huh. So where does that end? Say I start writing so I sound less to the editor and others ‘like chatGPT’ (and I never figured out what that means), then chatGPT digests that and then what? I adapt again? Do I eventually come back around to those phrases once they have rinsed out of the large language model?

2) Editors are shaping the language around chatGPT.

Motivated by a co-author’s suggestion, I wrote a short reflection which recently came out in a careers column. I much appreciate the journal recognizing this as an important topic and that they have editorial guidelines to follow for clear and consistent writing. But I was surprised by the concerns from the subeditors on my language. (I had no idea my language was such a problem!)

This problem was that I wrote: I’ve been mistaken for a chatbot (and similar language). The argument was that I had not been mistaken — my writing had been. The debate that ensued was fascinating. If I had been in a chatroom and this happened, then I could write `I’ve been mistaken for a chatbot’ but since my co-authors and I wrote this up and submitted it to a journal, it was not part of our identities. So I was over-reaching in my complaint. I started to wonder: if I could not say ‘I was mistaken for an AI bot’ — why does the chatbot get ‘to write’? I went down an existential hole, from which I have not fully recovered.

And since then I am still mostly existing there. On the upbeat side, writing the reflection was cathartic and the back and forth with the editors — who I know are just trying to their jobs too — gave me more perspectives and thoughts, however muddled. And my partner recently said to me, “perhaps one day it will be seen as a compliment to be mistaken for a chatbot, just not today!”

Also, since I don’t know an archive that takes such things so I will paste the original unedited version below.

I have just been accused of scientific fraud. It’s not data fraud (which, I guess, is a relief because my lab works hard at data transparency, data sharing and reproducibility). What I have just been accused of is writing fraud. This hurts, because—like many people—I find writing a paper a somewhat painful process.

Like some people, I comfort myself by reading books on how to write—both to be comforted by how much the authors of such books stress that writing is generally slow and difficult, and to find ways to improve my writing. My current writing strategy involves willing myself to write, multiple outlines, then a first draft, followed by much revising. I try to force this approach on my students, even though I know it is not easy, because I think it’s important we try to communicate well.

Imagine my surprise then when I received reviews back that declared a recently submitted paper of mine a chatGPT creation. One reviewer wrote that it was `obviously Chat GPT’ and the handling editor vaguely agreed, saying that they found `the writing style unusual.’ Surprise was just one emotion I had, so was shock, dismay and a flood of confusion and alarm. Given how much work goes into writing a paper, it was quite a hit to be accused of being a chatbot—especially in short order without any evidence, and given the efforts that accompany the writing of almost all my manuscripts.

I hadn’t written a word of the manuscript with chatGPT and I rapidly tried to think through how to prove my case. I could show my commits on GitHub (with commit messages including `finally writing!’ and `Another 25 mins of writing progress!’ that I never thought I would share), I could try to figure out how to compare the writing style of my pre-chatGPT papers on this topic to the current submission, maybe I could ask chatGPT if it thought I it wrote the paper…. But then I realized I would be spending my time trying to prove I am not a chatbot, which seemed a bad outcome to the whole situation. Eventually, like all mature adults, I decided what I most wanted to do was pick up my ball (manuscript) and march off the playground in a small fury. How dare they?

Before I did this, I decided to get some perspectives from others—researchers who work on data fraud, co-authors on the paper and colleagues, and I found most agreed with my alarm. One put it most succinctly to me: `All scientific criticism is admissible, but this is a different matter.’

I realized these reviews captured both something inherently broken about the peer review process and—more importantly to me—about how AI could corrupt science without even trying. We’re paranoid about AI taking over us weak humans and we’re trying to put in structures so it doesn’t. But we’re also trying to develop AI so it helps where it should, and maybe that will be writing parts of papers. Here, chatGPT was not part of my work and yet it had prejudiced the whole process simply by its existential presence in the world. I was at once annoyed at being mistaken for a chatbot and horrified that reviewers and editors were not more outraged at the idea that someone had submitted AI generated text.

So much of science is built on trust and faith in the scientific ethics and integrity of our colleagues. We mostly trust others did not fabricate their data, and I trust people do not (yet) write their papers or grants using large language models without telling me. I wouldn’t accuse someone of data fraud or p-hacking without some evidence, but a reviewer felt it was easy enough to accuse me of writing fraud. Indeed, the reviewer wrote, `It is obviously [a] Chat GPT creation, there is nothing wrong using help ….’ So it seems, perhaps, that they did not see this as a harsh accusation, and the editor thought nothing of passing it along and echoing it, but they had effectively accused me of lying and fraud in deliberately presenting AI generated text as my own. They also felt confident that they could discern my writing from AI—but they couldn’t.

We need to be able to call out fraud and misconduct in science. Currently, the costs to the people who call out data fraud seem too high to me, and the consequences for being caught too low (people should lose tenure for egregious data fraud in my book). But I am worried about a world in which a reviewer can casually declare my work AI-generated, and the editors and journal editor simply shuffle along the review and invite a resubmission if I so choose. It suggests not only a world in which the reviewers and editors have no faith in the scientific integrity of submitting authors—me—but also an acceptance of a world where ethics are negotiable. Such a world seems easy for chatGPT to corrupt without even trying—unless we raise our standards.

Side note: Don’t forget to submit your entry to the International Cherry Blossom Prediction Competition!

Cherry blossoms—not just another prediction competition

It’s back! As regular readers know, the Cherry Blossom Prediction Competition will run throughout February 2024. We challenge you to predict the bloom date of cherry trees at five locations throughout the world and win prizes.

We’ve been promoting the competition for three years now—but we haven’t really emphasized the history of the problem. You might be surprised to know that bloom date prediction interested several famous nineteenth century statisticians. Co-organizer Jonathan Auerbach explains:

The “law of the flowering plants” states that a plant blooms after being exposed to a predetermined quantity of heat. The law was discovered in the mid-eighteenth century by René Réaumur, an early adopter of the thermometer—but it was popularized by Adolphe Quetelet, who devoted a chapter to the law in his Letters addressed to HRH the Grand Duke of Saxe Coburg and Gotha (Letter Number 33). See this tutorial for details.

Kotz and Johnson list Letters as one of eleven major breakthroughs in statistics prior to 1900, and the law of the flowering plants appears to have been well known throughout the nineteenth century, influencing statisticians such as Francis Galton and Florence Nightingale. But its popularity waned among statisticians as statistics transitioned from a collection of “fundamental” constants to a collection of principles for quantifying uncertainty. In fact, Ian Hacking mocks the law as a particularly egregious example of stamp collecting statistics.

But the law is widely used today! Charles Morren, Quetelet’s collaborator, later coined the term phenology, the name of the field that currently studies life-cycle events, such as bloom dates. Phenologists keep track of accumulated heat or growing degree days to predict bloom dates, crop yields, and the emergence of insects. Predictions are made using a methodology that is largely unchanged since Quetelet’s time—despite the large amounts of data now available and amenable to machine learning.

What’s up with spring blooming?

 

This post is by Lizzie.

Here’s another media hit I missed; I was asked to discuss why daffodils are blooming now in January. If I could have replied I would have said something like:

(1) Vancouver is a weird mix of cool and mild for a temperate place — so we think plants accumulate their chilling (cool-ish winter temperatures needed before plants can respond to warm temperatures, but just cool — like 2-6 C is a supposed sweet spot) quickly and then a warm snap means they get that warmth they need and they start growing.

This is especially true for plants from other places that likely are not evolved for Vancouver’s climate, like daffodils.

(2) It’s been pretty warm! I bet they flowered because it has been so warm.

Deep insights, I know …. They missed me but luckily they got my colleague Doug Justice to speak and he hit my points. Doug knows plants more than I do. He also calls our cherry timing for our …

International Cherry Prediction Competition

Which is happening again this year!

You should compete! Why? You can win money, and you can help us build better models, because here’s what I would not say on TV:

We all talk about ‘chilling’ and ‘forcing’ in plants, and what we don’t tell you is that we never actually measure the physiological transition between chilling and forcing because… we aren’t sure what it is! Almost all chilling-forcing models are built on scant data where some peaches (mostly) did not bloom when they were planted in warm places 50+ years ago. We need your help!