Extinct Champagne grapes? I can be even more disappointed in the news media

Happy New Year. This post is by Lizzie.

Over the end-of-year holiday period, I always get the distinct impression that most journalists are on holiday too. I felt this more acutely when I found an “urgent” media request in my inbox when I returned to it after a few days away. Someone at a major reputable news outlet wrote:

We are doing a short story on how the climate crisis is causing certain grapes, used in almost all champagne, to be on the brink of extinction. We were hoping to do a quick interview with you on the topic….Our deadline is asap, as we plan to run this story on New Years.

It was late on 30 December so I had missed helping them but still had to reply that I hoped that found some better information because ‘the climate crisis is causing certain grapes, used in almost all champagne, to be on the brink of extinction’ was not good information in my not-so-entirely-humble opinion as I study this and can think of zero-zilch-nada evidence to support this.

This sounded like insane news I would expect from more insane media outlets. I tracked down what I assume was the lead they were following (see here), and found it seems to relate to some AI start-up I will not do the service of mentioning that is just looking for more press. They seem to put out splashy sounding agricultural press releases often — and so they must have put out one about Champagne grapes being on the brink of extinction to go with New Year’s.

I am on a bad roll with AI just now, or — more exactly — the intersection of human standards and AI. There’s no good science that “the climate crisis is causing certain grapes, used in almost all champagne, to be on the brink of extinction.” The whole idea of this is offensive to me when human actions are actually driving species extinct. And it ignores tons of science on winegrapes and the reality that they’re pretty easy to grow (growing excellent ones? Harder). So, poor form on the part of the zero-standards-for-our-science AI startup. But I am more horrified by the media outlets that cannot see through this. I am sure they’re inundated with lots of crazy bogus stories every day, but I thought that their job was to report on ones that matter and they hopefully have some evidence are true.

What did they do instead of that? They gave a platform to a “a highly adaptable marketing manager and content creator” to talk about some bogus “study” and a few soundbites to a colleague of mine who actually knew the science (Ben Cook from NASA).

The loneliest kind of lonely

This post is by Lizzie. I took the photo on Tuesday at a Sforza fortress in Bellinzona. The museum currently also has an installation by Zimoun, who also did `36 ventilators, 4.7m3 packing chips.’

I recently read Tree Story: The History of the World Written in Rings by Valerie Trouet and recommend it. It’s a quick and fun read about dendrochronology — the study of tree rings to reconstruct climate (good recap of dendro here from recent What’s Going On in This Graph?) — weaving in the author’s scientific life and lots of the best tree ring stories (I think my favorite story may still be related to the climate reconstruction of the period around Genghis Khan’s empire). It was recommended to me the Tree Spotters National Phenology Network group in Boston, who I was chatting with when back in Boston in May.

After years of fascination with tree rings (including the intrigue of their divergence problem) I was finally thinking of dipping my toe in the topic, starting somewhere I had a little bit of scientific overlap: shifting growing seasons (I work on that!) and tree growth measured via radial ring width (I don’t work on that).

We’ll see what happens with that, but for now I am still trying to understand this massive field, and I feel I am coming from another pitch. Dendrochronology is focused on extracting climate signals from tree rings; as an ecologist I work a lot more on how trees — from seeds to younger saplings (that we can manipulate in experiments) to older trees — grow and reproduce. The dendro folks definitely do not think much about reproduction and I can’t fully blame them as even I set that aside in some ways with trees (when I am really interested in growth-reproductive trade-offs I work on annual or biennial plants), but they also think about growth from a climate focus, whereas I think about growth somewhat differently.

Recently someone in my lab said ‘of course we know drought and temperature matter to tree growth, that is not interesting.’ I perhaps would not go that far, but it captured the contrast. As a community ecologist tree growth to me is limited by lots of things, but especially other trees. For species to coexist in a stable manner my field says within-species competition must limit growth more than between-species competition (with this balance occurring most in the most climatically favorable years, which should differ among species for coexistence) — this whole world of tree-to-tree competition is where the story is at. It’s how we get diverse forests, and certainly it relates a lot to how trees grow.

Competition is something dendrochronologists spend a good while trying to avoid (at least for the trees that they core). They want trees growing on the edge of a ledge of stone, alone in the world (and I guess we assume alone forever) and they ‘standardize’ out the mess of the early years when we think most of the important competition between trees might happen. They fit a growth curve and take the residuals. Then they take all the time-series they have for individual trees in a similar (but climatically caustic) area and put them together so they get a ‘chronology.’ I knew all this before, but it was better explained in Tree Story. And the process of figuring out which species to use to reconstruct which climate variables and where to find these trees battling the climatic elements also explained.

But what the book didn’t explain and I still don’t get is where the field goes from its current approaches. I think I understand why the standardizing approach would have been good when the field started a hundred years ago or so, or maybe even 20-30 years ago, but I am not sure I see how it can be the path forward (so much so that the database for tree ring data it is all available as already `standardized’). I’d be interested if someone could convince me this is the way to go statistically.

Andrew has a paper on this (Schofield et al. 2016), which I also really like, but it doesn’t seem to be taking dendrochronology world by storm, especially compared other papers using Bayesian methods and tree rings, which seem to often use pre-standardized time series (e.g., Tingley & Huybers 2010, but correct me if I am wrong). This is partly social science, it would be a lot of work to take away the standardization methods and I can see the cost:benefit analysis many might be doing. It probably feels worth more now to build up the spatio-temporal methods. But it also makes the field hard to approach for people like me who see their part of the science lost to all the detrending and averaging.

Bayesian Believers in a House of Pain?

This post is by Lizzie

A colleague from U-Mass Amherst sent me this image yesterday. He said he had found ‘multiple paper copies’ in an office and then ruminated on how they might have been used. I suggested they might have been for ‘a group of super excited folks at a conference jam session!’

This leads back to a rumination I have had for a long time: how come I cannot find a MCMC version of ‘Jump Around’? It seems many of the lyrics could be improved upon with an MCMC spin (though I would keep: I got more rhymes than there’s cops at a Dunkin’).

My colleague suggested that there is perhaps a need to host a Bayesian song adaptation contest ….

Messy data crashes into us

This short post is by Lizzie. 

Excellent job title alert! University of Nebraska-Lincoln has an open-rank position in messy data. Thanks to my former student, Dan, for sharing this.

In other news, I am in France (near the toy shop in this photo) where the radio also goes on strike. This means that at 8am last Thursday when I went to listen to the top-of-the-hour news I instead heard ‘There is a light and it never goes out, There is a light and it never goes out …’

As the last days of January dawn… the 2nd International Cherry Blossom Prediction Competition arrives!

Just in time for February, it’s the return of the great annual International Cherry Blossom Prediction Competition! (This post is by Lizzie.)

Help scientists like me better understand the impacts of climate change and win cash prizes by predicting when the cherry trees will bloom in four cities across the globe. The competition is open to all with prizes for closest prediction as well as other categories.

Interested? Check out the website with all the details, including data, rules and how to enter here.

Last year over 80 contestants from across four continents entered with a variety of prediction approaches. You can read more about last year’s competition here.

A big thanks to the American Statistical Association, Caucus for Women in Statistics, Columbia University’s Department of Statistics, George Mason University’s Department of Statistics and Posit (formerly RStudio) for their support, and partnerships with the International Society of Biometeorology, MeteoSwiss, USA National Phenology Network, and the Vancouver Cherry Blossom Festival—as well as Mason’s Institute for a Sustainable Earth, Institute for Digital InnovAtion, and the Department of Modern and Classical Languages. Sponsors and partners will be updated on the website.

Organizers: Jonathan Auerbach and David Kepplinger (George Mason University) and Elizabeth Wolkovich (University of British Columbia)

Important questions

This post is by Lizzie.

I have been robbed twice in my life: first as a grad student when I was traveling and most of my belongings were in a rental car, and second when some teenagers crow-barred the apartment I was staying in. The first time I was robbed I learned a some useful information that made getting robbed the second time easier: when you tell people you’ve been robbed they often say the most useless things. People who purported to be my friends would reply immediately with, ‘what type of car was it?’ or ‘where exactly did you park?’ While there are definitely predictors that increased the probability of me being robbed, I didn’t fully see the point of these questions.

Eventually my Mum or some other soul wiser than me explained this to me. They’re asking these questions to construct a model of the world in which they don’t get robbed. I sort of get this: it can be traumatic to have most of your worldly possessions stolen. I hope they didn’t also know that it can be painful to have your friends be so lame and self-centered.

I was trying to explain something similar the other night when a friend and colleague caught me after a group zoom call to ask me something. She had met someone who was a graduate student when I interviewed at their graduate institution for a faculty job. I always enjoyed meeting grad students on interviews; they were so much fun, asked interesting questions and made me hopeful about the world. I interviewed at many places and have run into some of these students since and always really enjoy it, so I thought this was where the conversation was going. No.

The student had recounted that at the end of my job talk a senior white faculty member got up and made an obnoxious and gendered comment that in no way related to my science.

I don’t remember this. I suspect I don’t because it was the least of my troubles back then. I remembered the guy though. In our one-on-one interview he asked me if I was planning to freeze my eggs. And in case you as a reader want to chock this up to a one-off, my notes from asking a female faculty member in the department are: “Everyone has had totally weird interactions with him, when she interviewed he told her all about his relationship with his wife. Someone else interviewing told him to **** off. And nothing slows him down, he’s gotten worse with age.”

What I told my colleague on the Zoom call, though, was what hurt was how people replied to it. How many people tried to brush it off in one way or another (even when I couched it in my reality — I liked this guy in many ways and people are multifaceted and complex, they are rarely all good or evil).

I was trying to recount this an hour later over beers but I didn’t get to finish the thought. One woman worked to come up with witty quips that I should have said back. And the one guy said, ‘oh, I know him. He likes to shock people. You know how you should handle this guy is ….’ which led onward to ‘I wanted to invite him a few years back and my colleague said ‘I dunno’ so then we invited him and someone known to be even crazier! And let me tell you about him. He got into a physical altercation with someone ….’

So, I just wanted to make a public service announcement to folks who want to reply in similar ways to people telling you about harassment they received — consider keeping these thoughts to yourself. And then maybe take some time to figure out why you so desperately jump at saying them.

In other news, look what I saw on my hike on Wednesday!

Wandering through Sforza castle

Weekend before last I spent a day in Milan to see an old colleague (I am on leave in Zurich just now so miraculous things like a dayhop to Milan are feasible). He used to be a research technician in my lab where he did the classical ecology lab tasks of things like identifying and tagging trees in forests, making thousands of observations of leaf unfolding in a growth chamber experiment (on paper! Though that part was neither my idea nor his) and nudged the lab forward through working on automating photo capture of leaf unfolding. The resulting images and time lapse videos were gorgeous, but they didn’t change how I did science.

But since then my Bayesian models have become far more generative—many thanks to my Stan collaborators–and I have started to realize some sad hard truths about my data and my science. The first is that you need a lot more data than I often have to fit some of the models I think underlie the processes I am studying. I work on winegrapes because they have way more data than other systems I am interested in (such as forests, which critically store carbon) and when I don’t have enough data to fit a temperature response curve for winegrapes I can skip trying it on more `natural systems.’ The second is that we also need way better data. A stats-department colleague said to me this year, ‘it’s not like you don’t have a lot of data, it’s like the data you have are quarters [and you need higher denominations].’ By later in the day when he next picked up the metaphor, but was now calling my data pennies. (Sigh.)

He’s not wrong. Ecology has a history of valuing what you can learn staring intently at a backyards pond or an artificial pool near Palermo full of water bugs. We need to understand the ‘natural history’ well to understand systems. Very true, but I sometimes wonder how we advance. A massive NSF project to collect lots of large-scale, NEON, has not revolutionized much. For my own research, I think we need to cover greater temperature variation — and work harder to know what happens at the extremes (you have to wait a long time for plants to do something at 5C, but maybe we need to wait for that; at the other end many warming chambers don’t function about 30 C well, but that’s well below what’s too hot for most plants), we need more replicates and we need to collect better data, at finer scales.

Which is part of why I went to Milan. To pick my colleague’s brain about how my lab can best break out. He’s now building up new FabLab for his company’s new North American headquarters in Chicago, and had some useful ideas.

At the end of our meeting he asked if he should go to grad school, which struck me. It struck me for a lot of reasons, one is that some undergrads in my lab, who I think could bridge new technologies to ecology, are getting scooped up by start-ups digitizing agriculture, and putting their undergrad degrees on a possibly never-ending pause.

The same weekend I saw a colleague from my long-lost NCEAS days. In between nerd-crushing/raving about our colleague Jim Regetz, we discussed the apparent disconnect between the number of PhDs being awarded (erm, not sure about that verb) and the number of job openings where a PhD is critical. Some of his former postdocs were starting a new company, trying to make theoretical ecological models more useful to the point of underpinning a for-profit company.

I imagine academia often feels to be falling behind, but this weekend I felt it a little more acutely. We’re supposed to have the freedom and metaphorical space to be racing ahead. But it doesn’t feel that way when I can easily see why students in my lab would ‘pause’ undergrad to race around North America to improve how we harvest wheat, when we churn out publications faster and faster at the expense of the time it takes to really advance science (it’s so much quicker to grab a p-value than to develop a model with parameters you care about, then step back and gape at that the uncertainty around the estimates; p-values are so happy to hide your meaningful parameters and their uncertainty from you), or similarly churn out PhDs without a clear idea of their job prospects (hello Canada’s `HQP’). I am not so worried about folks in my lab, we train strongly in computational methods and how to design and answer useful questions—skills industry and beyond needs, but I worry about the future, and I could certainly train in this area better if there was more pressure, recognition, and support for it in ecology.

On the good news side, I enjoyed my take-out pizza from Milan for two glorious dinners!

Why not spend your February modeling cherry blossoms?

I am emerging—momentarily—from teaching to announce (this post is by Lizzie) …

The 1st International Cherry Blossom Prediction Competition!

We are pleased to announce a new international prediction competition “When will the cherry trees bloom?” Help scientists better understand the impacts of climate change (and we have prizes)! The competition is open to all.

Us competition organizers are providing all the publicly available data on the bloom date of cherry trees we could find. Competitors will use this data, in combination with any other publicly available data, to create reproducible predictions of the bloom dates at four locations around the globe.

The competition is open throughout February 2022 and seeks statisticians and data scientists of all levels, from experts to students just beginning to use statistical software. Complete submissions include a short narrative and a link to a publicly accessible Git repository.

For complete details or to contact the organizers, please visit https://competition.statistics.gmu.edu. A recording of the kickoff event is available on the competition website.

A big thanks to the American Statistical Association, Caucus for Women in Statistics, and George Mason University’s Department of Statistics and the Columbia Department of Statistics for their support, and partnerships with the International Society of Biometeorology, MeteoSwiss, USA National Phenology Network, and the Vancouver Cherry Blossom Festival—as well as Mason’s Institute for a Sustainable Earth, Institute for Digital InnovAtion, and the Department of Modern and Classical Languages.

Organizers: Jonathan Auerbach and David Kepplinger (George Mason University) and Elizabeth Wolkovich (University of British Columbia)

My problems with Superior

This post is by Lizzie.

I was once in a faculty meeting where a colleague attempted to explain that people are complicated. They are rarely simply good or bad. The same person can be intellectually brilliant in some regards and a revolting racist at the same time (Jim Watson was the subject in this case). I am not sure she made a dent in the perceptions of everyone who needed perhaps more than a dent in their perceptions, but it was a an excellent try.

This sort of complexity is sometimes a major pain. It means there are no easy answers where it would be really handy if we had them. And then we could organize things as simply good or bad, right or wrong. Such as research in human genetics, with its revolting and recurring history of eugenics, alongside its power to help us better diagnose and and potentially develop gene therapy treatments for a suite of diseases (such as Cystic Fibrosis).

And this sort of complexity is a major part of what was missing for me from the book Superior: The Return of Race Science by Angela Saini (reviewed by Andrew here).

I have a couple disclaimers about my feelings towards this book.

  • I read this book many months ago (in the spring of 2021 I believe) and am only now, at the end of the year, getting around to writing this up. So my memory is not as fresh as I would like. Part of the delay was feeling busy at work and part was not wanting to be negative about a book that in many ways takes on a hard subject and makes a compelling case, sharing information we should all know along the way.
  • I am a biologist. I am a community ecologist, so fairly removed from the genetics realm, but I find genetics work fascinating and my partner works in ecophylogenetics, so I am not as far removed as some. This makes me biased, as the book takes aim at biologists, but it also means I understand the science behind some of the claims.

I recommend Superior. It’s well-written and easy to read. It’s a super quick read of a topic we all should think and talk more about: race science, and how science can slip into supporting racism, nationalism and a suite of other evils that society cannot seem to rid itself of. It also captures a good dose of the sadder side of the process of science, including how the power of publication pushes people to over-reach and over-state.

I also disliked Superior for a few key reasons:

1. It seems to skip over all the really difficult questions in the topic of how we do science on human genetics, with no attempt to acknowledge that they are difficult questions or provide any answers. I spent a good half the book waiting for this to crop up. But it never did. My best guess is that Saini thinks we should simply stop doing 90% of the research we’re doing in human genetics.

There’s a related illogical mix of how the science is presented. She explains that the greatest human genetic diversity is found in Africa (which I presume all biologists know and I hope most people learn this early in school nowadays), then explains human migration patterns and that they show how often different populations of people in human history have interwoven. This is cool science! And stuff we learned through studying human genetics of different groups of people in different places. But then she condemns recent studies using similar methods, without fully defining the problem that made the old studies useful and the new studies evil. She seems to decry studies of ethnic groups without ever mentioning the utility of say, studies of the Ashkenazi Jews, which helped biologists find the genes linked to Cystic Fibrosis, Tay-Sachs and many other diseases. Science does good and bad. That’s the messy, tricky problem that I thought this book seemed to mostly ignore.

There’s actually an interesting issue here to me in the evolution of genetic/genomic science. From sometime in the 80s through sometime in the 2000’s genetics was really focused on finding single genes that did important things: caused Huntington’s disease for example, or changing the hair color of certain rodent groups. But at some point we ran out of finding those and I would say we didn’t find all that many. Most phenotypic traits might be more like height — which genetically is controlled by at least 50 genes, probably many more, in complicated ways that interact with the environment you’re raised in. And it takes a lot of data to find this all out. So finding out the genetic component of most traits is going to require a lot of data (and better methods) and wading far deeper into this tricky subject of how to do it ethically. But polygenic inheritance doesn’t appear much in the book and it’s not in the index.

Back to good and bad, scientists seem to be mainly good or bad also in this book. They are either out telling the world we’re all one people, or slipping rapidly and ignorantly into race science. The ones who find human history is a complicated story of migration and intermixing are never those charging for publications or doing anything scientifically questionable. David Reich “surprises [the author] with his gentleness. … He is unfailing polite, pausing only to message his wife.” He is never a researcher running a powerhouse lab potentially in part by colluding with other well-funded labs to outcompete the labs that won’t collaborate with him (as detailed here). I think he’s likely both, but scientists don’t seem to be so complicated as that in the book.

There’s no blurry line the author ever leads the readers to look at and wonder what to do about.

2. I thought statistics were not consistently reported. They were explained in depth when they supported the author’s argument and glossed over otherwise. For example, the author (who is British-Indian) writes “as much as 95% of [human genetic] variation is within population groups. Statistically this means that while I look nothing like the white British woman who lives next door to me in my apartment building, it’s perfectly possible for me to have more in common with her than with my Indian-born neighbor who lives downstairs.” Sure, it’s perfectly possible (for many reasons, given how little info we’re given here), but statistically it’s more likely that Saini will have more in common genetically with her Indian-born neighbor because of that other 5% which is never discussed. She introduces uncertainty intervals only for a recent study painfully trying to link a gene variant, which appeared to sweep through some European, North African, Middle-Eastern and Asian groups at a mean date of 5800 years ago, to the brain, but we never discuss the uncertainty for past human migration that mixed populations. I get that she’s making an argument, but I wanted it more carefully made. Complaining that scientists warp statistics for their own agenda is less convincing when an author does the same.

3. I am biased as a biologist myself, but I felt like the author had a holier than thou tone, looking down on us racist little biologists who could not understand the complexity of the social construct of race. I sort of liked this for my own personal benefit: it struck home for me the idea of why we need to all acknowledge our privilege and bias. I was annoyed with the author for not making it clear to me that she knew she was as much part of the problem, and that we’re all in it together.

At the end of the book she writes of biologists who use racial categories in their research: “They should at least know what race is” (author’s emphasis). This comes just after she writes about a:

…fairly young, diverse, international team, not all stuffy or old fashioned. And [the anthropologist studying them] noticed that all the scientists were routinely using racial categories not only to select their subjects but also to confidently pick out statistical differences between these racial groups.

So as [the anthropologist] observed them, she asked each scientist she interviewed one simple question `How would you define race?’

Not one of them could answer her question confidently or clearly.

I bet a bunch of them bombed answering this in a way that I’d be horrified by, but I didn’t trust the author at this point to think much more than that this is a damn hard question that would require a long answer about the social science of it, with a dash of biology (including the current inadequacy of both genetic data and statistics), and the mess of check-boxes on forms.

It’s a question that I wish Saini had given me a straight answer to in the book, and I don’t feel she did.

One answer is that race is a social construct and cannot be defined genetically. Saini seems to suggest the ills of racism spring from this attempt. But the ills of racism to me do not come from whether we define race socially or genetically, but in the idea that there could ever be a hierarchy of races. If we could define races well genetically, I doubt we rid ourselves of this inane idea.

Superior seems to want to take down the basic ideas of racism by showing that we are all one people, that we are highly intermixed and genetically similar. We are! But we are also genetically diverse. And, like we value cultural diversity, we should also value — not try to obscure — that diversity. In my work and the work of many researchers, genetic diversity is often what saves us; it’s the monocultures that fail us. I think Superior fails by not acknowledging how fantastic genetic diversity is, how much it has given us (and will in the future), and instead offers the false promise that if we all recognized race isn’t genetic then we wouldn’t be racist anymore.

Event frequencies and my dated MLB analogy

Apparently, it’s blog day!

This post is by Lizzie, and I am requesting analogy help (by the way, thanks for your recent help on how to teach simulation to students).

Yesterday morning I watched a little Metro-Vancouver parks worker trundling along in their tractor, as they gathered up the debris strewn across the beach from our recent storm. The storm had been fantastic fun to cycle home during and I snapped some photos on my ride that do not at all do justice to how riled up the ocean looked (one shown). It also triggered the now-almost-normal stream of requests to link “the severe weather effects we are seeing in [insert place] and how this relates to global warming.”

Which led me to trot out my now very old analogy to explain why we cannot generally attribute any one specific weather event (a specific storm, frost, heat wave etc.) to climate change: consider a MLB player, let’s call her Barry…. For the beginning years of her MLB career she was a pretty good hitter and every so often hit a home run. In the later part of her career she starts taking steroids and hits many more home runs on average. You can’t attribute any particular home run to Barry’s steroid use, but you can associate the changing frequency ….

I didn’t come up with this analogy. I copied it from someone who copied it from someone … and on and on until we find someone who thinks he invented it, but I bet he just forgot where he heard it.

And I like it! People generally get the connection and they are sometimes willing to let go of their urge to pressure me to stay, ‘Whoa! What a storm that was yesterday. That storm was caused by climate change, folks.’ And the steroids fits nicely with our juiced-up climate system so it’s often a good segue into what’s changing in our climate system.

But my analogy feels really out of date! It feels old and I think I lose people who try to remember back when or figure out what I am talking about. I am wondering if anyone has (and is willing to share) a better one they’re using, or wants to propose one I can use.

Research on heat extremes is moving towards terms such as ‘nearly impossible in the absence of warming‘ or ‘virtually impossible without human-caused climate change‘  so maybe I can shelve my example someday? But I am not ready for that. (For anyone waiting on rapid attribution of the PNW storm, I suspect World Weather Attribution is working on it.)

Climate change as a biological accelerator

Or, how leafout is like driving to grandma’s house. Or, how a 90 minute fake data simulation solved a problem my lab had spent over 3000 hours on. Or, why did almost none of my co-authors like my sentence about how ‘climate change steps on the biological accelerator’?

Or… wait for it — crows are black because crows are black.

This post is by Lizzie and it’s about a paper I wrote with Andrew, Jonathan Auerbach, Cat Chamberlain, Dan Buonaiuto, Ailene Ettinger and Ignacio Morales-Castilla on one explanation for why biological responses to temperature are declining in recent decades. If you’re most interested in the paper, I suggest you just read the paper as it’s much shorter than this post (1100 versus 2100 words). This post travels through the paper’s origins (including fake data simulation!) and its trip through the friendly and peer review process, with a quick overview of the paper’s findings too.

About six years ago, a paper was published called `Declining global warming effects on the phenology of spring leaf unfolding,’ which led to lots of excitement in my tiny field of plant phenology (for North American readers: phenology is the timing of recurring life history events, such as leafout, flowering, when birds lay eggs etc.).

The paper showed that what we call a `temperature sensitivity’ (change in days per degree of temperature, as measured by linear regression) was getting smaller in magnitude recently. For example, if back in the 1980s birch tree leafout would advance 5 days per 1 degree of temperature increase, today it might be just 4 days per degree. If the trend continued, someday trees’ leafout would stand still in time perhaps, or maybe reverse and start leafing out *later* in warmer years.

I was at a meeting in Turkey reading through the supplement trying to figure what could be up with the paper. I am always worried about time-series analyses, and especially using the data they did (PEP 725) which varies in quantity and location over time. But I didn’t find anything obvious and when I talked to my colleague, Mark Schwartz, he knocked out a few additional hypotheses I had.

The biological hypothesis for this effect was obvious to everyone in the field, and laid out in the paper: temperate plants’ leafout generally responds strongly to spring temperatures and that has been the dominant controller on leafout. But, underneath the hood, plants also cue to daylength and winter cool temperatures (`chilling’); someday, if spring warming comes so early that the days are really short, or if chilling gets too low, then plants will wait a little until they leafout. And this will lead to a smaller (in magnitude) temperature sensitivity. Functionally, the biology currently suggests that with short days or low chilling plants will require more spring warming to leafout; this higher required spring warming is the most proximate delay of leafout, with chilling or daylength causing the higher threshold.

This was all well known from experiments. For decades (centuries probably), biologists have been taking dormant tree branches in the winter and putting them in little boxes where we can set light and temperature to different amounts to test and re-test this model. It’s effectively a fancy version of bringing in pussy willows to your warm house in the spring, except imagine you put some in the window, some in a cool corner, some in a warm corner etc..

At about this time, I was actually starting a meta-analysis of such experiments to estimate the effects across species of spring warming, winter cool and daylength, and about four years later was scratching my head when the results didn’t jive with the declining sensitivity paradigm in the field. And at this point a good number of high profile papers had documented this effects (here’s just a couple from PNAS, and  Nature Climate Change). Working with some excellent folks in my lab (four of the co-authors listed) we’d been taking the model from our meta-analysis and checking what it predicted for regions seeing declining sensitivities, and we could not match predictions of `declining sensitivities’ unless we warmed up the world 4C or higher (it’s warmed about 1 to 2 C in Europe). Staring again at the underlying observational data I realized that a 1 degree warmer day when it’s cool is not the same as a 1 degree warmer day when it’s warmer.

Makes total sense, right?

No, it probably doesn’t and that’s about where I was for a bit. I quickly wrote up simulations that showed my hunch could be correct, but I struggled to explain what was going on. I spent a while calling it a ‘non-stationarity in the unit of day’ issue.

And around then, in pre-Covid times, I swung through New York in a land where people used to hang out in the their offices, grad students in stats even used to offer in-person open sit-down-here-next-to-me-and-tell-me-your-statistical-troubles stats advice hours. I spent a while explaining my model to Andrew as simply as I could, explaining what is often called the ‘bucket model’ of leafout: plants need a certain amount of spring warming to fill the bucket, when the bucket is full, the plant leafs out. With climate change, days are a little warmer, so they fill the bucket more each day — so if you use a metric, such as temperature sensitivity, with day in the denominator, then it has to go down as the world warms even though plants require the same thermal sum to leaf out (are you getting my ‘non-stationarity in the unit of day’ yet?). The whole thing works without ever invoking daylength or winter temperatures (chilling). I was wobbling through this explanation — I had a couple cobbled together graphs about the underlying effect of temperature on plant development, and then the ‘bucket model’ and then connecting back to the linear regression, all of which I was dragging Andrew through in the lounge of the stats department.

When he left to get something in his office I wandered over to open office hours (also in the lounge of the stats department) to chat with Jonathan Auerbach, who is great fun to work with, and started dragging him through my problem. Andrew returned and we started de novo coding my simulation code (of course much nicer, since Jonathan wrote it). However, while I generally pegged my simulation code to a smaller (more realistic) range, Jonathan wrote up his to cover a big range of temperatures — 5 to 30 C.

As the simulated data showed up on the screen both Andrew and Jonathan had the eureka moment — “any process observed or measured as the time until reaching a threshold is inversely proportional to the speed at which that threshold is approached.”

Getting to leafout is just like driving grandma’s house, as Andrew explained it. In this case, day of leafout is akin to how long it takes to get to grandma’s house, and temperature is akin to average speed: the relationship is inherently non-linear so comparing the effect of 1 degree warming at different points along the speed (temperature) axis will give different slopes. We all know that if you drive at 50 miles per hour, driving 5 miles per hour (10%) faster will have far less of an effect on your arrival time than if you were driving at 10 miles hour and drove 5 miles (50%) faster.

But somehow I, and most everyone in my field, did not see this connection.

Instead we felt intuitively (and damn strongly if you ask me) that the slope should be constant. If you ask people about the underlying model, they will mention it’s a threshold process, so we all seemed to agree on that, but also felt strongly that using a linear model was fine. We also have a strong intuition about what should happen to the slope of a linear regression if you raise that threshold (the way short days or low winter chill should) — it should go down in magnitude. But all these intuitions are just plain wrong. I fell victim to them like everyone else.

What’s been interesting to me is how hard it’s been to get people to question that intuition. In our paper we don’t say this is definitely the cause of declining sensitivities, as we don’t know. We do state that it’s a simple explanation for it, based on the biological model we all claimed to agree on. And we did suggest a pretty simple correction to try — just log your data before you run your linear regression (and we showed that logged data did not show any sign of a declining sensitivity). But we still got a lot of fascinating responses when sending the paper out to review. While some jumped on board, at least half did not, and many of them seemed to jump up and down beside the boat demanding we all stay with them on the land.

I was surprised how much folks dug in. Here’s some of the most common responses:

  1. Mainly I found people came up with new models we should try that would show the declining sensitivity: these included trying a shifting window to `find’ the best temperature window (aka, highest correlation), decreasing winter chill that drives a higher spring warming threshold. We’ve basically done all of these and only recreated declining sensitivities with extreme effects of daylength. The rest of the models don’t produce it (even though I get that we all feel they should). And even then the log estimates picked up the change more clearly than a linear model.
  2. Many folks did not like that we suggested a log transformation (one reviewer wrote, “why this particular data transformation was used is not apparent. There are many other data transformations that could have been used as alternatives, and these could be explored for how they would affect the results”) even though the log is the natural transform of an inverse, which we wrote. I never fully got a good explanation for this. It seems a weird response to me coming from biologists. But it seemed often to go with ideas for a new model (my previous point).
  3. Multiple reviewers said that we only did the analysis for two of the original seven tree species in Fu et al. 2015 paper and thus our analysis was uncompelling. (True! I was too lazy to fit the type of thoughtful model I would want for the uneven data for all seven species, so we just did the two with the most even data where we could hold the data constant over time and space; we tested it with another species when asked and it too showed that the declining sensitivity result went away).
  4. Lots of folks looked at the supplement and said the model we proposed was too confusing, even though it’s just the math of a simple bucket model I am otherwise told is very simple.

I tried to show people how simple the model was, sending along code:

data <- data.frame(leaf_date = numeric(0),
                   cum_temp  = numeric(0),
                   mean_temp = numeric(0),
                   threshold = numeric(0),
                   delta     = numeric(0))
threshold <- 1000 # thermal sum for leafout
for(delta in c(5, 10, 15, 20)) { # this is warming added on
    for(sim in 1:1000) {
	temp <- delta * (1:100) + rnorm(100, 0, 50)
	leaf_date <- which.min(cumsum(temp) < threshold)
	cum_temp <- sum(temp[1:leaf_date])
	mean_temp <- mean(temp[1:leaf_date])
        data <- rbind(data, data.frame(leaf_date, cum_temp, mean_temp, threshold, delta))
    } 
}

Once someone understood it, then they often would go to point 1 — they no longer thought that leafout happened after a certain thermal sum (if even they agreed strongly with that in an earlier email). It was fascinating!

I think though that my favorite reply was this one: “This MS provides a simple explanation for the observed non-linearity in biological temperature sensitivity… The MS develops a model to prove this point and it is very nicely written. Nevertheless, I found this somewhat uncompelling, because that the explanation appears to be the same as the observation (e.g. crows are black, because crows are black).”

There were also reviewers who wanted to skip over the issue we were discussing and find interesting new biology in the results. At PNAS one reviewer was adamant that we talk in depth about differences we had not deemed very different (slope of -0.17 versus -0.20, both with large uncertainty intervals).

I am happy to say the paper is finally published, and the last reviewer pointed us to this fascinating paper dating back several years, before Fu et al. 2015: “On the uncertainty of phenological responses to climate change, and implications for a terrestrial biosphere model,” which shows a MODEL that produces the declining sensitivity problem, and the reviewer asked, “If the model structure and parameters are fixed in time (suggesting the biology is also fixed), how can the temperature sensitivity change (which would imply that the biology ISN’T fixed)? Is the temperature sensitivity metric itself flawed?”

I’d like to thank Faith Jones for the artwork, my friend and colleague Caroline Tucker, for alerting me the paper was out (I wrote this post a little while ago), which she figured out via Twitter (thanks to this Tweet, thanks also to Alexa Fredston for the tweet). And I’d like to thank They Might be Giants for the paper’s theme song.

‘No regulatory body uses Bayesian statistics to make decisions’

This post is by Lizzie. I also took the kitten photo — there’s a white paw taking up much of the foreground and a little gray tail in the background. As this post is about uncertainty, I thought maybe it worked.

I was back east for work in June, drifting from Boston to Hanover, New Hampshire and seeing a couple colleagues along the way. These meetings were always outside, often in the early evenings, and so they sit in my mind with the lovely luster of nice spring weather in the northeast, with the sun glinting in at just the right angle.

One meeting was sitting on a little sloping patch of grass in a backyard in Arlington, where I was chatting with a former postdoc, who now works for a consulting company tightly intertwined with US government. When he was in my lab he and I learned Bayesian statistics (and Stan), and I asked him how much he was using Bayesian approaches. He smiled slyly at me and told me a story about a recent meeting he was at where one of the senior people said:

“No regulatory body uses Bayesian statistics to make decisions.”

He quickly added that he’s not at all sure this is true, but that it encapsulates a perspective that is not uncommon in his world.

The next meeting was next to the Connecticut river and with a senior ecologist, who works on issues with some real policy implications: how to manage beetle populations as they take off for the north with warming (hello, or should I say goodbye, New Jersey pine barrens), the thawing Arctic, and more. I was asking him if he thought this statement was true, which he didn’t answer, but set off on a different declaratory statement:

“The problem with Bayesian statistics is their emphasis on uncertainty.”

Ah. Uncertainty. Do you think uncertainty is the most commonly used word in the title of blog posts here? (Some recent posts here, here and here.)

In response to my colleague I may have blurted out something like ‘but I love uncertainty!’ or ‘that is a great thing about Bayesian!’ and so the conversation veered deeply into a ditch, from which I am not sure that it ever recovered. I said something along the lines of, isn’t it better to have all that uncertainty out in the middle of the room? Rather than trying to fit in under the cushions of the sofa as I feel so many ecologists do when they do their models in sequential steps, dropping off uncertainty along the way (often using p-values of delta AIC values of 2 or…) to drive ahead to their imaginary land of near-certainty? (I know at some point I also poorly steered it towards my thoughts on whether climate change scientists have done themselves a service or disservice in shying away from communicating uncertainty; I regret that.)

We left mired in the muck that so many of the ecologists around me feel about Bayesian — too much emphasis on uncertainty, too little concrete information that could lead to decision making.

So I pose this back to you all: what should I have said in response to either of these remarks? I am looking for excellent information, and persuasive viewpoints.

I’ll open the floor with what I thought a good reply from Michael Betancourt for the first quote: fisheries, and that Bayesian gives better options to steer policy. For example, if you want maximum sustainable yield without crashing a fish stock, you can more easily suggest a quantile of catch that puts you a little more firmly in ‘non-crashing’ outcome.

This system too often rewards cronyism rather than hard work or creativity — and perpetuates the gross inequalities in representation …

This post is by Lizzie. I started this a while ago, but Andrew’s Doll House post pushed me to finally get it up on the blog.

The above quote comes from a recent article on the revelation that the person Philip Roth decided should write his authorized biography has a history of sexual harassment accusations (I mean, the irony…). It reads more fully, “this system too often rewards cronyism rather than hard work or creativity — and perpetuates the gross inequalities in representation that disfigure the American literary landscape”, but I think it certainly applies to lots of other `landscapes’ including the one within which I exist much of the time: ecology and evolutionary biology (or EEB).

It relates to a quote that has been rolling around in my head for many months now, from a student in my lab: “It feels like the message is that the careers of these [double-digit] people were worth less than the career of this one [purportedly] brilliant man.”

And I really did not have a good response to this student other than ‘ummm.’ I have asked around for a better response from my colleagues around me and I am not sure any of us have one.

It’s a relevant question, that could be posed for many publicized and less publicized events in EEB. Two recent publicized examples include the just-risen star of Jonathan Pruitt, a spider behavioral biologist, who has been accused of fabricating most of his high-profile data, passing out said data for junior folk to write up with him as senior author, with multiple papers now retracted or expressions of concern issued. On the rising-star front is Denon Start. He’s had recent accusations of suspicious data, but before those he had two awards abruptly (to me at least) rescinded from the Canadian Society for Ecology and Evolutionary Biology (CSEE). You can fall down the Twittersphere to try to figure out why, but CSEE never said.

There are similarities and differences in these cases galore, but the one similarity that has grabbed me is how much those around me, especially my faculty colleagues, seem to pin as much blame as possible on each man — I get the feeling 99.9% would be a good amount to many, leaving just enough room to acknowledge ‘we could all and should all do better.’ I agree a lot of blame falls on the perpetrators, but I also think that, by not leaving enough room to blame ourselves and the community we create, we open the gates for the behavior to continue.

Both of my examples had remarkable publication records (potentially falling under The Armstrong Principle), but they also were promoted by many people to get where they were (and are perhaps, Pruitt at least seems to still be employed by McMaster). Academia seems good at passing along and promoting people where we should perhaps have cause for concern. This to me means that either we are too out of it to know what is happening, or we look the other way because it seems easier, and it protects the careers of the rising star and all those connected to their star. There’s always some talk about the former — how can we build a better community where senior people know what is going on and can intervene? But what I want to talk about is: How do we hold a community of researchers accountable when they may have known there were concerns?

I suggest a step forward to is to hold the letter writers more accountable. Academia does function on reputation and a big part of your reputation is formalized in your letters of reference. American letters are often so positive as to feel almost useless. But they’re not. When they are short, brief or otherwise feel perfunctory, they can say a lot. Here’s a letter that might raise eyebrows:

Dear committee of special award or position:

I have known [X] in [this way] for [this long]. S/he/they have published [Y] papers, taught (or TA-ed) Z classes.

If you have more questions about this applicant, I encourage you to call me at the number below.

Sincerely,
Dr. Especially-eminent

So, that’s something we could start to do if you ask me. I also suggest the following after-the-fact potential actions:

(1) If you read a glowing letter for someone whom shortly after you hear had major concerns, go look at that letter. Did you miss something? If it seems like you didn’t, why not call up the letter writers and ask about the disconnect? Letter writers feel more okay writing these letters because there is generally no consequence for their careers.

(2) We could formalize some of this. Department chairs who receive glowing letters and there are issues later could be expected to contact the department who employs the letter writer and express some concern. We effectively write reference letters as part of our service, so it’s part of our job; if we’re doing that part of our job poorly shouldn’t it count against us somehow?

(3) The other thing I suggest we all do, after the fact, is be very cautious in how much we push back on whether the system needs to change by focusing on the perpetrator or fears that someone or some organization will be sued. I don’t think it feels like you’re saying you don’t want the system to change when you focus so strongly on the individual perpetrator or when you say, as many did to me, ‘what if [insert some society or senior person] is sued?’ These are both important things to do, but when they are most of what you do — then you just did trade in these action items of fear and blame for an actual closer look at the system.

Which brings me back to my student. Was my student reading the message correctly? ‘That the careers of these [double-digit] people were worth less* than the career of this one brilliant man’ in the particular case we were discussing.

It’s a good question to ask, if you ask me, as there are some implicit weighting and numerical assumptions here. Every time we don’t question the letter writers, or worry about what will happen to some established person or society, I think we effectively do send the message my student felt they had received. I even heard a colleague recently say we should not scrutinize too strongly the high fliers with the many, many publications, because ‘what if we discourage them? What if we lost [insert name of new NAS member] to that extra scrutiny?’ I have two replies to this. One is that if they are that great they will stand up to a little extra scrutiny.

And the other is to think more on how implicitly we undervalue those we lose when we say this. Ask people to explicitly count up those we lost along the way — who drop out, leave or otherwise are valued less, and how we value the creativity and exciting science we never got to see from them. We either think we aren’t losing many, or that they’re not worth the one great man we saved.

*In earlier version of this post I mistakenly wrote “worth more than the career of this one brilliant man,” which led to much understandable confusion.

What is the landscape of uncertainty outside the clinical trial’s methods?

I live in the province of British Columbia in the country of Canada (right, this post is not by Andrew, it is by Lizzie). Recently one of our top provincial health officials, Dr. Bonnie Henry, has received extra scrutiny based on her decision to delay second doses of the vaccine. The general argument against this is the one I have heard from Dr. Fauci of the US, who has been various levels of adamant that you do what the trial did (I would say very adamant, very adamant, very adamant, then slightly less adamant after the kerfuffle with the UK). You don’t deviate from the methods of the trial.

This has got me wondering what the landscape of uncertainty looks like as you move away from the methods of a clinical trial. And what progress we’ve made — if any — on this in the last couple of decades, when I first realized how stark the divide between inside and outside the trial methods is for many.

Over 15 years ago I was helping take care of a 50-year-old family member who had cancer and was struggling to get through a 6 week regime of radiation + chemotherapy at a major cancer institute in Boston. She had gotten through the first couple of weeks okay, even driving herself the 6-8+ round-trip hours from her home to the institute five days a week for her daily radiation appointments. But things got progressively worse in the third and following weeks (when I was her trusty chauffeur and companion). By her fifth week she was in and out of the ER with various major issues and was receiving various infusions to attempt to prop up her system so she could survive the next dose of radiation. Every day before radiation she needed a series of tests followed by a visit to her oncologist to get approval for that day’s dose of radiation, and this did not seem out of the ordinary for the later weeks of high-dose radiation therapy.

At one visit, when things were going particularly poorly, the radiation oncologist was brought in to consult on whether to continue treatment. He was advising for continuing, though it would be hard. It took a lot of her energy to speak, so she was often quiet, but on this day she asked him: ‘why do I have to do this? I have done this for most of the 30 visits, why do these last few matter so much?’ And he told her the truth — “because we only have data on the people who get the full dose. We don’t know what happens if you don’t take the full dose, or you take a few days off before continuing.” It was very helpful. I remember she said something along the lines of ‘okay,’ and we drove home in semi-shock, but at least we knew why they were pushing for this now. It was always her choice, but until then neither of us realized how gaping the uncertainty was between, say, going for 27 of your total 30 radiation visits, and going for all 30.

Clearly, the ethics matter, and that’s especially clear with a highly infectious and deadly disease like Covid. I assume that many of these deviant public health officials who have delayed second doses have done the simple SIR model math and figured out that: (n higher number of people vaccinated at X% efficacy given Y weeks of delaying the second dose)*(black box uncertainty as you deviate away from the trial methods)=likely more lives saved. Henry has cited studies showing >90% efficacy for the three weeks after the first dose of the Moderna and Pfizer vaccines, so I suspect she’s feeling good that her X in extending the second dose to four months is still fairly high and thus has some internal estimate on the landscape of uncertainty beyond the methods of the trial, and there’s growing data on this.

But if you listen to various interviews with Fauci and other public health officials, I start drifting into memories of discussions that start with, ‘what is the variance of a fixed effect? It’s either 0 or infinity.’ Now, I don’t mean exactly that — but I do mean there seems to be a large gap in perspectives here. In one — the clinical trial methods must be followed to a T, until a new or properly vetted trial of any deviation is approved, conducted and reviewed. And in another — some adjustments happen given the potential for lives saved despite the uncertainty and that ‘population health data’ is then used to make further adjustments on the fly (in conjunction with other ways of viewing the clinical trial data you have).

These debates have made me wonder what progress have we made addressing this uncertainty from both a bioethics, and data collection and design standpoint? I am not (at all) a bioethicist but the rigid adherence to the trial methods doesn’t feel terribly ethical to me, and I think Covid has highlighted that. So I wonder how much has changed in last 10, 20 or 30 years of how those who deviate or ‘drop out’ of clinical trials are handled as datapoints. Are they required to be tracked? Or is it better to save money by focusing only on those who follow the trial perfectly? Is there an incentive for research or new methods or databases that compile these deviants to start fleshing out that landscape of uncertainty beyond the clinical trial methods? Or is everything beyond basically zero, or maybe infinity? Or maybe somewhere in between.

What is/are bad data?

This post is by Lizzie, I also took the picture of the cats.

I was talking to a colleague about a recent paper, which has some issues, but I was a bit surprised by her response that one of the real issues was that it ‘just uses bad data.’ I snapped back reflexively, ‘it’s not bad data, it just needs a better analysis.’

But it got me wondering, what is bad data?

I think ‘bad data’ is fake or falsified you don’t know is fake/falsified. Fake data is great, and I think real data is a fine thing. I don’t know that data can be good or bad, can it? It can be more or less accurate or precise. I can think of things that make it higher or lower quality, but I don’t think we should be assigning data as ‘bad’ or ‘good,’ and labeling it as such to me suggests a misguided relationship to data (which I do think we have in ecology).

The paper uses data from the Living Planet Index (LPI):

…is a measure of the state of the world’s biological diversity based on population trends of vertebrate species from terrestrial, freshwater and marine habitats. The LPI has been adopted by the Convention of Biological Diversity (CBD) as an indicator of progress towards its 2011-2020 target to ‘take effective and urgent action to halt the loss of biodiversity’.

These data have been used a lot to show wild species populations are declining. The new paper purports to show that most wild populations are not declining. The authors use a mixture model to find three parts to their mixture (I don’t know mixture models so encourage anyone who does to take a look and correct me, or just comment generally on the approach, but as best I can tell they looked a priori for three parts to the mixture) and then they took out the ‘extreme’ declining ones and find the rest don’t decline. Okay. Then they wrote a press release saying most populations aren’t declining and we should all be hopeful. Yay.

But the LPI data are best described (by Jonathan Davies) as polling data. We can’t measure most wild species populations and we tend to measure ones from well monitored areas, which are often less biodiverse (I could start saying which animal populations are like which human groups that answer polls a lot, but I will resist). The LPI knows this and they generally do a weighting scheme to try to correct for it (which isn’t great either), but this paper doesn’t seem to try to do much to correct the data. One author wrote a blog about it, noting:

We also note that these declines are more likely in regions that have a larger number of species. This is why the Living Planet Index uses a weighting system, otherwise it would be heavily weighted towards well-monitored locations.

I wish these researchers and others in biodiversity science would use some of the thoughtful-stratification and other approaches used on polling data to try to give us a useful estimate of the state of global biodiversity, instead of press-release-friendly estimates. If anyone needs a project, the data are public!

To all the reviewers we’ve loved before

This post is by Lizzie (I might forget to say that again, when I forget you can see it in the little blue text under the title, or you might just notice it as out of form).

For the end of the year I am saluting the favorite review I received in 2020.

This comes from a paper that included a hierarchical model where we partially pooled by plant species. We had a low and variable replicate number per species, but a pretty good sample size across all species, and we wanted to estimate effects of experimental treatments (things like ‘warm temperature’ and ‘cool temperature’) across species.

We were working on invasive species (species native to somewhere else, in this case Europe, that have been introduced and grow quite well in that somewhere new — in this case North America). There’s a lot of interest in whether evolution post-introduction happens, more specifically evolution that helps the plants do so well somewhere new. We were looking for it in their germination response, by growing seeds we collected in North America or Europe (the seeds’ ‘origin’) in different conditions.

We didn’t find much of an effect of origin. We found effects of our treatments and a few other things, but we didn’t find an origin effect. Maybe because there isn’t a big origin effect, maybe because of the species we picked, or maybe because of lots of things.

In the main text we showed our model estimates, and showed raw data plots in the supplement. I generally think you should try to show raw data in the paper when possible, but this is the first time I got this response from doing it. The reviewer writes:

From [your main text model estimates figure] the main takeaway point that I could garner is that increasing temperatures impact growth and germination speed across species… I find Figure [in the supplement showing the raw data] to be interesting, because the raw data [often for the species with really low sample sizes] suggests to me that you may actually have some origin differences for some species in certain environmental contexts, which may not have come through so clearly in your global, many- leveled models [then details on what specific treatments and for which species this reviewer has discovered some trends].

It’s great to have robust models, but I think it’s worth taking a look at your data and ask whether some more interesting or nuanced stories might come out.

This is the first time I recall when I felt a reviewer was actually taking me by the hand and leading me back to the center of the garden and saying, ‘might you please consider this alternative path? Instead of going left at the fork, perhaps if you go right … you could get the origin effect we all so want to see.’

And to all the reviewers who’ve shared their thoughts
Who now are someone’s else’s reviewer
For helping me to grow
I owe a lot I know….

What George Michael’s song Freedom! was really about

I present an alternative reading of George Michael’s 1990’s hit song Freedom! While many interpret this song as about Michael’s struggles with fame in an industry that constantly aimed to warp his true identity, it can also be interpreted as a researcher progressing in a field where data ownership and data ‘rights’ are still hotly contested.

Heaven knows I was just a young boy
Didn’t know what I wanted to be
I was every little hungry schoolgirl’s pride and joy
And I guess it was enough for me

In these first lines the researcher describes the heady days of early grad school, where most folks were still in their twenties, enjoying a cadre of new friends, and still not sure of their future.

To win the race, a prettier face
Brand new clothes and a big fat place
On your rock and roll TV
But today the way I play the game is not the same, no way
Think I’m gonna get me some happy

The researcher describes a few things here: moving on to their postdoc, the excitement of their first few conferences where they got to present their exciting new results, but also hints at a big change in how they are approaching science recently. A change that could bring great happiness.

 I think there’s something you should know
(I think it’s time I told you so)
There’s something deep inside of me
(There’s someone else I’ve got to be)
Take back your picture in a frame
(Take back your singing in the rain)
I just hope you understand
Sometimes the clothes do not make the man

The researcher is nervous about what they will share. They know it is not the view of many in the field and they hope others will understand. Even though they may wear the tevas-with-socks and pleated shorts of an ecologist, they have views they fear will not be widely accepted by their community.

All we have to do now
Is take these lies and make them true somehow
All we have to see
Is that I don’t belong to you and you don’t belong to me

Here the researcher sings out their truth! They stare down the lies they have repeatedly heard, including:

  • Data you collected are owned by you and you should hold onto them possessively.
  • If you publish your data or don’t tightly guard it, it will be stolen by others, and then your career may be ruined.
  • People build entire careers on reusing other people’s data and they get more fame and recognition than those who toil away collecting data and are not recognized for their efforts.
  • Your data can never be fully understood without your presence, and thus should probably not be used without you around in some way.
  • If we all publish the data we collected regularly we will never have good data again, and we will never ever have long-term data because people will stop collecting long-term data.

The researcher sings out to their colleagues that these are lies and that for science to progress, data should be free, that data don’t ‘belong’ to any of us. They encourage their colleagues to let go of possessiveness (it never makes you happy!) around data.

Freedom (I won’t let you down)
Freedom (I will not give you up)
Freedom (Gotta have some faith in the sound)
You’ve got to give what you take (It’s the one good thing that I’ve got)
Freedom (I won’t let you down)
Freedom (So please don’t give me up)
Freedom (‘Cause I would really)
You’ve got to give what you take (really love to stick around)

Here they sing out for data freedom (“Freedom!”), alternating with pulls they have felt from colleagues who believe data sharing may destroy the field (“I will not give you [data] up”).

Heaven knows we sure had some fun, boy
What a kick just a buddy and me
We had every big-shot goodtime band on the run, boy
We were living in a fantasy

The researcher again looks back fondly on their PhD, remembering days in the field when they pulled on their waders, grabbed their plastic bucket and collected data (for example, see opening images here), and then published their first exciting papers.

We won the race, got out of the place
Went back home, got a brand new face for the boys on MTV (Boys on MTV)
But today the way I play the game has got to change, oh yeah
Now I’m gonna get myself happy

The researcher remembers wrapping up their PhD, submitting a great Dance Your PhD (here’s a favorite example), and then returns to their refrain on realizing that their field must change. Both for the field and for personal happiness.

The chorus repeats ….

All we have to do now
Is take these lies and make them true somehow
All we have to see
Is that I don’t belong to you and you don’t belong to me
Freedom!
Freedom!
Freedom!
It’s the one good thing that I’ve got

Sing it with me! Data freedom! Data freedom! If you love ‘your’ data set it free!

And if you’re an ecologist or in any similar field with a contingent of folks who speak some of the lies mentioned above I encourage to ask for examples. Ask for the list of people whose careers have been ruined by data sharing, ask also for the list of happy people who publish data they collect—and try to actually figure out what the distribution of these ruined versus non-ruined people looks like. If they tell you someday your field will be destroyed, ask them for examples of other fields where people have been made to share data (think GenBank, parts of medicine, please help me expand this list!) and what actually happened.