The mantra and mania of data sharing

This post is by Lizzie. The photo is from a photos folder I found from my PhD called ‘favorite stake photos.’

I wrote this post before seeing Andrew’s post for today.

When I was a grad student I spent a remarkable amount of time wandering around Sweetwater National Wildlife Refuge visiting 56 shrubs. Over and over again. I visited the shrubs almost every day. It wasn’t wandering, it was a structured, efficient route into the site, hitting each `replicate,’ then down the hill, up the next. Some days I was opening or collecting pitfall traps under the shrubs, some days I was vacuuming the shrubs (both of these tasks were to collect arthropods), other days I measured soil respiration, collected soil samples under the shrubs, I took clippings of the shrubs. There was also climate data to collect under the shrubs, little litter bags I was variously adding and removing under the shrubs. Vegetation sampling! I have forgotten much of it but I recorded a daily log so it can come flooding back. It felt like a lot of work.

In the end I published four papers about those 56 shrubs (which became 54 after a fire). Stuff about invasive grasses and carbon cycling, and bugs and stuff. Solid work, science maybe inched forward? Maybe it stepped right, but because of all the data I think it inched forward.

After grad school I joined a sort of think-tank that had been funded by NSF to promote data synthesis in ecology. Ecology needed it. It was (is?) a bit of a stuck field with a cacophony of individual studies in different places with different shrubs (or in lakes, forests etc.) — ‘boots and bucket’ ecology I heard it called.

You put on your boots, grabbed your bucket and — voila! Ecological science. The think tank was a renegade endeavor, trying to make sense of all the individual studies, by looking across them for patterns and maybe even testing some theory now and then. There was a tension between ‘boots and bucket’ ecologists at the time and ‘synthesis’ ecologists. According to the ‘boots and bucket’ tribe the `synthesis’ ecologists were stealing all their data for flashy papers. They were upsetting the order of things. Some said they were getting it all wrong because if you didn’t collect the data, you didn’t know the system enough, you could never figure anything out (yes, let’s all take a minute and think about where a field with this idea would be headed). Others said data would stop being collected and everyone would just do ‘synthetic ecology’ and never have much data.

This world was swirling far above and away from me and my 56 shrubs at Sweetwater, but at the think-tank they gathered all the new postdocs and told them about the power of data sharing. Science advances if we share data! Think of the questions we could answer if all the data were shared and organized! It just takes a hour or so to post your data. Go ahead, post your PhD data!

I was totally in. I posted all my PhD data and I dove in on the power of data sharing and wrote a paper about it for climate change biologists. I found papers showing that the massive improvements in pediatric oncology (for leukemia it went from 4 to 94% survival) could be attributed in part to data sharing. I read up on GenBank and drooled at a field so close to mine in topic but so far away in data sharing — and also trounces ecology in finding important interesting science IMHO. I felt like scientists should take an oath to advance science, and if they took that oath, then clearly they would see that they have to share “their” data. We’re trying to mitigate climate change people! Share your data!

Fast forward 20 years and all the scare tactics of the anti-data sharing folks have not come true. There’s no drop in data. (Though I got this argument recently from a marine biology postdoc, who then retreated a little from the premise when I asked for data on the declining data — given, and she did manage to agree, that this had been happening for 20 years at least so shouldn’t we see the pattern? — she then said data is only be propped up by PhD students who have to collect it for their PIs, so I guess suggesting a radical shift in how data are collected? And some verifiable decline in other data types? Through I didn’t try to steer the argument anymore.) Journals require data sharing. Granting agencies do. I think the synthetic ecologists won.

But a bunch of folks — beyond that one marine biologist postdoc — missed the message. I have been running into major governmental and non-profit data-collecting agencies that will not share data over the last two years.

For today, I will tell you about just one of them.

It’s the Canadian Forest Service (CFS). My lab recently contacted them for a big tree growth responses to climate across western North America analysis we’re doing. We have a lot of data, because these data in the US are generally public. They’re either on the ITRDB or they were uploaded with papers (there are certainly some that are not shared, but I like think they all will be, as the USFS and related US agencies do usually have a mandate to share data), but I happened to know that the CFS usually does not share data without co-authorship. They don’t share plot level data, they don’t share tree ring data, they don’t share data unless you sign an agreement with them and guarantee them co-authorship (and some other weird stuff that sort of sounds like they control whether you can publish what you find or not, but I think they have had to back off on that, so now there’s just related smushy language I suspect).

We asked anyway. I told my lab it’s important to ask and not just assume (even if everyone has told you that you will not get the data without co-authorship) and here’s the reply I got:

We are particularly interested in collaborating with researchers who bring expertise in Bayesian methods to help advance our analyses and explore future growth projections. With that in mind, we would like to explore the possibility of a scholarly collaboration with you.

Entering into collaboration would help streamline access to tree-ring data across Canada by removing certain barriers. Some of the data you requested are under restricted use, with licenses granted only to CFS researchers. Others require external requestors to obtain authorization from the original data owners—a process that can be time-consuming. Additionally, some datasets (highlighted in red below) have not yet been published, and we are actively encouraging collaborative projects that incorporate these data.

To provide further context, it is common practice for National Forest Inventory (NFI) data to have access restrictions, particularly for raw or highly detailed data. These restrictions are in place for several important reasons. NFI plots are located on both public and private lands. Disclosing precise locations could compromise data integrity or infringe on landowner privacy. Some datasets also contain sensitive ecological or proprietary information. Also, NFI data are designed to provide an unbiased representation of forest resources. Controlled access helps prevent misuse or misinterpretation, especially given the complexity of the data and the ongoing updates and revisions. NFI data support national and international reporting obligations, policy development, and collaborative efforts across jurisdictions. Ensuring consistent and validated use is essential. Finally, while publicly funded, the collection and processing of NFI data represent a significant investment (CFS & NSERC). Responsible dissemination protects this investment and ensures proper attribution and use.

I am working on a reply to this and open to all ideas/suggestions. I’ll give you what I have so far, vaguely in order of the arguments they have given (which is probably not the best order).

  • I appreciate their reply and understand their perspective, but requiring collaboration for access to data slows scientific progress, reduces equity and diversity in access to data, and has never been shown to be helpful or beneficial to science to publishing robust results, and thus is something my lab has a policy against (we do).
  • If you have data you have had for a while and not published (7 of the 30 datasets we asked about), but would like the data analyzed, then publish the data. This is the best way to get data analyzed and then it will likely be analyzed by different teams of researchers so CFS would get maximum insights from the data.
  • Sensitive data can be fuzzed, jittered or otherwise changed enough to meet privacy standards but still allow others to use the data. Certainly for our purposes, given the grid-size of the climate data we’re using it is hard to imagine this would not be possible.
  • The best way to get data cleaned, corrected and properly interpreted is to share it widely. The more eyes on the data, the quicker these issues can be spotted and fixed. Further, lack of access suggests there is something to hide, which is extremely concerning.
  • It is precisely because these data are used for national and international policy that making them public seems critical. (Can anyone help me here? This seems so obvious that I am not sure how else to say it.)
  • Data is far more widely used (and cited) when publicly shared. More papers and research seems like a better return on the Canadian taxpayers’ investment, no?
  • If you really want to charge for the data, then charge for it — but make access of those data available to all who can pay.
  • Fundamentally, there is a large number of researchers — Canadian and otherwise — who would use the CFS data and don’t because of this policy. These are excellent researchers who simply either do not have time for the efforts of collaboration with one team that requires collaboration in exchange for data when all other teams make the data public or do not want to support this process because it slows progress in science and the more researchers who sign onto it, the more it is tacitly condoned. Lots of good scientists I know — myself perhaps soon to be one — will not use CFS data because of the current access policies. Or, if they use it once, they won’t use it again.
  • Taxpayers paid for these data to be collected, they should get to see it and use it how they please. And with the way things are going, I would add that — if the commitment is to data quality — the more people who can access and download it now, the better. Political regimes of the worst kind often remove and restrict data.

I don’t think CFS researchers have anything to hide. They run a really nice database of their data (I know, because you get to search around it to request the data) and they are helpful and sharp when I meet with them or we correspond over email. I also don’t think widely incorrect papers or policies have been prevented by this restricted access. But I think that some researchers have been fed a steady diet to make them fear these possibilities and I am not sure how to disabuse them of this version of the world.

I hang out with a lot of people who share their data — climatologists often share it, and folks related to them (e.g., dendrochronologists in the US), phenology people usually are better (though I have had recent issues) — and I hang out with folks who don’t.

The people who share data are happier. They don’t spend time telling me all the horrible things that will happen if they share data. They don’t spend their time worrying about it. They just share their data and move on.

It’s like people who spend all their time talking about work-life balance. I find them much less happy than the people working until 11pm some days — those folks are often also the ones tango dancing until 1am the next night, or leaving on multi-day kayak trips or getting in a Truck Surf hotel to tool around Morocco surfing. The ones talking about work life balance seem to set on searching for something I think they would find if they stopped searching for it so desperately.

 

76 thoughts on “The mantra and mania of data sharing

  1. I am not sure what you want to achieve. Could you state your objective clearly? This would help us to check whether your message is hitting the right notes.
    For example, if you want to vent about data sharing, we can probably find a few more reasons. If you want to persuade them to share data that they are currently unwilling to share, I recommend killing them with kindness. If you think they are thinking about the matter rationally, use cold, hard reasoning. If you think they have concerns that you are currently unaware of, practise active listening. I could go on, but clearly defining your goal and assumptions about their state of mind I believe would really help build your case.
    And now to pull an Andrew Gelman: composing this message is basically modelling the entire data generation process, incorporating both point estimates and uncertainty around them! (The Andrew Gelman in here is that there is a not initially obvious connection between ‘the story’ and statistical modelling. I submit this observation respectfully.)

      • Lizzie:

        I’m with you on that. I remember once asking a team of researchers for their (anonymized) data from a psychology experiment, and they refused to share, not because of confidentiality or anything like that, but just because they were concerned that I’d use the data to dispute the claims in their published paper.

    • That’s glib, but also ignorant. With even a moment’s thought, you should realize it isn’t the trees getting confidentiality, it is the land-owners who allowed CFS to survey forests on their land. That’s how the US Forest Service FIA (Forest Inventory & Analysis) works: plot-level data were not shared, and now plot-level data but not actual plot locations are available. They have an additional concern of land managers targeting their plots for different management and thus biasing their population estimates, so they don’t even let NPS know the precise locations of FIA plots in parks, even though NPS surveys their own additional FIA-compatible parks because they need greater sample intensity to make estimates at smaller scales (parks or management units, v state & regional for FS).

      I’m a strong advocate for open data and FAIR principles (Findable, Accessible, Interoperable, and Reusable). My earlier career combined identifying patterns in existing long-term data with conducting short term manipulative experiments to test potential mechanisms. But several of Lizzy’s bullets go against my current experience on the agency side. I’ll try to write a cogent reply later.

      • Tom:

        I fully agree that I’m ignorant of forestry and I’m sure you’re right that there are cases where there are legitimate confidentiality or data quality concerns, so good point. I’ve seen a zillion examples where people give bogus reasons, or no reasons at all, for not sharing data, so that’s what I usually assume. But, fair enough, this isn’t always the case.

  2. Clearly those who have figured out work-life balance – tango dancing until 1am, surfing in Morocco – don’t have kids. Is that the solution?

        • John:

          I agree that the problem is real–not for me but for my grandparents who had to work all the time, for sure. I just think the phrase, “work-life balance,” doesn’t help.

    • Most of the people I talk with who seemed concerned about work-life balance are younger and do not have kids so I am comparing among that population — kid-free 20 (to sometimes 30) somethings who talk a lot about work life balance but never seem to be rushing off to dance class (or https://trucksurfhotel.com/) vs those who never mention it and may not work a 9 to 5 schedule but appear to be doing a lot of non-work living.

      • That type of agreement doesn’t come out of nowhere so I am guessing that they have either been badly burned by sharing data or they are protecting non-government interests e.g. they sample on private land (individual or business) or indigenous land.

        But it’s pretty tone deaf to ask Canadians to hand over data solely under the provisions you want when your president is threatening to annex Canada as a 51rst state. American exceptionalism and entitlement at it’s finest.

        • The “Sweetwater National Wildlife Refuge” is in San Diego so I did have some reasonable grounds to believe that Lizzie was American. I apologise to Lizzie for the tone deaf comment – I should have investigated further.

          But as someone working in Canada, if not being Canadian herself, she must be aware of the issues around Indigenous data sovereignty that is probably informing a lot of limitations around data usage.

          The taxpayer may pay for data collection but ordinary citizens volunteer or, in some cases, are compelled to provide information and their rights for privacy and anonymity should be the priority.

        • But by my own bumbling, I think I proved a point. When you are only provided with supplied information and don’t see the full context of the data, you can draw conclusions that are wrong-headed and potentially offensive.

        • As a Canadian I would generally agree with this. However, as Chris points out, Lizzie is at UBC so this criticism doesn’t quite wash. If the issue was truly third party confidentiality (or agreements) the data likely wouldn’t be for sale by co-authorship; it would be a hard no, and that would be okay. However, if the data is available to researchers (on certain extortionate conditions) then there should be no embargo on its dissemination and a researcher at a Canadian research institution ought to have unfettered access. Of course this goes both ways. Any research produced using the public purse ought to be readily accessible by tax payers or any other researcher.

          Unfortunate that we don’t live in more reasonable times.

        • It’s not entirely about confidentiality, although that can be part of it, it’s about correctly representing the people the data is collected from in a way that is fair, accurate and culturally appropriate.

          That’s mostly likely why the forestry people want co-authorship so that if some wrong-headed comments are made about some particular group then the forestry people can clear up that ignorance, or ask for input from that group to clear it up.

          If the taxpayers are paying for the collection of the data then they should also have an expectation that the data is fairly and accurately used when describing any aspect pertaining to any of those taxpayers.

  3. The prevailing myth of science is that we are all in it together to beat back the forces of ignorance and injustice. Perhaps “we are all in it together” is the prevailing myth of every endeavor. Indeed, maybe collaboration is so rare and delicate a thing that we unthinkingly celebrate it when it occurs. A nonny mouse’s statement that “it is pretty tone deaf to ask Canadians to hand over data solely under the provisions you want when your president is threatening to annex Canada as a 51rst state. American exceptionalism and entitlement at it’s finest” seems to have some relevance.

        • Sure, let’s reflect.

          You started with a really low-brow, off-topic swipe at American exceptionalism, made up a bunch of silly reasons why someone might withhold data, threw in some pearl-clutching about indigenous tribes, and seemed to have entirely missed the fact that the discussion is about tree data and not people.

        • A reflection might be “the organisation might have genuine and legitimate reasons, some that might actually be legal, why they can not release the data in the form that I want it. Rather than disparage them in public for not releasing the data how I want it, and anyone else who doesn’t release data freely, I might take a less combative approach and try and understand what their limitations are and how we can both get what we need.”

        • > try and understand what their limitations are and how we can both get what we need

          That is exactly what the original post is about—requesting feedback about how to appropriately respond.

  4. Quote from the blog post: “The people who share data are happier.”

    Maybe some people could perform some research on that to see whether this is truly the case, and of course subsequently make the data available. Perhaps this may also lead to dozens of possibly sub-optimal conclusions that will be drawn from looking for correlations between the countless of factors and combinations. Or perhaps significance tests will not be valid anymore when countless of them are performed on the same data set by countless of people who do not know who is analyzing what in the data-set.

    Anyway, regarding the happiness and data sharing. Some researchers might invesitigate this further and provide “objective” data. They could perhaps only measure hapiness and data-sharing, but not other things that may influence and/or correlate with these things.

    Then they could make the data availabe and state things like “well this suggests, or even proves, that X, Y, Z”, look here are “the data”. Then they could also state things like “well, these are the best data we have on this stuff so it would only make sense if we follow the advise coming from this all”. Or they could state things like “well, science is a process so perhaps we’ll find something more nuanced or different in X years time, but for now we should really listen to and stick to ‘the data’ of our recent study”.

    And in the mean time they could also additionally state things like “we should really ‘collaborate’ more in science, so we propose large-scale data collection and research procedures where a small group of people decide to, for instance, study the relationship between happiness and data-sharing and only measure certain things, and not others”.

    Science! Data!

    • Quote from above: “Perhaps this may also lead to dozens of possibly sub-optimal conclusions that will be drawn from looking for correlations between the countless of factors and combinations. Or perhaps significance tests will not be valid anymore when countless of them are performed on the same data set by countless of people who do not know who is analyzing what in the data-set.”

      I need some more thought and investigation on this all, but I thought it might be useful to mention here that I am pondering proposing a new open practices badge titled “no data-dredging or -fishing”. It would be an image of a dredger-ship on the water with something that represents “data” below the water and a big red no-circle depicted over the total image.

      This possible new open practices badge would be used for researchers who 1) gathered new data specifically connected to the designed and executed research, 2) carefully conclude things, and are generally aware of what variables they did and did not measure and how this might affect reasoning and conclusions, and acknowledge and communicate this in their writing.

      I am a bit worried though that this open practice badge might be abused in the future by some third party somehow deciding they are the ones who should be handing these badges out, or making money from them in some way. Or perhaps even the mere attention towards “a badge” in itself and not thinking and talking about other things that may be way more important might not be wanted.

      So perhaps it might be good to think about some things some more…

    • Professor P. Hack closed the door
      Of his office at the university on the second floor
      Before searching what has not been found before
      By hacking and dredging until the data were sore

      Then all of a sudden on a particular day
      He was told things aren’t supposed to work that way
      His methods and tactics led things astray
      He was told Science is not some game to play

      Some people proposed some new rules
      There are now new methods, and new tools
      There are new popular people, and new schools
      But perhaps there might also be new fools

      If professor P. Hack opened his door
      Of his office at the university on the second floor
      And hung up a sign that read “data in the drawer”
      Would that truly be less poor

      If his colleagues entered the office
      And hacked and dredged like a novice
      Would it be truly less thoughtless
      Would it truly solve the madness

      Perhaps the open data sign
      Results in thinking what yours is now mine
      But is there really such a line
      And if there is one, is it incredibly fine

      Maybe the data, hacking, and dredging are essentially still the same
      Even with these new rules of how to play the game
      Sure, it may be more indirect and there is no one person to blame
      But maybe some things have just been given a different name

      • The p-hacking is really irrelevant, anything goes when developing a model. The problem is there is no expectation these models will ever be tested on new data. The keystones of science are replication and surprising, yet accurate predictions. Without that its just people deciding whatever “makes sense” in the best case. And there’s no quality control for the worse behaviors either.

        • Quote from above: “The p-hacking is really irrelevant, anything goes when developing a model. The problem is there is no expectation these models will ever be tested on new data.”

          If I am not mistaken, I can clearly remember reading about how a researcher might look and search some data and then forget about it and later on analyze the data again or something like that as an explanation (or excuse?) for how p-hacking or data-dredging occured. And then preregistration was mentioned as a way to help solve this issue.

          But it seems to me that this described scenario is all of a sudden not a possible problematic scenario when people talk about open data, and how this might possibly be used (or abused?). It seems to me that a researcher might go through some open data set and then forget about it and later on analyze the data again (with or without preregistration).

          Or what about multiple researchers going through a data set and only those that find something significant end up reporting on it. Is this all similar to, or essentially the same as, p-hacking or data-dredging or whatever term is most appropriate to use? And what about all the “exploratory” findings that can be found via these open data sets, which might then still be reported and viewed and used like “exploratory findings” in the past only now they might seem more valid because of the large number or participants or such things. Perhaps this latter point ties in with your comment about models and testing on new data.

          It seems all very similar to me, which seems like a way in which certain problematic issues might also occur with “the new rules of the game” so to say. That’s part of what I was trying to make clear, or ponder and wonder about.

        • I don’t care if they p-hack, its like reading tea leaves, interpreting dreams, or taking mushrooms to come up with ideas.

          It is supposed to be a hypothesis generating step. The error is treating the first step like its also a validation. It simply is not. The validation is independent replication (for methods) and predictive skill (for models/hypotheses/explanations/theories)

  5. I think Statistics Sweden is the gold standard here.

    A (whole) lot of data is available directly from the website, no questions asked, and provided in a resolution that preserves privacy (e.g. wage data is binned, I believe).

    All data is available on request. For a small fee, intended to cover labour costs I believe, an administrator at Statistics Sweden puts together the file containing what you need. I think you must promise to follow certain security protocols.

  6. LIZZIE
    I want them to share the data without requiring co-authorship or other signed agreements. 

    I think a nonny mouse has correctly identified the issue.

    Given everything from indigenous land claims to the seemingly perpetual softwood lumber dispute with the USA, on-going battles on Vancouver Island about old growth forests —few years ago  a friend was chaining herself to a tree, IIRC— not to mention Trump’s more recent tariffs, you are running into  some nasty political issues. 

    One assumes that senior management at the Canadian Forest Service and on up to the  Deputy Minister and Minister of Natural Resources Canada, if not the Prime Minister, is antsy as all–get–out about losing control of potentially sensitive data even to researchers in Canada. If I understand your lab’s policies, you would make the CFS data generally available if you publish. I suspect that this is anathema to the Canadian Government at the moment.

    I would guess that you have almost no hope of gaining access to the data without the “co-authorship or other signed agreements”.  I doubt if the Canadian Government would want to have a data release to your lab covered under the Security of Information Act but given the level of paranoia that probably prevails in Ottawa at the moment,  who knows.

    With the current state of affairs in the USA, I side with Canadian Forest Service. Heaven only knows what the equivalent of RFK Jr at whatever is the US version of Natural Resources Canada is would do with totally innocuous data.

    • Not sharing to avoid getting “scooped” is one thing. It’s unfortunate but how our society works.

      Not sharing because it will be “misused” just means you have no confidence that your analysis is better than some other one.

      Indeed, 90+% are just one arbitrary model out of effectively infinite equally plausible choices.

      Solution: Predictive skill. If RFK Jr’s model has worse predictive skill on future data, shouldn’t be hard to convince people yours is better.

      • Anon:

        To say that RFK Jr. has a “model” is absolutely ridiculous. He’s a politician. He might have a policy. He does not have a model in the sense that it would have “predictive skill.” Your argument is like saying, oh, I dunno, maybe that there’s no need to lock the door because if a burglar comes you could reason with him not to steal your stuff.

        • RFK Jr is just a placeholder for whoever it is thought will “misuse” the data.

          Perhaps I don’t understand the threat. What exactly can someone do besides analyze it differently and come to a different conclusion?

      • No, you do not seem to grasp the idea that release of this data is not a matter of scientific communication but potentially serious international dispute.

        Lizzie’s research and the data she requests may be innocuous and to our national good but no Canadian Government is likely to want to release data that the US regime might torture into policy.

      • No . I do not think you understand. We, in Canada, consider RFK jr.is a “”%$$89 idiot, a war criminal, and a danger to the continent. I assume Mexico feels the same way.

        • Its a doth protest too much situation. Lizzy can’t have the data because a hypothetical tree-ring version of RFK Jr might also get to reanalyze it and come to a different conclusion.

          I looked into tree ring data like 5 years ago, it was pretty dubious stuff involving choosing “good” trees and lots of averaging. Almost more assumptions than data. I suspect the real reason is the dataset is just too noisy on its own and questionable assumptions are filling the gaps.

        • Anoneuoid –

          You say this…

          Lizzy can’t have the data because a hypothetical tree-ring version of RFK Jr might also get to reanalyze it and come to a different conclusion.

          … as if it’s the only possible explanation. I’m not going to defend not sharing data, you speak to a legitimate issue – but the fact that you find can ses only one explanation is a effectively a fallacious argument from personal incredulity. Repeated, it becomes a confirmation of a bad faith bias.

        • I think people tend to be older here so don’t really “get” internet posting, which is best when kept as concise as possible. Ie, no one* completely reads the long rants.

          A week or so ago I had an issue that someone though I was hiding something due to only quoting the relevant part of a link that I shared. Including a direct quote and link to a page with a few paragraphs is not generally considered a way to hide something, since it requires trivial amount of effort to check the context at the source. Further there are multiple readers, one of whom is likely to click and scan the contents at the link.

          Likewise, in the interest of conciseness, the reader is expected to understand that the privacy, etc concerns were left unmentioned because it would clutter the comment to include them. I.e., I considered including a line like “they are better off sticking to the privacy angle”, but left it out because it is redundant information in the context of the thread.

          * The use of “no one” is also shorthand for “It is rare for someone to” and is not meant to be taken literally.

      • When I am in some minor despair
        Reading about some data to, or not to, share
        That eventually, perhaps inevitably, involves politics
        About which I give no two licks
        I can always remind myself
        To reach for my tools on the shelf
        So I can turn a “three-ring circus”
        Into a “tree-ring circus” on purpose
        Cutting through it all like a chainsaw
        To reach the core like an oil platform offshore
        Not to get to the oil, and not to count the rings
        But to go-with-the-flow, and to cut the strings
        Some might say it might turn a frown
        Upside down
        Just like a clown, that’s just messing around
        In this particular circus, in this particular town
        But that’s not what seems to me
        To be what it is, or what it might be
        Perhaps it’s something I just can’t see
        Perhaps it’s something I just need to do to feel free

    • While I understand the rationale you describe, and am willing to believe that it describes the mindset of the CFS, I do not believe it is a good reason to restrict access to data (modulo the privacy issues which, as Lizzie notes, can be addressed in various other ways). I do not disagree that, particularly under the current US regime, such data could easily be misused and misrepresented. Rather, the current regime is going to do whatever it wants anyway. The availability of data might alter the surface features of whatever justification they deign to provide, but it will be equally vacuous regardless. So I do not believe that any harms are prevented by restricting access to data; it only harms good science and the potentially good policy that could come from it.

      In that sense, the CFS policy can be seen as a victory for the current US regime, given that their goal is to restrict public access to data and knowledge.

  7. I think the co-authorship requirement is misguided, but I do think that it’s reasonable _and good practice_ to talk with the collectors of the original data to understand how to use it correctly and all the idiosyncrasies. I have talked to people about data sets I have on ICPSR, and I’m glad to so that they don’t end up creating confusion.
    People should actually like it when other people use their data since citations to data sets are citations. And of course you put them in the acknowledgements.

    If there really is individual tree level data that they can’t share for privacy or safety reasons (and I could imagine scenarios where someone targets a tree or when a private landowner would worry about activists learning something about their trees) then there do need to be agreements. This is like getting the individual transcript data for NELS.

  8. Call their bluff.

    “Thank you for your response. We are not interested in pursuing a collaborative scholarly endeavor with you. Consequently, please provide the data that is not under restricted use licensing, is published, and is available without authorization from the original data owner. In addition, we are willing to invest the time required to obtain authorization from original data owners, so please provide the relevant contact information. Thank you for being open to advancing science by providing data that is not specifically restricted.”

    • Uh, it is not a bluff.

      If that data was gathered by the CFS, it does not matter who the original sources were, it is the property of the Canadian Government. If I remember this correctly, that means it is the property of his majesty King Charles III in right of Canada. Canadian law is just a bit different from US law.

      please provide the data that is not under restricted use licensing,

      It is quite possible that most or all of the data is “protected” if not more highly classified. Canadian Gov’t people tend to do this reflexively.

      Most of that request is either something that would be laughed at or would take a few years under the Access to Information Act to process. The Access to Information Act is not terrible good if you need information. One of our many failings.

        • No, they mean what they say. As Lizzie says, some data is publicly available and I imagine CPS people are happy to provide it. From my own experience, the researchers likely would be cracking open champagne bottles on finding that someone actually wanted their data.

          It’s a matter of getting clearance to release the other data. A breech of the Security of Information Act can, potentially, be a 25 year prison sentence.

          Besides, as I mentioned in an earlier post, we do not trust the US Gov”t to not distort any data. At the moment, Canadian civil servants are probably reluctant to release their telephone numbers.

  9. I’m a data sharer because when the USA federal government funded science mine often came from NASA, which has insisted on it for a while, and my PD’s checked… the only problems data sharing has ever caused me involved uploading, cross-checking, other QA/QC, etc. that I had to do w/in 3 months.

    But I’m hung up on the “data sharing makes you happy – data shielding makes you sad” idea. Maybe the causality is reversed, and happy people share – bitter ones shield. Of course, there could well be a positive feedback loop or or two in play…

    Whatever it is, I am SURE that procrastinating people use ellipses. Gotta go write a quiz before dinner

  10. How about a tool where the user inputs a paper, then all papers citing that one are listed.

    Next, for each figure in the original paper all the figures in citing papers are checked for replications in the general (indirect + direct) sense. Then the methods are checked for variations to further label the direct replications. The author lists could also be compared for “independence”: some function of shared authors, institutions, etc.

    There would be an uncurated “auto” (llm) mode but also allow user submissions. Then (if there was enough interest) crowdsourced curation in a wiki style can be applied.

    It sounds like a dream. It would be so useful for theorists to filter out all the unvalidated observations, then focus their models on fitting what remains.

    The main problem with p-hacking, hiding your data, and such is the noise it generates. The speculations are mixed in with the validated observations. This way people can keep publishing that stuff (which they obviously like to do), without impeding those who require higher standards for their work.

    • That’s exactly the data! Thanks for the paper; that’s helpful. Can you specify where they say they are sharing it?

      (I think that they think they are sharing it by the fact that you can have access to it if you agree to all their requirements.)

      • Hi Lizzie,
        I found some quotes on page 2 of the paper I linked:

        ” Tree-ring research has benefitted from international open
        access research databases for decades, with data archived in the
        International Tree Ring Data Bank (ITRDB; Grissino-Mayer and
        Fritts 1997; Ols et al. 2018; Zhao et al. 2019).”

        It is a moral failure to exploit the data-sharing of others while refusing requests for your own data.

        Also from page 2:

        “To answer global questions, data can be pooled for greater coverage across large-scale networks. For example, several recent studies relying on broad-scale tree-ring data networks have described patterns of forest growth response to global change. These studies include subcontinental- to hemispheric-scale analysis of climate tree growth relationships […].”

        Here again they are extolling the value that can be gained from international data sharing while refusing to share themselves.

        I also have a general comment about this thread. It’s pretty clear from all the vague statements in support of the CFS policy that this community is not really on board with full data sharing. These comments all sum up to “bad people might do bad things,” which is ridiculous in the case of tree ring widths. If you can sufficiently anonymize private information about humans and still have useful data, you damn sure can do it for trees!

  11. I am wondering if anyone posting about Indigenous land and Indigenous data sovereignty could clarify exactly what they think is the connection? I see two options these people could be posting about: (1) The Canadian Forest Service has access to important data on tree ring width (let’s make it sound more exciting and call it tree growth) that Indigenous people would like and this policy prevents Indigenous people from getting said data or (2) Indigenous people have asked the CFS to restrict data access so that researchers working on climate change cannot access data on tree growth from Indigenous lands and somehow the way to access it is to give CFS people co-authorship. What am I missing here?

    • I can’t speak for anyone else, but First Nations, Inuit, and Metis in Canada have a long history of being moved off land at gunpoint the day after the government or a big business found something to use that land for (or being moved at gunpoint because someone in Ottawa looked at numbers and thought it would be cheaper). And the Crown does not have much legal basis for authority over them, there were no treaties in much of the west of Canada. So many First Nations, Inuit, and Metis are suspicious of proposals to collect data about them and their land, because historically that data was used by outsiders to hurt them.

      • Hey, to be fair, I think we usually only used bayonets. /s

        I think many commentators here are not Canadians and do not realize the complicated political situation that can surround something as seemingly innocuous as tree ring growth. Personally’ as an Easterner (Ontario), I don’t understand the forestry issues all that well but I do see the CFS’s point.

        One really does not want questions in parliament about data which one has lost control of. It really does not matter if it is DFO in Newfoundland or CFS/NRC in BC.

        • This “the natives will get their feelings hurt” sounds like a BS excuse to me. The vitriolic anti-Americanism displayed by Canadians on this blog explains the lack of data sharing all by itself.

          The US should ban all data sharing from tax-payer funded research with Canadians.

        • Anonymous
          I’m not sure why I bother responding to you, but can you seriously complain about Canadians showing anti-American sentiments? Trump has threatened to take them over (somehow) and enacted serious economic harms on them (despite prior negotiated agreements), yet you somehow think they should feel kindly towards us. There’s a role in the Administration for you – just wait for the next firing and there will be an opening.

        • Dale, I never voted for Trump, but Trump trolling Canada is not the excuse for the vitrol displayed by Canadian academics towards American academics you think it is.

        • The message I get from Canadians on this blog is they want to play hardball. OK, let’s play hardball. Cut them off from access to US academics, research, opportunities, and data.

        • To all: Chill out on the U.S./Canada war! If any of you don’t want to collaborate with researchers across the border, that’s your call. But not sharing data is a negative-sum game. We’re talking about research that should benefit everyone, not weapons technology or other national security topics.

          To put it another way: I’ll share my data with anyone under no conditions (assuming there are no legal restrictions). Even if researchers won’t share their data with me, I’ll still share my data with them. If I could get organized enough, I’d just post all my data, so I’m not sharing it “with” anyone in particular; I’m sharing it with the world. I’m not always organized, so I share data with anyone who asks. I guess that if I was overwhelmed with requests, I might need a different policy, but as it is, I share with everyone who asks. I’d even share data with Mark Hauser, the disgraced primatologist who famously wouldn’t even share his data with his own collaborators. I’d even share data with Dan Ariely, the unfortunate psychologist who keeps coauthoring papers based on data that seem not to exist. I’d even share data with Mary Rosh, who reported on an entire survey for which there’s no clear evidence that it actually exists. Etc.

          But, sure, if any of you–Canadians, Americans, whatever–don’t want to share your data, that’s your call. As the saying goes, you’re under no obligation to share your data and I’m under no obligation to believe whatever claims you make based on secret data.

        • Actually the “the vitriolic anti-Americanism” on display was by me, and I am not Canadian, nor even North American.

          As far as I can see the Canadians replying have been incredibly polite.

          I come from a public health background where it’s important (but not always possible) that groups of people should have the opportunity to analyse and/or interpret information that is collected about themselves as collaborators. Otherwise, all that happens is that their data is interpreted through the eyes of white upper-middle class, middle-aged people who bring their own opinions and prejudices to their interpretations. It’s a lot more than “not hurting their feelings”, it’s about reporting fairly, accurately and in context, with a view to improving people’s lives, not diminishing them.

      • I can’t speak to Canada, but what Sean says reminds me of disputes over data privacy in the US. Representatives of Native American tribes have been pretty vocal in debates over how the US Census Bureau preserves data privacy, presumably because they feel like their ability to function autonomously has been f’ed with enough and don’t have a lot of patience with further mandates they perceive as threatening their data. If there are similar tensions over autonomy in Canada that may explain part of it.

        • It’s sort of a different scenario in the US, because the data are population estimates. No one but the Census Bureau has access to raw measurements, they only release noised versions. But tribal members have spoken up against the application of methods like differential privacy out of concern that the government is further disenfranchising their ability to govern independently. I don’t know if there are parallels in Canada, but if the data are documenting reserve land, maybe there are similar tensions about who should own that.

        • Lizzie wrote:

          “Did the US Census Bureau ask for co-authorship”

          Ya this thread gets weirder by the minute.

          Jessica, I’m hoping you will be willing to help me out since no one else has tried to connect the dots. Five commenters have now raised “what if?” questions about the privacy of indigenous tribes, what has that got to do with tree ring widths?

          The annoying part of this for me is that the battle over tree ring data has quite a sordid history going back nearly 30 years. Back then, authors had really good reasons to not want outsiders to see their data, especially the raw squiggles that look nothing like smooth hockey sticks. Now the climate wars are way behind us (a total rout for the consensus side, if you are keeping score) and most of that data was eventually archived and is now available through one of the open databases. Back in the day, the excuses for withholding data – tribes, private land, contractual obligations, intellectual property, etc. – often looked a lot like the ones the CFS is currently using, and the coauthorship requirement was part of that as well. Just like back then, you would see the supposed requirements quietly disappear if you let yourself be extorted into taking on a coauthor.

        • Matt, rather than speculate, I suggest 1) reading the excerpts from the email from CFS, 2) reading their publications, and 3) talking to them. CFS have given three general examples of good reasons not to share data, but everything depends on the specifics (eg. the difference between “this data set is unpublished because Communications has not finished the press packet” and “this data set is unpublished because of the three people who understand what will have to be done to make it useful, two are retired and the other is on maternity leave”). I repeat, with emphasis, that in the quoted email CFS does not refuse to share data, but says that some data can only be shared with third parties if they have been vetted, and some data can only be shared with third parties if the owners give permission.

        • Sean wrote:

          “CFS have given three general examples of good reasons not to share data…”

          Nonsense. There is no such thing as a “general example of a good reason not to share data.” Let’s review what the e-mail said.

          1. Disclosing precise locations could compromise data integrity or infringe on landowner privacy.

          Still looking for some brave soul to connect the dots from tree ring widths to privacy. Anyone?

          2. Some datasets also contain sensitive ecological or proprietary information.

          Since we all have the same goals here, the CFS could just include a note that “this parcel contains rare carnivorous plants that have been subject to exploitation. Location data for these populations are withheld from the public by Canadian law.” That way Lizzie knows to scrub out any references to those plants in her paper, which is all irrelevant to the tree ring widths anyway. How would this situation change if a CFS coworker was a coauthor, since Lizzie’s nefarious coworkers could still go out and harvest those valuable plants?

          3. NFI data are designed to provide an unbiased representation of forest resources. Controlled access helps prevent misuse or misinterpretation…”

          Now suddenly we have replaced both peer review and the entire scientific method – both of which are structured to “prevent misuse or misinterpretation” – with “controlled access.” “Misuse” sounds suspiciously like “used to make a claim we don’t endorse.”

          But the most telling thing in the e-mail was this:

          “Entering into collaboration would help streamline access to tree-ring data across Canada by removing certain barriers.”

          This is where they are telling you that if you accept the coauthorship requirement, all the stuff they spitballed goes POOF!

  12. I do not think they are refusing to share data without collaboration, but they are saying that some specific data is licensed for internal use only, and some specific data is unpublished. Those sound like two good reasons for not sharing data: either it is not yours to share, or it hasn’t been put in a state that a newcomer can do sensible things with it. If you send me a draft paper on the condition that I do not share it, and someone asks me for a copy, that someone should not be upset at me when I say “you will have to ask Lizzie for permission” and lecture me about abstract principles.

  13. Just to add to this discussion, let me say that analysis of tree rings is really hard! Matt Schofield, Richard Barker, Ed Cook, Keith Briffa, and I published a paper a few years ago which was pretty much all about how difficult it is to draw conclusions about historical climate from tree-ring data. The problems are as follows:
    – Not a lot of trees in the dataset.
    – Lots of computations so anything but the simplest models take awhile to run.
    – The transfer function relating historical climate to tree-ring data varies over time, over space, it even varies by tree.
    – It would be great to include more trees, but then the model would have to get more complicated to handle the additional variation.
    The whole thing’s a goddam mess.

  14. Are the comments here related to this:

    The Canadian province of Nova Scotia is facing pushback for what some have called “draconian” restrictions as it tries to limit wildfire risk in extremely dry conditions.

    Last week, Nova Scotia banned all hiking, fishing and use of vehicles like ATVs in wooded areas, with rule breakers facing a C$25,000 ($18,000) fine. A tip line has been set up to report violations.

    The Canadian Constitution Foundation, a non-profit that defends charter rights in the country, called the ban a “dangerous example of ‘safetyism’ and creeping authoritarianism”.

    Tens of thousands of residents are under evacuation alerts in eastern Canada as the country experiences its second worst wildfire season on record.

    Nova Scotia Premier Tim Houston says human activity is responsible for almost all wildfires in the Atlantic province – official statistics from 2009 say 97% of such blazes are caused by people.

    https://www.bbc.com/news/articles/cn8533np061o

    It’s interesting to think about if there are were humans starting fires, and no powerlines, etc. In that case, the only known cause is lightning. So the logic goes: “no known lightning, no powerlines, it was a hiker”.

    In general think about all the fires humans start, including the controlled ones. Every lighter, running any kind of hydrocarbon powered device (eg, your car runs on a series of small combustions).

Leave a Reply

Your email address will not be published. Required fields are marked *