Help teaching short-course that has a healthy dose of data simulation

catsinNH

This post is by Lizzie. I hope you like the cats photo from this summer. I do.

I am looking for help. I decided to change my term course (12-14 weeks-long) on `introduction to Bayesian modeling with some hierarchical modeling’ (no, that’s the not the official title, but that is the gist) to a three-week intensive. I have been thinking about doing this for a couple years, but finally decided to do it now. My feelings in teaching the term-length versus short course is that: (1) a lot of students discover during term that Bayesian approaches are a lot more work than what they can get away with in frequentist methods and lose interest, but are still stuck in class for weeks — this way, when they can find this out, the course will soon be over! (2) As a corollary to (1), taking a Bayesian course (to me) mainly means finding out if you want to dive in and do it more (as you will never learn enough in a term course to be off and running fully), so with a short course, more students will take it and find out if they want to learn more. (3) More students will take a short-course and provide more support for those who want to continue on (building a useful community). (4) More students will take the course and learn how to simulate data, and I want more people to learn this. There’s lots of other reasons, but those are my big ones.

The downsides are that students will likely arrive unprepared and I won’t have time to go through the basics the way I do in 12+ weeks and that generally, there will be less content. Also, no more analyzing your own data as a term project that I support. People will miss that.

The students will be mainly MSc and PhD students in ecology or evolution (but not in molecular evolution, think more: people studying populations of birds and how their phenotypes might evolve) who will hopefully not be taking many other courses. Most will work in R and have some background in statistics, but not have been too challenged to fully understand the statistics they’re using.

I have two days of 3.5 hours each per week for three weeks (about 21 hours) and plan to cover:
Week 1: Data simulation for linear regression; what are priors and some ways to check them
Week 2: Fitting a model in rstanarm to simulated data; diagnostics
Week 3: Introduction to hierarchical models and posterior predictive checks

What I could use help on:
– Suggested classroom activities and problem sets
– Good example datasets or vignettes and generally any other good resource that could help me teach this material
– Advice on how to structure or approach a short course to make it work well
– Recommended background info to point students to

I have some ideas of examples I might use, like simulating data to show how much bigger a sample size you need to estimate an interaction in week 1 perhaps, and the Olympic figure skater example (judges, skaters from Regression and Other Stories) as homework for week 3, but I am really not sure so all advice and ideas would be most welcome.*. I am especially hoping to emphasize data simulation throughout, so ideas there are extra welcome. Please let me know in the comments any ideas/thoughts you have (or you can email me if you much prefer). Thanks in advance.

* One thing that I am NOT looking for is a big debate on the values of trying to teach all the way to hierarchical when people may not really have the basics.

What is spatial epidemiology, anyway?

Every time I talk or teach about spatial epidemiology, I find myself confronted with the difficulty of defining what it is. More specifically, I have a hard time defining what my version of it is, why I do research in this area, teach about it, and just think about it a lot of the time. I also worry that students who came for maps, GIS, fancy statistical models, and all that good stuff will be a bit disappointed when they get my version, which has some of that but is also more eclectic and navel-gazey.

Sometimes, I think about changing the name of the class to something like “relational epidemiology”, “spatial and contextual epidemiology” or just health geography – but I’m not a geographer and I’m not totally sure what health geography is, either.

At the end of the day, spatial epidemiology is interesting and important to me because it is relational in nature. Maybe this just reflects the way my brain has been poisoned by training in the social sciences and infectious disease dynamics, which are all about relationships and interpersonal dependence. But if we called it relational epidemiology, what would the most important relationships be?

  • Relationships between individuals, e.g. in a classic social network.
  • Relationships between people and the environment, i.e. climate change and other types of human-driven ecological change.
  • Relationships between areas of the physical environment, e.g. dispersal of dust and other pollutants through the air, movement of bacterial and viral pathogens via water sources.
  • Hierarchical relationships between social units, i.e. neighborhoods within cities.
  • Within-individual change over time, e.g. the progression of chronic illness, natural aging.

This defines the problem space for what I think of as being the super-group of “relational epidemiology” topics. Then we have a set of ideas or approaches that act as useful frames through which to view these ideas: Spatial analysis clearly falls under this heading, but so do network analysis, time-series analyses, non-spatial hierarchical models, individual-based models, and on and on. These also touch on other well-established fields like ecology, social epidemiology, environmental health, sociology, economics, political science, and on and on.

Making Choices

One of the early lectures in my online spatial epidemiology course is titled “Making maps means making choices”. I like this one because it gives me the opportunity to feel smart by reiterating a point that has been made many times before: Spatial approaches to public health are powerful because they are decidedly non-neutral. Maps have the pleasing appearance of something settled and clear, but we know they obscure more than they show. A disease map includes the information on risk and relationships we want to highlight. The stuff that is left out is implicitly understood to be less important than what is left in. This makes it just like any other model, statistical, mathematical or otherwise.

I guess this is why I keep calling the class spatial epidemiology rather than something more expansive that could allay some of the mildly guilty feeling I get about teaching a version of this class that is heavier on ‘spatial thinking’ (whatever that is) than ‘spatial methods’. (Honestly I’m not even sure what exactly belongs in that set or doesn’t – but that’s for another day).

When I say it’s a course about spatial epidemiology, to me that ultimately means that space is the starting point rather than the destination. In other words, if we put things on a map or estimate a model of the distances between individuals with different attributes or outcomes, then we have to ask why the patterns we see are the way they are. We get to tangle with all the wooly questions about relationships and interdependence, but we start from a place that most people grok on at least some level.

This can be done as effectively through other lenses: social network analysis, ethnography, agent-based modeling and others. But to me the reason space is particularly powerful for building a relational perspective in epidemiology and public health is that you can put anything on a map: Everything that is within the concern of public health can be pinned down to some location on a map. Whether or not that location is meaningful is another question, but at least it gives us some place to start.

Ok, so what?

I don’t know – up here in Michigan we’re on spring break (it’s above freezing!), and I’m taking a few minutes to think about why I do the things that I do. But more than that, spatial analysis feels like one more slippery set of tools or concepts among the ones I care about. Asking why I care about spatial epidemiology is not that different from asking why I think Bayesian statistics, transmission models, hierarchical analysis, and many other things that sound kind of well-defined but aren’t are good and important things other people should care about.

Teaching about these things, but also publishing on them and writing grants to get people to pay for the work, forces us to articulate what they are all about. But it might be helpful sometimes to zoom out and admit to ourselves and everyone else that these are all fuzzy concepts, more like a question we have to continually ask and answer rather than one that has a fixed meaning.

And maybe you already knew that – but I wrote this to remind myself for the next time I forget.

(Thanks to Krzysztof Sakrejda and Joey Dickens for ideas & feedback! h/t to Justin Lessler et al. for their great paper “What is a hotspot anyway?” that got me thinking about this.)