Edward Kennedy on the Facebook/Instagram 2020 election experiments

The first batch of papers from the Facebook/Instagram 2020 election studies were published about a year ago. I thought it might be interesting to give people a bit of a view into the role of statisticians and methodologists in this kind of project.

This project had several methodologists involved. I reached out to Edward H. Kennedy, Associate Professor of Statistics and Data Science at CMU, in part because he’s the most deeply embedded in the world of academic statistics. Edward has made extensive contributions to the causal inference literature and has also engaged in empirical collaborations in multiple fields, including criminology and biomedicine. Edward coauthored the papers based on randomized experiments that assigned users to chronological feeds, removed reshared content from their feeds, or downranked content shared by “likeminded” others. Below is our exchange via email from August & September 2023.

DE: How did you get involved in this collaboration?

Edward Kennedy: It was pretty straightforward and lucky on my part – in the midst of the pandemic (August 2020), Drew Dimmery emailed me and asked if I’d be interested in working with him, Facebook, and some political scientists, on a project where heterogeneous treatment effects could play an important role. I knew of Drew and the political scientists involved, and I was really excited to apply methods I’d recently been studying theoretical guarantees for, “in the wild”. I remember our first call about the project pretty vividly, which I took while walking around nearby Frick Park with my then 1.5 year old – his daycare was closed that year so we took many walks in the park during that period.

DE: There are a lot of connections (sometimes neglected) between problems in survey sampling and in causal inference. Here these come together because the main estimands are all average treatment effects for a broader population, and the experiment has a biased sample of people from that population. Can you say a bit about how you thought about the choices to (a) designate those population ATEs as the main quantities of interest and (b) then actually choose estimators for those quantities?

EK: My sense is the population ATEs were clearly of more interest in this particular case. However the question of whether to target sample versus population effects is a really interesting one — I actually have a paper with Siva Balkrishan and Larry Wasserman coming out soon about minimax optimal estimation of sample effects. On the one hand, in-sample effects can sometimes be estimated more accurately and under weaker conditions (e.g., allowing covariate dependence / non-random sampling), whereas on the other hand, population effects may be of more primary substantive interest, albeit requiring stronger assumptions to identify.

DE: There has been substantial interest in methods for detecting and estimating heterogeneous treatment effects, and you’ve contributed to this area. These papers involved using systematic methods for looking for these effects, with limited evidence of heterogeneity in effects on key outcomes. How do you think about these results? How should researchers designing a study consider interest in HTEs at the design stage?

EK: Great question – I think a lot of design issues are understudied and could benefit from more research. Perhaps the most obvious answer is that if conditional effects are of specific interest, then one may want to over-sample more subjects of certain types than would be expected from a simple random sample. But in modern continuous/high-dimensional settings I think more work could be very useful here.

DE: With new statistical methods, sometimes it can take a while for them to be accepted and intuitive to empirical researchers in an area. Maybe modern approaches to HTEs are an example. A fun moment in the peer review file for the paper on downranking “like-minded” sources is when one reviewer writes, “I know that post-hoc subgroup analyses are passé, but I was really hoping for a heterogenous treatment effect surrounding age.” How do you think about the role of statistics and statisticians here in shifting (or just being responsive to) what quantities applied researchers want?

EK: Participating in the give-and-take between statisticians and substantive scientific researchers is one of the most fun parts of being a statistician. It’s really exciting to try and decode what precisely the scientific question is in substantive scientific work, and translate that to a formal statistical problem, which then very often can lead to exciting new theory & methods. I’m a big fan of taking scientific questions at face value, and not trying to force them to align with, for example, a coefficient in a potentially misspecified parametric regression model.

DE: Related to both having many subgroups and many outcomes, these papers use methods for controlling false discovery rates. Here on this blog, we might think about taking other approaches that would jointly model effects on these outcomes, thereby borrowing information across them, and perhaps reducing the need for post-estimation adjustment of p-values and intervals. I wonder if you have any thoughts about the approach to multiple-testing taken in these papers (and that seems to be becoming more common in social sciences) and alternatives.

EK: I think both types of approaches can be useful, depending on the context. The hierarchical Bayesian approach can sometimes rely on quite strong modeling assumptions, which one might want to try to avoid in some settings — for example, if the assumptions are not accurate, then potentially severe bias could arise.

DE: Other fields, like life sciences, have more of a tradition of having dedicated statistician authors, but this is less common in the social sciences. You’ve been involved in both empirical work in the life sciences and social sciences (including not just this project, but also work in criminology for example). How does the role of the statistician compare? Are we going to see more of this kind of division of labor in the social sciences?

EK: All of my collaborative projects have been pretty different/unique. In these [Facebook] studies, I served in more of a consulting/advising role, pre-analysis, whereas in other work I have been more explicitly & intimately involved with data analysis. I do think more generally that there is a major gap in the social sciences for statisticians to fill, and that more statistician involvement could be very useful for everyone involved.

DE: Your point that here you were more in a consulting role, without hands-on data analysis, is perhaps relevant in the context of Michael Wagner’s reflections on observing this collaboration as independent rapporteur. To summarize, he sees this as an independent, rigorous project, but also not a model for future research. This is in part because of limits on external researchers’ access to the data, but also more generally their need to rely on internal partners to even figure out what might be possible. To what degree to see this collaboration as a model for future research on/with tech giants? Are there other models — perhaps from other fields — that might work instead?

EK: I’m honestly not sure whether this style of collaboration will be replicated. It’s one of the things I’m most curious to follow going forward – in some ways the collaboration was very unique, and may not be feasible for other research groups, but on the other hand perhaps it could act as a model in some cases. For what it’s worth, though, it worked really well from my perspective – it seemed clear to me that everyone involved was very invested and wanted to work together to do the best job possible. I second Drew Dimmery who said: “Everyone wanted to get this right!” [DE: link]. This motivation/attitude came through in all my discussions with folks both in academia and at Meta.

DE: In the studies we’ve seen so far, there are a lot of outcomes for which we can’t reject the null of no effect. Now maybe this is fine because there is useful evidence against large effects (as quantified by standard CIs or by equivalence tests). My own guesses of likely effects for many of these outcomes weren’t zero, but were also small enough that these studies had low power to detect them. Maybe I had unusual pre-results beliefs. How did you or do you now think about power and precision in these studies? Is there a role for statisticians to help in eliciting and summarizing experts’ prior beliefs for use in research design?

EK: I didn’t have a great sense a priori of power and precision for these particular effects; my contribution was mainly focused on providing guidance in implementing flexible but robust statistical methods for heterogeneous effect estimation, which come with strong guarantees on MSE and inference, for example. But surely there are ways to do more to incorporate prior beliefs, or other structure, to get more precise results. I really enjoyed reading your post on Gelman’s blog detailing your predictions; it would be fun to think about how to use this kind of information at the design stage. There are for sure some interesting problems to work on there.


Thanks to Edward for this exchange. We delayed this post a bit thinking we might link to the as-yet-unreleased paper he referred to above and then I forgot about it for a bit. But that’s OK, this blog is no stranger to posting on delay!

One final bit of commentary from me. Some of my questions were about choices at the design stage, and I think Edward’s answers are consistent with the idea that perhaps the statistics literature has neglected design, compared with analysis. This made me think about systematic reasons we could have a shortage of work on design (rather than analysis) of experiments. For example, do applied stat papers on design have a harder time because they can’t as easily show better performance in real data the way some new estimator or predictive model could?

[This post is by Dean Eckles. Because this post is about a collaboration with Meta, I want to note that I have previously worked for Facebook and Twitter, received funding for research on COVID-19 and misinformation from Facebook/Meta, and coauthored papers with Facebook/Meta researchers. See my full disclosures here.]

6 thoughts on “Edward Kennedy on the Facebook/Instagram 2020 election experiments

  1. Dean, Edward,

    Thanks for the discussion. I appreciate the long blog delay!

    Also I wanted to pick up on this statement above from Edward:

    The hierarchical Bayesian approach can sometimes rely on quite strong modeling assumptions, which one might want to try to avoid in some settings — for example, if the assumptions are not accurate, then potentially severe bias could arise.

    There are a bunch of interesting things to chew on here . . . really, it’s worth its own post. Maybe we can discuss further. Here are a few issues:

    1. Most obviously, all methods use assumptions, so the statement, “if the assumptions are not accurate, then potentially severe bias could arise,” is generic and applies to all approaches, not just hierarchical (or non-hierarchical Bayes). Speaking generally, assumptions can be used to derive a method (as with Bayesian inference) or can be thought of as inducing bounds on the applicability of a method. One way to see this is that you can take any Bayesian inference, look at the resulting computational procedure and its outcome, and throw away all the assumptions. Now you have an “estimator” whose frequency properties can be evaluated without reference to its Bayesian assumptions. From the other direction, just about any non-Bayesian estimator or procedure can be retconned, at some level of approximation, as a Bayesian inference, and you can look at its implicit assumptions. From either direction, you can find examples where inferences are more or less robust to the model assumptions, and often the level of robustness depends not just on the model and the method but also on the specific data.

    2. Getting to the nitty gritty of hierarchical Bayes and related methods, I’m genuinely unclear about how to think about what are “strong assumptions.” For example, if you do a lot of partial pooling, that could be seen as a strong prior or a strong assumption compared to a no-pooling model; on the other hand, no-pooling is equivalent to a super-strong prior on the group-level variance parameter, and if you think about it that way, hierarchical Bayes uses much weaker assumptions, as it lets the data estimate the amount of pooling, within reason. Aleks and I discussed some of this in very general terms in our 2007 paper, Bayes: radical, liberal, or conservative?

    I think there’s more to be done here in the context of specific models. One direction that could be useful is to look at summaries such as effective number of parameters (see for example this paper and chapter 7 of BDA3). I don’t claim to have anything approaching a complete understanding here.

    • As mentioned by Drew below, part of the context here is a large and heterogeneous collaboration. I wonder if this leads to use of methods that are more likely seen as objective and conservative (whether those labels are helpful or not). I could also see the other direction: If you can really get 30 experts to sign onto using particular periods for an analysis, then maybe those priors do indeed reflect a broad, unobjectionable consensus about the problem.

      • Dean:

        I get that there can be “sociological” reasons why a certain approach might be used, perhaps because it is traditional or sometimes the opposite, because there’s a desire to try something new. But Edward was not just saying that; he’s also making the statistical statement that “The hierarchical Bayesian approach can sometimes rely on quite strong modeling assumptions, which one might want to try to avoid in some settings — for example, if the assumptions are not accurate, then potentially severe bias could arise.” If he’d said, “Some of our collaborators are not familiar with the hierarchical Bayesian approach . . .”, that would be a different story.

  2. Great interview! I can add to a few of the threads here.

    I agree with Andrew’s general point that one can consider Bayesian / hierarchical approaches can just be considered “an estimator” and evaluated however one wishes, but I think the challenging aspect with this is that in the case of causal inference, “model fit” can be a bit hard to pin down due to the fundamental problem of causal inference. Anything like “looking at model fit” is almost entirely going to be about “outcome heterogeneity” (i.e. E[Y(0)|X=x]) rather than the actual model performance we care about, which is on _effect_ heterogeneity (i.e. E[Y(1)-Y(0)|X=x]). There’s _some_ relationship there, but it’s rarely straightforward.

    Working internally on HTE estimation, we tried to work on some diagnostic approaches (https://pubsonline.informs.org/doi/abs/10.1287/isre.2021.0343) that could be helpful, but I think there’s no silver bullet. The problems here are, I think, less severe with weakly regularized linear-based methods than full-on ML, so I don’t think HLMs would have been totally out of place so long as the priors weren’t too strong.

    In general, I’m no stickler for unbiasedness (and in other settings have proposed using biased shrinkage estimators in large-scale RCTs), but I think the case of US2020 required some additional care in a few ways. (1) Unbiasedness was very important for solving some social problems in the collaboration (e.g. what are the priors? how should we choose them?) (2) We had a lot of estimands we pre-registered (p20 of https://osf.io/brh3x), and the DR-learner provided a consistent way to estimate them all with appropriate flexible covariate adjustment (3) we had a ton of outcomes, so we just didn’t have enough time to carefully evaluate fit and iterate to find modeling choices for every single one.

    As it was, we pre-registered an ensemble of plugin models to DR-learner and let cross-validation select how to combine them. Point (3) was the core of the R package I wrote for doing all of this: we could pre-register the entire procedure of HTE estimation and use it consistently for each outcome (see, e.g. https://ddimmery.github.io/tidyhte/articles/methodological_details.html#replicating-analyses). I imagine we’re leaving potential variance gains on the table from more carefully selecting models, but we just didn’t have the time. In the replication materials, we have full model diagnostics on all of the models we used, so maybe someone can improve on what we did (these were far too much to even include in our existing 300p supplementary materials), but you have to apply for access, unfortunately: https://socialmediaarchive.org/record/43?ln=en

    In particular, the way I see it is that Edward’s DR-learner was a great match for us because it gave us the freedom to try out aggressive ML approaches to reduce variance (using automated ensembling through SuperLearner), but due to the fact that it was ultimately just AIPW in an RCT, we could be confident that nothing would go off the rails if a sub-optimal model were selected. This is surprisingly not the case for many methods!

    > How should researchers designing a study consider interest in HTEs at the design stage? […] perhaps the statistics literature has neglected design, compared with analysis

    I think this is a really important point!

    I have a paper that specifically comes at HTEs from this side: https://arxiv.org/abs/2010.11332. It asks the question: if we defined a model (KNN-regression) up-front, what would be the ideal way to assign units to treatment to make that model work well? In short, the answer is the algorithmic problem of MAXCUT. I think there’s a lot more that could be done in this space, too. For instance, we sketch out what the probabilistic form of this design would be, but we didn’t write out the full sampler for it, which would be a cool addition (and make this more usable on practical problems).

    At least in ML, I think design is often not prioritized as much as it should for cultural reasons as Dean said. Reading things like Donoho’s recent paper (https://arxiv.org/abs/2310.00865) illustrates this pretty well, I think. The idea of taking a dataset and code off of the shelf, modifying the method and seeing how well you do is very much in the blood of modern DS. Most of the causal work in ML is very much in this vein as well, but I don’t think it translates as well in this setting. There are a handful of these papers in ML journals each year (and maybe an increase over time?), but not that much. One could also consider the bandit literature here (as in response adaptive design), but most of that literature is so deep in the mathematical weeds it has little relationship to reality.

    • Thanks for this detailed reply. It would be great to follow up on questions about treatment effect estimates and calibration more.

      Just picking up on one little bit (“what are the priors? how should we choose them?”), I wonder if more explicit elicitation of the team members’ prior beliefs about these effects would have resulted in valuable changes to the studies. For example, perhaps after eliciting people’s priors for these effects, it would have turned out these studies would be too small or noisy to result in much of a Bayesian update. (This might have motivated, e.g., dropping more treatments, or focusing on different outcomes.)

  3. This is a sidenote, I guess, but I felt disappointed reading the “like-minded” paper in the context of the peer review file which was (refreshingly) published. Some really reasonable suggestions by the reviewers seem to have been ignored. To take one example, a reviewer gives very thoughtful comments about how one should interpret the results and suggests changing the title to avoid implying that the study suggests like-minded sources aren’t polarizing. The authors appear to have ignored or rejected this suggestion.

    I suppose that illustrates another benefit of public peer review: you can get perspective on published results whether or not authors/editors formally incorporate that perspective.

Leave a Reply

Your email address will not be published. Required fields are marked *