Unifying Design-Based and Model-Based Sampling Inference (my talk this Wednesday morning at the Joint Statistical Meetings in Toronto)

Wed 9 Aug 10:30am:

Unifying Design-Based and Model-Based Sampling Inference

A well-known rule in practical survey research is to include weights when estimating a population average but not to use weights when fitting a regression model—as long as the regression includes as predictors all the information that went into the sampling weights. But it is not clear how to apply this advice when fitting regressions that include only some of the weighting information, nor does it tell us what to do when analyzing already-collected surveys where the weighting procedure has not been clearly explained or where the weights depend in part on information that is not available in the data. It is also not clear how one is supposed to account for clustering in such analyses. We propose a quasi-Bayesian approach using a joint regression of the outcome and the sampling weight, followed by poststratifcation on the two variables, thus using design information within a model-based context to obtain inferences for small-area estimates, regressions, and other population quantities of interest.

No slides, but I whipped up a paper on the topic which you can read if you want to get a sense of the idea.

6 thoughts on “Unifying Design-Based and Model-Based Sampling Inference (my talk this Wednesday morning at the Joint Statistical Meetings in Toronto)

  1. This makes me think about the variety of situations in which the weights either are or aren’t (approximately) a deterministic function of x.

    Sometimes they definitely aren’t, like when the group making the weights has a bunch of valuable private information. This was the case in the COVID surveys in collaboration with Meta, including the one we ran.

    • Dean:

      Yes. And in addition to that sort of example, there are lots and lots of surveys that tell you to use the weights but don’t say where the weights come from. Sometimes you can reverse-engineer it by regressing log weight on background variables to see what you can explain. Other times the weighting is known but it uses non-census variables so you can’t just postratify. Or the weights can depend on variables that are relevant to the data collection but that otherwise you don’t really care about, so you don’t want to include them in x.

  2. “A well-known rule in practical survey research is to include weights when estimating a population average but not to use weights when fitting a regression model—as long as the regression includes as predictors all the information that went into the sampling weights.”

    What about doubly-robust approaches? Isn’t some variation on this what those techniques do? (Not an expert on this, just wondering what your opinion is on these techniques.)

  3. Andrew,

    Glad to see you still working on this–I’m remembering the pieces with comments and rebuttal you did several years ago on survey weights in Statistical Science. I appreciate the point re. clustered surveys as ppl. often ignore the issues with OLS estimates on clustered surveys.

    Best,
    Josh

Leave a Reply

Your email address will not be published. Required fields are marked *