Art Owen informed me that he’ll be teaching sampling again at Stanford, and he was wondering about ideas for students gathering their own data.
I replied that I like the idea of sampling from databases, biological sampling, etc. You can point out to students that a “blood sample” is indeed a sample!
Art replied:
Your blood example reminds me that there is a whole field (now very old) on bulk sampling. People sample from production runs, from cotton samples, from coal samples and so on. Widgets might get sampled from the beginning, middle and end of the run. David Cox wrote some papers on sampling to find the quality of cotton as measured by fiber length. The process is to draw a blue line across the sample and see the length of fibers that intersect the line. This gives you a length-biased sample that you can nicely de-bias. There’s also an interesting out there about tree sampling, literally on a tree, where branches get sampled at random and fruit is counted. I’m not sure if it’s practical.
Last time I found an interesting example where people would sample ocean tracts to see if there was a whale. If they saw one, they would then sample more intensely in the neighboring tracts. Then the trick was to correct for the bias that brings. It’s in the Sampling book by S. K. Thompson. There are also good mark-recapture examples for wildlife.
I hesitate to put a lot of regression in a sampling class; It is all too easy for every class to start looking like a regression/prediction/machine learning class. We need room for the ideas about where and how data arises and it’s too easy to crowd those out by dwelling on the modeling ideas.
I’ll probably toss in some space-filling sampling plans and other ways to down size data sets as well.
The old Cochran style was: get an estimator, show it is unbiased, find an expression for its variance, find an estimate of that variance, show this estimate is unbiased and maybe even find and compare variances of several competing variance estimates. I get why he did it but it can get dry. I include some of that but I don’t let it dominate the course. Choices you can make and their costs are more interesting.
I connected Art to Scott Keeter at Pew Research, who wrote:
Fortunately, we are pretty diligent about keeping track of what we do and writing it up. The examples below have lengthy methodology sections and often there is companion material (such as blog posts or videos) about the methodological issues.
We do not have a single overview methodological piece about this kind of work but the next best thing is a great lecture that Courtney Kennedy gave at the University of Michigan last year, walking through several of our studies and the considerations that went into each one:
Here are some links to good examples, with links to the methods sections or extra features:
Our recent study of Jewish Americans, the second one we’ve done. We switched modes for this study (thus different sampling strategy), and the report materials include an analysis of mode differences https://www.pewresearch.org/religion/2021/05/11/jewish-americans-in-2020/
Jewish Americans in 2020: Answers to frequently asked questions
Our most recent survey of the US Muslim population:
U.S. Muslims Concerned About Their Place in Society, but Continue to Believe in the American Dream
A video on the methods:
https://www.pewresearch.org/fact-tank/2017/08/16/muslim-americans-methods/This is one of the most ambitious international studies we’ve done:
Here’s a short video on the sampling and methodology:
https://www.youtube.com/watch?v=wz_RJXA7RZM
We then had a quick email exchange:
Me: Thanks. Post should appear in Aug.
Scott: Thanks. We’ll probably be using sampling by spaceship and data collection with telepathy by then.
Me: And I’ll be charging the expenses to my NFT.
In a more serious vein, Art looked into Scott’s suggestions and followed up:
I [Art] looked at a few things at the Pew web-site. The quality of presentation is amazingly good. I like the discussions of how you identify who to reach out to. Also the discussion of how to pose the gender identity question is something that I think would interest students. I saw some of the forms and some of the data on response rates. I also found Courtney Kennedy’s video on non-probability polls. I might avoid religious questions for in-depth followup in class. Or at least, I would have to be careful in doing it, so nobody feels singled out.
Where could I find some technical documents about the American Trends Panel? I would be interested to teach about sample reweighting, e.g., raking and related methods, as it is done for real.
I’m wondering about getting survey data for a class. I might not be able to require them to get a Pew account and then agree to terms and conditions. Would it be reasonable to share a downsampled version of a Pew data set with a class? Something about attitudes to science would be interesting for students.
To which Scott replied:
Here is an overview I wrote about how the American Trends Panel operates and how it has changed over time in response to various challenges:
Growing and Improving Pew Research Center’s American Trends Panel
This relatively short piece provides some good detail about how the panel works:
https://www.pewresearch.org/fact-tank/2021/09/07/how-do-people-in-the-u-s-take-pew-research-center-surveys-anyway/We use the panel to conduct lots of surveys, but most of them are one-off efforts. We do make an effort to track trends over time, but that’s usually the way we used to do it when we conducted independent sample phone surveys. However, we sometimes use the panel as a panel – tracking individual-level change over time. This piece explains one application of that approach:
https://www.pewresearch.org/fact-tank/2021/01/20/how-we-know-the-drop-in-trumps-approval-rating-in-january-reflected-a-real-shift-in-public-opinion/When we moved from mostly phone surveys to mostly online surveys, we wanted to assess the impact of the change in mode of interview on many of our standard public opinion measures. This study was a randomized controlled experiment to try to isolate the impact of mode of interview:
From Telephone to the Web: The Challenge of Mode of Interview Effects in Public Opinion Polls
Survey panels have some real benefits but they come with a risk – that panelists change as a result of their participation in the panel and no longer fully resemble the naïve population. We tried to assess whether that is happening to our panelists:
Measuring the Risks of Panel Conditioning in Survey Research
We know that all survey samples have biases, so we weight to try to correct those biases. This particularly methodology statement is more detailed than is typical and gives you some extra insight into how our weighting operates. Unfortunately, we do not have a public document that breaks down every step in the weighting process:
Most of our weighting parameters come from U.S. government surveys such as the American Community Survey and the Current Population Survey. But some parameters are not available on government surveys (e.g., religious affiliation) so we created our own higher quality survey to collect some of these for weighting:
How Pew Research Center Uses Its National Public Opinion Reference Survey (NPORS)
This one is not easy to find on our website but it’s a good place to find wonky methodological content, not just about surveys but about our big data projects as well:
We used to publish these through Medium but decided to move them in-house.By the way, my colleagues in the survey methods group have developed an R package for the weighting and analysis of survey data. This link is to the explainer for weighting data but that piece includes links to explainers about the basic analysis package:
https://www.pewresearch.org/decoded/2020/03/26/weighting-survey-data-with-the-pewmethods-r-package/
Lots here to look at!
It’s been awhile since I’ve taught a course on survey sampling. I used to teach such a course—it was called Design and Analysis of Sample Surveys—and I enjoyed it. But . . . in the class I’d always have to spend some time discussing basic statistics and regression modeling, and this always was the part of the class that students found the most interesting! So I eventually just started teaching statistics and regression modeling, which led to my Regression and Other Stories book. The course I’m now teaching out of that book is called Applied Regression and Causal Inference. I still think survey sampling is important; it was just hard to find an audience for the course.
” it was just hard to find an audience for the course.”
This is great! Back in the old days the faculty and people in the profession decided what courses were necessary for given activities and the students were required to take them. Totally old school, I know! In the new student centered concept of academia, if “The Role of Teddy Bears In Mental Health” is the course that will provide an audience, that’s what we teach! My how education has improved! I’m looking forward to the social benefits of Student Centered Education!!
You are living in a nightmare world constructed entirely inside your own head. Please go outside
While I wholeheartedly disagree with the suggestion that students in general are shying away from difficult subjects these days, Chipmunk’s comment made me wonder: What is the ‘right’ content of a lesson?
In ancient Rome, there were traditional disciplines that students who could afford it would receive. In medieval and Renaissance continental Europe, there were three traditional disciplines/degrees that could be obtained (law, medicine and theology). Usually a young man* could not pay for his own education, and since there were not many scholarships or grants available, teachers had to convince the source of the money (usually the parents) of their worth. When Prussia introduced compulsory education for children/young people, it was to produce soldiers who could read and write. Education was funded by the government. A little later, with the industrial revolution, education was re-discussed, as it increased the flow of productivity, but also allowed workers to become more assertive (demanding wages, occupational health and safety, democracy, social insurances, etc.).
It seems to me that there is a tradition in education that the content of teaching has often been chosen for the students rather than by the students. Only recently (given how long humans have been teaching each other) has there been a large-scale shift towards considering the needs of the students. This certainly has a positive effect on motivation to engage with the subject. It is also likely to reduce the tendency of the sciences to stick to orthodoxy/ their traditions (for better or worse). Whenever I find myself in a teaching situation, it is a much nicer experience for me, tge educator, when the audience is engaged than if it is not. Sadly, while the students will be more inclined to push the class in directions that are relevant to them, the majority will at the same time try to avoid the difficult topics, whether they are important or not.
All in all, it seems to me that some democratic elements in education will further the cause**. On the other hand, I conclude that the teacher cannot just follow the wishes of his audience. Once again, it seems that the optimum lies between two poles.
*At first I wrote ‘young people’, but that would be a misrepresentation of who was being educated at the time. It would also diminish the struggle of the brave women and men who fought for women’s access to education – a struggle that is still going on in some places.
**I deliberately omit discussions of the exact goals of education and claim everybody has some intuition of the subject.
I would argue that the driving force behind this process is the increasing availability of affordable teaching materials, starting with the invention of high quality writing paper, the printing press that made written text available to the general public and the translation of books from the lingua franca, latin, to vernacular. There is also the increasing centralization all over Europe that led to lack of literate administrators.
“In medieval and Renaissance continental Europe, there were three traditional disciplines/degrees that could be obtained…”
Not really. Universities aren’t really representative of the middle ages. Usually, to persue a scholarly career, one had to become a monk or nun and would then enjoy relatively high degrees of scholarly freedom. In general, the literacy rate was incredible low, basically only the clergy knew how to read and write, even the majority of the nobility didn’t. This is not really surprising. To copy a bible on parchment one needs about 200 sheeps and a lot of time. Paper really made a difference. Maybe they should have given cuneiform a try, though it would have been difficult to get those beautiful illustrations on clay tablets.
@Anonymous:
Agreed. When I was writing about the Middle Ages last night, I was thinking more of the Renaissance. For example, Martin Luther switched from studying law to theology (against his parents’ wishes). You are right, in the so-called Middle Ages education was very much centred around monasteries. However, I would disagree with a general “freedom of research” for monasteries. In some places, monks and nuns were indeed free to pursue their scholarly interests (e.g. the works of Hildegard of Bingen were a milestone in medical research). I know that on the Iberian peninsula there was a period of great peace and cooperation between Jewish, Christian and Muslim scholars under Muslim rule. It ended when the Christians pushed the Muslims out.
Fun fact: I saw a drawing of an early university lecture. In the front row, the students were looking excited, taking notes and seemingly discussing the content of the lecture. In the back rows the students were asleep, staring at holes in the ceiling and totally distracted. Sound familiar? ;)
The mention of whales reminds me of a talk I heard at Berkeley about detectability functions, in this case estimating whale populations from sightings of whales by pairs of observers on a survey ship. The pairs included a biologist and a Norwegian whaler, and the main conclusion seemed to be that whalers see more whales than biologists do.
This report https://escholarship.org/uc/item/3r6120x8 on estimating age structure of salmon catches is interesting for what it says about the practical sampling difficulties as well as the maths of various approaches
Another example is validation sampling where you have a variable Z on everyone and a way to get a more accurate value X on a subsample (eg self-reported illness from memory vs medical records). Here the question is which people to sample to get the most information, and it’s especially interesting when Z is reasonably accurate.
Wasn’t there an issue that the act of sampling arctic polar bears itself was biasing the results? Ie, the bears viewed it as a threat so would move inland to avoid the humans, which then also made their (the bears’) lives more difficult.
Nothings coming up on a search but that would make sense. You could also see how being sloppier around camp (even in previous years) could increase the number of bears, etc.
In mammalogy, the effect of trapping in catch and release studies is well known. These effects can be especially important in mark and recapture studies. Mammals, being generally pretty smart, can learn pretty quickly that traps are bad (frightening encounters with large predators from which you escape alive, but you never want to go through that again) or good ( a source of easy to get tasty food, warm bedding for the night, and none the worse for wear), and thus become trap-shy or trap-loving. Either effect can bias a study. The same effects can occur in purely observational sampling (people can be repellent or attractive), as Anoneuoid notes for polar bears.
My courses in sampling were a while ago: Kish’s textbook taught by Kish (and Robert Groves, Graham Kalton, and Martin Frankel) in windowless classrooms in Ann Arbor.
What I remember was a lot of detail on variance estimation (what’s called the Cochran approach in the post). But then I went into marketing research, where the problems of nonresponse, bias, and shifting panel samples over time were paramount. These were covered, but at the back of the book. The problems I saw weren’t much about the variance of a properly executed sample. The problems were how to correct for an inevitably flawed sample. (e.g. Pew’s data showing how response rates have tanked in our lifetimes).
But, as I said, my academic work was a while ago, and I hope the curricula have shifted appropriately.