What am I reading now?

Marshal Zeringue asked, and I replied:

I’ve been reading Personal Days, by Ed Park, which based on the reviews (and the first couple of chapters, which is what I’ve read so far) is a remake of Joshua Ferris’s Then We Come to the End, last year’s hilarious and claustrophobic satire on office-cubicle life. Ferris’s and Park’s books read like a cross between Geoffrey O’Brien and Don DeLillo, only funnier. Or like a fast-forward Richard Ford without the smugness. I’m impressed how novel-writing technique has improved in recent decades. John Updike was pretty slick, but Ferris and Park (and, for that matter, Ford) really seem in total control of their material, even in comparison to the masters of the previous generations. Sure, there was Nabokov (and, in his own way, James Jones), but that’s about it from back then. Now there seem to be a lot of novelists who really know what they’re doing in this way. (I think Jonathan Coe could have total control of his material too, if he really felt like it. He seems like Mailer or (Martin) Amis in his desire to shatter his own smooth surfaces.)

I’m also reading Fateful Choices: Ten Decisions that Changed the World, 1940-1941, by British historian Ian Kershaw. He goes into the historical evidence on how the leaders of Germany, Japan, the United States and the other WW2 participants made some of their seemingly inexplicable decisions. In addition to giving background on the historical personalities involved, Kershaw’s book is fascinating in how it focuses on the decision-making process within each country. In the writing style as well as in content, this book reminds me of A. J. P. Taylor’s classic Origins of the Second World War.

Mavericks of the past

Phil Klinkner writes:

History doesn’t repeat itself, the saying goes, but it does rhyme.

To me [Klinkner], the recent House defeat of the financial bailout bill echoes the defeat of the national sale tax in 1932. The Depression dried up federal revenues, so the Hoover administration proposed a national sales tax to raise money. Business and the leadership of both parties favored the bill, but the public was overwhelmingly opposed. Liberal Republican Fiorella LaGuardia led a bipartisan revolt against the bill. House Speaker John N. Garner actually left the speaker’s chair to go into the well and plead with his fellow Democrats to pass the bill. Garner normally had tight control on his party, but not this time. The bill was defeated 153-223.

In both cases, an unpopular Republican administration put forward a proposal to deal with an economic crisis, supported by the Democratic leadership in the House and the vast majority of the business community. Nonetheless, a bipartisan populist revolt sent it down to defeat.

And, Phil forgot to mention, James Garner was Maverick.

Nothing to do with statistical modeling, causal inference, or social science (except to the extent that all human endeavor is related to statistical modeling, causal inference, and social science)

When I was about 9 years old, I read just about every book of fairy tales in the library. 398.2 in the Dewey decimal system, I remember it well.

I also remember reading The Scarlet Pimpernel when I was about 15 and telling my mom how cool it was, and she said she really liked it too, she read it back when she was about 10. That kinda pulled the rug out from under me–not my mother’s intention, I’m sure–since I was always the one who was supposed to do things at an early age (perhaps the product of being the fourth child). Anyway, after going through the fairy tales section I made my way to science fiction, starting with Asimov etc. and then reading whatever else was on the library shelves.

At some point (maybe age 15 or so) I came across Barry Malzberg, who wrote stories in a perverse sort of barely readable formula which, if unpublished, would be hard to take seriously but because in print was somehow compelling. Either important, must-read stuff or despair squeezed out like toothpaste from a tube, who knows. At some point it becomes easier to read about this stuff than to read it directly, which has led me, decades later, to subscribe to the New York Review of Books instead of reaching for the latest Richard Ford, etc. Malzberg wrote a compellingly despairing memoir-like book of criticism which I bought when it came out around 1982. Right now, the segment of Engines of the Night that moves me the most is the surely-manipulative-but-in-this-case-it’s-gotta-be-ok “Cornell George Holey Woolrich: December 1903 to September 1968,” but the theme of Malzberg’s less dramatic struggles with commercial realities is compelling enough, and was so at the time also, despite my own continuing lucky separation from any financial needs.

I recently looked up Malzberg on Amazon and found he’d recently published Breakfast in the Ruins, an expansion of Engines of the Night. (Great titles.) I ordered the new book and read it through. He has tics in his writing, we all do, but it flows much better than his fiction, or maybe it’s just that I’m less interested in reading about spacement than about about middle-class, middle-aged literary life. Malzberg’s absolute lack of smugness (but not lack of self-regard) endlessly invigorates. (But it’s funny to read of his frustration about being ignored by the Alfred Kazins of the world, after reading of Kazin’s own frustrations about keeping his untidy life going while trying to get by as a book reviewer.)

The reason I recently looked up Malzberg–I’d liked (actually, enjoyed, although this doesn’t quite seem like the nice thing to say) Engines of the Night but hadn’t thought much about it in recent years–was that I’d recently read a dismissive piece on it in a collection of essays and reviews by Thomas Disch. Disch is another one: some of his stories are excellent (in particular, Casablanca is simultaneously poignant and nasty) and the poems are delightfully readable (my favorite, of course, is The Joycelyn Schrader Story), but what really reads like candy to me is his criticism. Especially his poetry criticism, where he mocks and mocks and mocks these poets who can’t get a real job and also throws in some judicious comments on those rare poets who have something interesting to say. Anyway, I recently from some source (probably Amazon again) learned that Disch had recently published two new books of criticism, one on poetry and one on science fiction–these were miscellaneous collections, not crafted work, but still, one takes what one can, and so I ordered them immediately and quickly wolfed them down when they did arrive. (To be more precise, a combination of almost no free time and a little bit of restraint allowed me to stretch these out for a month or so.) They were indeed treats and I thought I’d do everyone (or a few people) a favor by blogging on them, so I looked up Thomas Disch on the web, only to find that he had died . . . two days earlier. Shot himself, which is just hard for me to picture given his jaunty writing persona available from his essays as well as from his most famous story, The Brave Little Toaster. So, no blog entry then, but in the meantime I’d ordered the Malzberg book, which I’ve just finished reading, having no discipline to revise my multiple comparisons paper on this train ride.

Stories are supposed to begin in the middle but end at the end, but somehow I’ve done a Pulp Fiction here and ended in the middle. Maybe I’ll go and reread Overlay or one of the others and see if there’s something there that I’d missed.

P.S. Phil tells me that he recently reread Ubik (which we read when we were about 20) and says that he was amazed at how sloppy it is. Somehow, sloppiness seems more objectionable in a book than in a movie, perhaps because it is clearer how it all could be fixed. At the same time, some of what I see as the key moments in Ubik–for example, Runciter’s annoyance that he has to waste time in small talk when he’s trying to get advice from his dead wife–I could see these moments disappearing if the writing became too smooth.

Exciting 1% shift!

Brendan Nyhan offers this amusing example of a newspaper hyping poll noise. From the LA Times:

Registered voters who watched the debate preferred Obama, 49% to 44%, according to the poll taken over three days after the showdown in Oxford, Miss.

That is a small gain from a week ago, when a survey of the same voters showed the Democratic candidate with a 48% to 45% edge.

A small gain, indeed.

Blogs as places?

Henry Farrell referred here to his blog as a “place.” Which seemed funny to me because I think of a blog as a “thing.” Henry replied:

That’s the way that I [Henry] think about blogs (or at least group blogs and blogs with comments) – places where people meet up, chat, form communities, drift away from each other etc.

My analogy was blog-as-newspaper, the self-publishing idea, and I’m not used to thinking of a newspaper, or even a listserv, as a place. I think there is an aspect of the analogy that I’m still missing.

P.S. See Mark Liberman’s thoughts in his blog here.

The one advantage that we have over the New Yorker is that we have Google and they don’t

John Cassidy writes:

If Barack Obama is victorious on November 4th, someone on his transition team should send inauguration tickets to Richard Fuld, the chairman and chief executive of Lehman Brothers.

This is meant to be ironic, I believe: Fuld was the straw that broke the camel’s back and brought down the house of cards that was the American banking system etc etc.

But this got me wondering . . . could Cassidy’s statement be literally true??

I remember from talking with Tom Ferguson that, while the superrich generally favor the Republican Party, the financial sector is one area that leans Democratic. So I looked up Richard Fuld and–hey–here he is:

Richard Fuld, Lehman Rothers – Chairman, CEO 1994-present
Total donations since 1978: $208,550
To Democrats: 63%
To Republicans: 16%
To Special Interests: 21%

In 07/08, he seems to have been covering his bets: $10K each to the Republican and Democratic Senatorial Committees, $4600 to Hillary Clinton, $2300 to Barack Obama, $4600 to Chris Dodd, $2300 to John McCain, $2000 to John Reed in Rhode Island, and $10K to the “Securities Industry and Financial Markets Association Political Action Committee.”

So maybe he’ll get an invite to the inaugural party no matter who wins.

P.S. I was kinda hoping Fuld had only contributed to Obama–that would make a more interesting story. (Or I suppose if he’d only contributed to Republicans all his life, then there’d be an even better story of Richard S. Fuld, Jr., as a sleeper agent for the Democratic Party.) Actually, though, he’s all over the map, basically giving to almost every big name in the tri-state area and then some, including Pete Dawkins, Pete Du Pont, John Glenn for President (remember that?), Brendan Byrne, Phil Gramm, Joe Lieberman, Bob Dole, Al D’Amato and several of his opponents, Joe Lieberman, Jon Corzine, etc etc.

Difficulties in communication with non-Bayesians

My blog discussion with Eyal Shahar (see comments #3 and onward here) reminded me of a persistent challenge I face when talking with outsiders about Bayesian statistics.

Shahar’s basic objection to Bayesian methods is that (a) they’re are based on subjective probability, betting, etc., and (b) that’s not scientific. As he or she put it, “No doubt that Bayesian statistics offer a coherent system of logical inference, but its problem is irrelevance to the business of science.” My reply is to diagree with (a)–I claim that Bayesian data analysis is no more subjective than any other approach, I do not in general think of prior probabilities as degrees of belief, and I am not in general interested in estimating the probability that a hypothesis is true. But somehow it was difficult to make this point.

I had a similar frustrating feeling when talking about 15 years ago with one of my Berkeley colleagues: he said that he didn’t believe in Bayesian statistics, but if he did, he would want to use the real stuff, i.e., subjective probabilities. But there’s nothing in the mathematics that requires the probabilities to be subjective. And, just because some people use betting as a justification for this approach, that doesn’t mean that I have to use that justification.

This is one reason why, in writing Bayesian Data Analysis, we minimized philosophy and focused on empirical justifications, i.e., the method works in lots of examples.

Models for cumulative probabilities

Dan Lakeland writes:

I am working with some biologists on a model for time-to-response for animals under certain conditions. The model(s) ultimately are defined in terms of a differential equation that relates a (hidden) concentration of a metabolic product to the (cumulative) probability that an animal will respond within a given time by changing its behavior.

Now mostly, in my experience, statistical models are models for averages, or particular quantiles of the dataset (medians etc). Most models attempt to predict something (like time to response) from something else (like say measured amounts of a drug). In this case, rather than predicting individual response times, we’re trying to predict shape of a distribution from measured exposure to a certain environment.

In this case, we are tempted to use some measure of the goodness of fit to try to guess what is going on internally within the animal. For ease of computation, I’m fitting this model with maximum likelihood methods initially (a Bayesian approach may come later if time allows).

What is your opinion on model selection methods in this type of scenario? Your book index has “model selection and why we avoid it” which sounds unhelpful, but the section on model selection was actually more helpful than the index implied. Is there anything you can add in this context?

My reply: I’m not quite sure what your question is, but maybe, if I can translate it into the social-science examples with which I’m more familiar, I can imagine you’re doing something like predicting what percentage of people will respond a certain way to an advertisement, or how low a price would have to be before half the people would buy something. Framed that way, these sorts of models are pretty common. In section 6.8 of ARM, we discuss the relation between certain models for individuals and for groups.

“Blue parents who name with red values”

Laura Wattenberg has a fascinating discussion of the one topic you think you’ve already heard enough about . . . Sarah Palin’s kids’ names. You really have to read the whole thing, but here’s the gist:

No naming event has ever filled my [Wattenberg’s] inbox with as many reader queries as the unveiling of Sarah Palin–mom to Track, Bristol, Willow, Piper and Trig–as John McCain’s running mate. “Any comment?” “I’ve never heard Trig as a name for anything but a math class.” “Is this ‘an Alaska thing’?'”

In a way, yes, it is “an Alaska thing.” If you had nothing to go on but the baby names and had to guess about who the parents were, you’d guess that that they lived in an idiosyncratic, sparsely populated region of the country…and that they were conservative Republicans. . . .

For the past two decades, a core set of “cultural conservative” opinions has served as a theoretical dividing line between “red” (Republican/conservative) and “blue” (Democratic/liberal) America. These incude attitudes toward sex roles, the centrality of Christianity in culture, and a social traditionalism focused on patriotism and the family. If you were to translate that divide into baby names it might place a name like Peter—classic, Christian, masculine—on one side, staring down an androgynous pagan newcomer like Dakota on the other. In fact, that does describe the political baby name divide quite accurately. But it describes it backwards.

Characteristic blue state names: Angela, Catherine, Henry, Margaret, Mark, Patrick, Peter and Sophie.

Characteristic red state names: Addison, Ashlyn, Dakota, Gage, Peyton, Reagan, Rylee and Tanner. . . .

Why is it the blue parents who name with red values? Because in baby naming as in so many parts of life, style, not values, is the guiding light. . . .

What’s up with Kazakhstan?

Chris Zorn pointed me to this graph and asked for my thoughts. I replied that I’d seen worse, but the use of two dimensions doesn’t help, and the comparison to the GDP of Kazakhastan is just weird. I mean, who has any idea what is the GDP of Kazakhstan??

Chris replied,

I’m teaching first-year Ph.D. methods on PoliSci this term, and we have a feature called “Graph of the Day,” where — for five minutes or so at the beginning of every class — the students all look at and comment on some graph from a paper, the press, etc. I used this one yesterday, and the response (from people with a grand total of three weeks of graduate education) was identical: “What’s up with Kazakhstan?”, and “Isn’t a reference point supposed to be *non-obscure*?”

Confusion about the changing positions of political parties in the U.S.

The states won by the Democrats and Republicans in recent elections are almost the opposite of the result of the election of 1896:

1896a.png

In their article, “Activists and partisan realignment in the United States,” published in 2003 in the American Political Science Review, Gary Miller and Norman Schofield describe this as a complete reversal of the parties’ positions. In their story, in 1896 the parties competed on social (racial) issues, with the Republicans on the left and the Democrats on the right. Then the parties gradually moved around in the two dimensional social/economic issue space, until from the 1930s through the 1960s, the parties primarily competed on economic issues. Since then, in the Miller/Schofield story, the parties continued to move until now they compete primarily on social issues, but now with the Democrats on the left and the Republicans on the right.

It’s an interesting argument but I have some problems with it. First off, it was my impression that the 1896 election was all about economic issues, with the Democrats supporting cheap money and easy credit (W. J. Bryan’s “cross of gold” speech) and the Republicans representing big business. At least in that election, it was the Democrats on the left on economic issues and the Republicans on the right.

Getting to recent elections, the evidence from surveys and from roll call votes is that the Democrats and Republicans are pretty far apart on economic issues, again with the D’s on the left and the R’s on the right. So, from that perspective, it’s not the parties that have changed positions, it’s the states that have moved. The industrial northeastern and midwestern states have moved from supporting conservative economic policies to a more redistributionist stance. Which indeed is something of a mystery, and it’s related to attitudes on social issues, but I certainly wouldn’t say that economic issues don’t matter anymore. According to Ansolabehere, Rodden, and Snyder, social issues are more important now in voting than they were 20 years ago, but economic issues are still voters’ dominant concern.

1896 vs. 2000 by counties within each state

Here are some more pretty pictures. First, within 6 selected states, a scatterplot of Bush vote share in 2000 vs. McKinley vote share in 1896. There are completely different patterns in different states! Nothing like as clean a pattern as the statewide plot above.

1896b.png

And here’s another plot, this time showing each county as an ellipse, with the size of the ellipse proportional to the population of the county (more precisely, the voter turnout) in the two elections.

1896c.png

Nowadays the Democrats clearly do better in the big cities (in these graphs, the large-population counties). In 1896 the pattern wasn’t so clear.

The recent role of population density

I asked Jonathan Rodden what he thought of the above graphs, and he replied, “I would like to see when this relationship developed, in which states, etc. My hunch is that suburbanization, especially after the race riots, significantly reduced the heterogeneity of cities. The era of Democrats winning 80 percent of the presidential vote in big cities seems fairly recent.” He also sent along these graphs of voting by population density:

rodden.png

As Jonathan noted, the pattern of high-density areas voting strongly Democratic is relatively new. (But I don’t buy the way his lines curve up on the left; I suspect that’s an unfortunate artifact of using quadratic fits rather than something like lowess or spline.) Also there seems to be some weird discretization going on in the population densities for the early years in his data. But the main trends in the graphs are clear.

Jonathan added the following comment: “The relatively high values on the left side of graphs in early years is due to Southern Democrats and some mining districts. Graphs of the UK, Australia, and Canada look very similar during the same period, with left voting concentrated in urban and mining districts.”

Mellow liberals and jumpy conservatives

Jamie points out this interesting article by Douglas Oxley et al. that appeared in Science last month. Here’s the abstract:

Although political views have been thought to arise largely from individuals? experiences, recent research suggests that they may have a biological basis. We present evidence that variations in political attitudes correlate with physiological traits. In a group of 46 adult participants with strong political beliefs, individuals with measurably lower physical sensitivities to sudden noises and threatening visual images were more likely to support foreign aid, liberal immigration policies, pacifism, and gun control, whereas individuals displaying measurably higher physiological reactions to those same stimuli were more likely to favor defense spending, capital punishment, patriotism, and the Iraq War. Thus, the degree to which individuals are physiologically responsive to threat appears to indicate the degree to which they advocate policies that protect the existing social structure from both external (outgroup) and internal (norm-violator) threats.

I myself am extremely sensitive to sudden noises, so make of that what you will . . . Seriously, though, this seems related to John Jost’s work on personality profiles and political affiliation.

Drew Linzer’s poll tracker

Drew Linzer writes:

I read your paper with some interest, as within the last week or so I’ve started analyzing the state tracking polls available on pollster.com using a simple Bayesian mean model, and posting my results here.

The model updates the Obama share of the Obama-McCain vote as new polls come in, and then calculates the posterior probability that that proportion is greater than 0.5.

I’m doing this as more of a hobby than anything else…was frustrated with analyses I’ve seen that strike me as overly complex, unstable, totally opaque, and frankly, fairly unbelievable — there’s a real chance Obama could get EVs in the 400s? come on. The R code I use to generate my graphs and predictions are also posted if you click the Data tab.

My goal was to create a model that was very simple, but made reasonable predictions. So, for example, I don’t have any time component — the model assumes the true state level proportion for Obama is constant. Maybe this is a good assumption, probably not, but also probably the actual within-state support numbers are not actually fluctuating as much as some of the predictions out there make it seem. The first time I set up the model, I just used the posterior from one poll as the prior on the next. Problem was that the variance of the priors got so small after a while that there didn’t seem to be enough flexibility to capture trends when they did seem to arise. So then I added a multiplier to increase the variance of the prior in proportion to how many days old the last poll was. I tinkered around with it a bit and came up with this that seemed reasonable.

sd.prior.flex <- sd.prior.flex * (1+(0.05*log(dat$daysold[i]+1))) Changing the 0.05 makes the trend line more or less sensitive to new polls. It's really just kind of acting as a smoother as the polls appear. The other thing the model doesn't have is any sort of cross-state correlation structure built in. Every state is treated as its own independent entity. This probably isn't very realistic either, but in these battleground states there also seems to be enough state-level polling going on to get decent enough within state estimates. Where I would like to take account of cross-state correlation is when I simulate election results. Doesn't seem like that would be too hard to estimate in a second stage after the trendlines have been calculated (or simultaneously in a more complicated model), I just haven't gotten around to it. Anyway, so that's basically it. the "predicted electoral vote" adds up the EVs for each candidate who my trend line has above 50%. And the simulation that produces the "probability of winning" and the histogram just draws from each state's most current posterior distribution 1 million times and adds up the number of Obama EVs. As I said, I don't know if the model is "good" but it is clean and transparent, relatively stable, and, to my mind, produces predictions that accord better with the available polling data.

My thoughts: First, I don’t think it’s impossible that Obama could get EV’s in the 400s, given the uncertainties in national forecasts. It’s not likely but it’s possible. Second, I think poll aggregation is fine (whether it be Drew’s method or Realclearpolitics or 538.com or whatever), but when it comes to forecasts, I think the best thing is a weighted average of polls and model-based predictions, with the model having two parts: (1) the national popular vote and (2) the states relative to each other.

Things I saw while waiting for the train

When the sign says the train will be 0:05 late, it won’t be 0:05 late. If it were going to be 0:05 late, they wouldn’t say anything at all. In reality it will be 0:30 late. But they won’t say it will be 0:30 late, because that would mean the train will be 1:30 late.

I got off a good line when I got on the train. I stepped in, saw a retirement-age couple already seated, and asked, Philadelphia? They said, yeah, that’s where they’re going, but they’re not sure either, they hope they’re in the right place. I said, yeah, I think this is right. (Pause) And, if there’s one thing last week’s news has taught us, it’s that you can trust a guy in a suit.

That got a laff.

Ads in the Newark train station

A big picture of a hot-dog guy holding a mustard-slathered beauty, next to the words: You Want Cancer With That? Medical research shows hot dogs increase your risk of cancer… [An ad for some law firm that’s suing food manufacturers.]

The Retreat at Princeton
Inpatient Alcohol and Drug Treatment for Executives & Professionals
www.RetreatAtPrinceton.com

P.S. I have no idea why, but this particular entry seems to attract a log of spam!

Approximate Bayesian inference using integrated nested Laplace approximations

The following is my discussion of the article, “Approximate Bayesian inference for latent Gaussian models by using integrated nested Laplace approximations” by H. Rue, S. Martino and N. Chopin, for the Journal of the Royal Statistical Society:

Statisticians often discuss the virtues of simple models and procedures for extracting a simple signal from messy noise. But in my own applied research I constantly find myself in the opposite situation: fitting models that are simpler than I would like—models that clearly miss important features of the data and, more importantly, important features of the underlying system I am modeling—because of computational limitations.

In some sense, “computational limitations” correspond to limited CPU time and memory. But in this age of gigabytes and more, it’s only fair to describe these as limitations on our computational procedures. I am routinely in the position of wanting to fit a model that can’t be fit using existing software, even though I know—know—that a simple enough algorithm must be out there to fit it using much less than the capabilities of a modern desktop PC.

The sorts of models I’m talking about include hierarchical models for parallel time series (for example, trends in public opinion in each of 50 states, or models for stochastically aligning tree ring data) and varying-intercept, varying-slope logistic regressions (that is, models where several coefficients can vary by group, in which case a covariance matrix needs to be modeled for the group-level structure).

In practice when fitting such models I lurch between various approximate methods based on point estimates, and full Gibbs-Metropolis which can be slow if not guided well. These two approaches can meet in the middle: approximations can be iteratively adjusted, leading ultimately to a Gibbs-like stochastic procedure, and Markov chain simulation can be made more efficient and reliable when guided by approximations that have been tailored to the problem at hand.

I welcome the article by Rue, Martino, and Chopin because it provides a more general way to construct these approximations. I suspect that, in addition to being a competitor to Gibbs and Metropolis, this approach ultimately can be used to make these stochastic algorithms more efficient.

As noted in the article, a challenge remains with problems with many hyperparameters, which are often themselves modeled hierarchically. As with the EM algorithm, it appears to be tricky to apply this method to a hierarchy with more than three levels, and I look forward to these researchers’ future efforts in this area. It might help to model the hyperparameters explicitly rather than to consider them as unconstrained in some potentially large space.

I conclude with a remark on the comment in Section 7 of the article, that MCMC is often perceived to be “exact” even though in practice it is not. Fifteen or twenty years ago, MCMC itself had to fight this misconception in another form. At the time, importance sampling was viewed as an exact method with MCMC as a sometimes necessary but unfortunate approximation. There was much discussion of how MCMC and importance sampling could work together, and ideas about starting with MCMC and then finishing up with importance sampling to get an exact result. Fortunately these ideas have subsided, as computational statisticians realized that actually existing importance sampling is not exact but can instead be viewed as just another iterative simulation method, and one that has no particular advantages over the Metropolis algorithm or other more clearly iterative approaches (Gelman, 1991).

Additional reference

Gelman, A. (1991). Iterative and non-iterative simulation algorithms. Computing Science and Statistics 24, 433-438.