10 Comments
User's avatar
Niclas's avatar

Very interesting. The curse seems even stronger when extended to log normal distributions where it's a power rather than a linear term!

Justin Purnell's avatar

Mathematical proof that MacKenzie Scott is better than Jeff Bezos

Savio's avatar

Seems like https://www.lesswrong.com/posts/nnDTgmzRrzDMiPF9B/how-much-do-you-believe-your-results also talked about this phenomenon without explicitly naming it.

Roman's Attic's avatar

Good post! I liked seeing this application of "regression to the mean" type reasoning.

Just out of curiosity, which specific EA groups (other than GiveWell) do you think aren't thinking about this enough? I think your model implicitly assumes that the average EA grantmaker looks at a list of interventions and throws all their money into the one with the highest expected impact based on past interventions, but my current understanding is that a lot of grantmakers/organizations are mostly looking for people interested in doing semi-new interventions that have a strong theory of change, high estimated impact (based on plausibility of success and estimated scope), and a capable team. Is your criticism mostly of people who are earning to give to existing charities directly?

titotal's avatar

Semi-new interventions with a high estimated impact are exactly the type of causes we would expect to be hit by the optimisers curse. For every semi-new intervention which gets proposed, there are probably 10 more that you did the estimates for and didn't get brought up. And we would also expect semi-new interventions to have higher uncertainty and higher chances of large errors than established causes like malaria prevention. As I showed here, this is a recipe for overestimation and bias toward uncertainty. Again I will repeat, the errors do not need to be caused by statistical noise in past trials, they can be errors anywhere in your estimation methodology.

I think basically every EA org which ranks charities on any basis should be thinking about how the curse affects their results. Givewell are by far the best on this, but discussion by other orgs seems extremely sparse.

Noah Birnbaum's avatar

how much do you think this affect AI x risk stuff?

titotal's avatar

I'm going to give a followup on this in a couple of weeks, I believe I have a pretty good curse-based argument for x-risk being overestimated. However, i think these dynamics make comparing cross-cause effectiveness pretty difficult, because it's affected by the distribution of "true" effectiveness, which is very hard to estimate for something like x-risk.

Roman's Attic's avatar

Thanks for your response!

I think there aren’t a ton of orgs that your critique actually applies to. I’m not even sure a lot of impact evaluation methods are sophisticated enough to have some of these problem outside of global health and development. A lot of nuclear war GCRR evaluations done by orgs seem to be much more like “let’s make a decision tree to try to guess how these 40 different agents and orgs react in response to what we’re trying to do and see if our intervention actually has a positive sign in a high enough number of them,” and I get the feeling that a decent amount of AI grantmaking is somewhat similar. Also, groups like Coefficient Giving seem to mostly be looking for orgs that have a decent chance of doing a lot of good to receive medium-sized grants (many groups face pretty large diminishing marginal returns beyond the grants they receive), rather than easily scalable functions to dump all your money into once you find the “highest EV”.

I think the critique you’re making most plausibly applies to biorisk funding and groups like Animal Charity Evaluators, but even then, I’m not fully sold that this is the case. I don’t know much about biorisk groups, but a lot of ACE’s work is about getting more information about existing interventions. They often make the groups they fund collect a lot more data about the work they do as part of the terms for receiving the grant, and a lot of the funding they give in the most speculative areas like fish, insect, and WAW is primarily about research. The conclusion of your critique could be something along the lines of “amongst the charities that ACE recommends, some of them aren’t nearly as high impact as others, and we don’t really know which ones those are” and like, yeah, sure. This is known, and part of the reason why speculative charities are funded is so we can learn more about them.

To briefly summarize the things I find most inaccurate about your critique: I think your modeling assumptions seem to treat EA evaluation orgs as mere impact estimators/auditors who try to estimate the impact per dollar of a whole bunch of different existing charity f(x)s and tell people about where all the money should go, and this is an incomplete view of the roles actually performed by charity evaluators (such as enforcing stricter research standards), the uncertainty that charity evaluators obviously have and try to communicate, and what charity evaluators are looking for (in theory they want the highest impact thing, but in practice trying to do the most good looks a lot more like allowing a whole bunch of smaller interventions to get started and attempt to do good).

titotal's avatar

You don't need sophistication for the curse to hit you. You just need ranking + decent chance of error. It sounds like in your nuclear risk example, you are still choosing some interventions over others according to some metric (say, percentage of positive sign reactions from other orgs). Even if you are completely unbiased in your estimate of each intervention, as long you have some error in how well you estimate that metric, the curse will kick in, and your top interventions will probably have overestimated metric values (in this case, less of the org outcomes will be positive than you thought).

I will be exploring portfolio strategies a bit more in my next post, they mitigate the problem a bit but do not eliminate it.

However, I do agree that getting more information is good! I'll show this more in the next post, but one consequence of the curse is that lowering the error in your estimates can pay off massively.

Sam Harsimony's avatar

Excellent post. It would be interesting to expand this model so that you're not selecting a single option, but investing in a portfolio of options with the goal of minimizing regret. In a perfect world, you'd put all your money on the best intervention, but when you're uncertain and risk-averse, it's better to spread your investment out.

This is a nice framework to consider the value of information too. How valuable is it to run a larger study to reduce the error bars? If smaller error bars wouldn't put an intervention in the top-3 for example, it may not be worthwhile to run the study at all.