What should be the correct reference class? If you want to know whether your income is high, should you compare yourself to all Americans, all Californians, all residents of San Jose, or just your neighborhood? If you want to estimate your risk of heart disease, should you look at the general population, all men, or all men over 40?
This question — which comparison group is “right” for assigning a probability to an individual case — is one of the oldest open problems in the foundations of probability. Venn identified it in The Logic of Chance (1866; 1876 second edition). Reichenbach gave it a name in 1949: the reference class problem. Hájek argued in 2007 that it afflicts every interpretation of probability, not just frequentism. Philosophers have been debating it for 150 years.
An answer has been sitting in the statistics toolbox for decades.
The structure of the problem
The philosophical reference class problem contains two separable questions: (1) given a grouping structure, how do you combine information across levels? and (2) which grouping structure is the right one? Hierarchical models do not solve the second question. They do give the natural answer to the first.
Move to a narrower reference class and the estimate becomes less biased — the group is more representative of the individual. But it also becomes more variable — you have fewer observations. Move to a broader class and variance drops, but bias increases. The “right” reference class is wherever these two forces balance.
This is the bias-variance tradeoff. And the solution is not to pick a level of the hierarchy. It is to combine them.1
Partial pooling
You have two extreme options. You can completely pool — ignore the grouping and use the population average for everyone. That’s low variance but high bias: it assumes your neighborhood is no different from the national average. Or you can refuse to pool at all — estimate each group entirely from its own data. That’s low bias but high variance: your neighborhood of 20 people gives you a garbage estimate.
Partial pooling is the compromise. The estimate for each group is a weighted average of the group mean and the population mean. When the group has a lot of data, the estimate stays close to the group mean. When the group has little data, the estimate shrinks toward the population mean, because the group mean is too noisy to trust on its own. And the degree of shrinkage depends on how much groups genuinely differ from one another: when between-group variance is large relative to within-group variance, the model trusts the group-level data more; when it’s small, the model pulls harder toward the population.
This is why it’s called partial pooling — it’s neither fully pooled (one estimate for everyone) nor fully unpooled (each group on its own), but a continuous interpolation between the two, with weights learned from data. No one needs to decide which reference class is “correct.” The model learns how much to borrow from each level of the hierarchy.
Multilevel models implement this logic. The machinery comes in several flavors. Frequentist mixed-effects models produce shrunken group-level estimates and can be understood as empirical Bayes — the prior is estimated from the data rather than being specified. Fully Bayesian hierarchical models put priors on the variance components and propagate uncertainty all the way through. For the purpose of the reference class problem, the distinction is secondary — what matters is the shared logic of partial pooling.
The formula for the simplest two-level case makes the weighting explicit (the generalization to deeper hierarchies is straightforward but notationally heavier):
Why the grouping hierarchy is problem-dependent
One thing the model does not tell you is which hierarchy to use in the first place. For income, geography dominates — cost of living makes national comparisons nearly meaningless, so the natural grouping is geographic (state, metro area, neighborhood). For disease risk, age-sex-ancestry is the natural hierarchy. There is no universal nesting structure, and the choice of hierarchy is substantive, not statistical.2
This is exactly what Venn and Reichenbach were worried about. Partial pooling provides the natural answer to problem (1). Problem (2) remains genuinely hard, but it is a modeling question — what features are predictively relevant — not a foundational crisis about the nature of probability.
The residual philosophical question
I don’t want to overclaim. Partial pooling is the best operational answer to the reference class problem, but it doesn’t fully dissolve the philosophical question. You still need to specify the hierarchy — which variables define the groups, and how they nest. Two analysts who choose different hierarchies will get different partial-pooling estimates, and the model itself cannot adjudicate between them without additional data or assumptions.
In this sense, the reference class problem has been reduced rather than solved. It has been reduced from “which group is correct?” to “which grouping structure is predictively useful?” — a substantive scientific question rather than a foundational paradox. That’s a lot of progress, even if it’s not a complete dissolution. And it leaves entirely untouched the deeper single-case problem — what a probability means for a specific individual — which is a question about the interpretation of probability itself, not about estimation.
Why is the synthesis still missing?
The bridge between the philosophical reference class problem and partial pooling in multilevel models is less explicit in the published literature than it deserves to be. There are papers that come close, and a few that say something strikingly similar, but the point has mostly appeared in side channels rather than in the central texts of either field.
The philosophical side has produced extensive work on the puzzle itself. Hájek’s The Reference Class Problem Is Your Problem Too (2007) remains the canonical source. Wallmann and Williamson’s Four Approaches to the Reference Class Problem (2017) discusses objective Bayesian solutions, combinatorial approaches, similarity-based methods, and mechanism-based approaches — but not hierarchical partial pooling. Roth and Tolbert’s Resolving the Reference Class Problem at Scale (2025) is especially interesting because it uses multicalibration from the algorithmic fairness literature to handle a high-dimensional, overlapping, non-nested version of the problem. That is closely related terrain but it’s different from multilevel partial pooling. And Wallmann’s A Bayesian Solution to the Conflict of Narrowness and Precision in Direct Inference (2017) is, in spirit, even closer to shrinkage: broader-class frequencies matter more when the narrow class is small and recede as the narrow class grows.
The statistics side has been using partial pooling as a practical answer for decades. Lindley and Smith’s Bayes Estimates for the Linear Model (1972) gives an early exchangeability-based hierarchical formulation; Efron and Morris’s Data Analysis Using Stein’s Estimator and Its Generalizations (1975) makes the shrinkage logic vivid in empirical-Bayes form; Gelman and Hill’s Data Analysis Using Regression and Multilevel/Hierarchical Models (2006) popularized the now-standard compromise between no pooling and complete pooling. The broader mixed-effects tradition, and the software ecosystems around Stan and PyMC, all rely on the same basic logic. But they do not usually frame it in these terms. They talk about shrinkage, borrowing strength, exchangeability, and the bias-variance tradeoff, not about Venn’s consumptive Englishman visiting Madeira.
The two communities are solving closely related problems with different vocabularies and have not fully recognized the overlap. Perhaps the most striking case is actuarial credibility theory. Credibility practice dates back to 1918, and in its modern Bühlmann form it has the same shrinkage structure as a simple random effects / empirical Bayes estimator. Actuaries have been navigating this tradeoff for over a century without using the term “reference class problem.”
There are also narrower literatures that come closer to stating the parallel. The closest approach comes from David Poole’s group at UBC. Poole’s Relations, generalizations and the reference-class problem: A logic programming / Bayesian perspective (2004) explicitly frames relational prediction as a reference-class problem and sketches a Bayesian approach with hierarchical priors and averaging over models rather than selection of a single best class. Chiang and Poole’s Reference Classes and Relational Learning (2012) develops the technical framework, showing how relational probabilistic models — especially those with latent properties — can be understood in terms of reference classes. Their solution is different from partial pooling across observed hierarchies described in this post, but the motivating problem is the same and the Bayesian machinery is closely related. Wang’s masters thesis Hierarchical Structure and Ordinal Features in Class-based Linear Models (2021), also supervised by Poole, implements a “bounded ancestor method” that constructs Dirichlet priors from parent classes and mixes them with observed data — a close analogue of partial pooling through a class hierarchy — explicitly motivated by the reference class problem. This line of work appeared in AI and relational learning venues, which may be why it hasn’t entered the conversation in either philosophy of probability or mainstream statistics.
From a different direction, Dawid’s On Individual Risk (2017) treats the reference class problem using exchangeability-based ideas that sit very near hierarchical modeling, without making the final step to partial pooling.
So the right claim is not that nobody has said this. It’s that the point has not been made central. The philosophy literature contains the puzzle in its classic form. The statistics literature contains the machinery. The Poole group has connected the two in AI venues. What is still rare is a clear statement, in the core vocabularies of philosophy or statistics, that once the grouping structure is fixed, this old philosophical problem becomes the same estimation problem that partial pooling was built to handle.
Statisticians learn hierarchical models as tools; nobody in a statistics seminar says, “This is a response to a 150-year-old problem in the foundations of probability.” And nobody in a philosophy seminar says, “The answer is lme4.”
The takeaway
If you’re a philosopher who worries about reference classes: the best answer is multilevel modeling with partial pooling. It won’t tell you which variables to condition on, but once you’ve chosen a hierarchy, it combines information across levels in a data-driven way. The “correct” reference class is not a choice among discrete alternatives; it’s a continuous interpolation, and the interpolation weights are learned from data.
If you’re a statistician: every time you fit a multilevel model, you are addressing a 150-year-old foundational problem in the philosophy of probability. Shrinkage toward the population mean is not just a variance-reduction trick. It’s an answer to the reference class problem — a question philosophers have been debating since Venn.
The bridge is there. It just hasn’t been walked across often enough.
Appendix: technical extensions
Exchangeability as the formal bridge: The deepest connection between the reference class problem and hierarchical models runs through de Finetti’s representation theorem. Judging observations within a group as exchangeable — their joint probability doesn’t depend on their order — entails that they can be represented as i.i.d. conditional on a latent parameter. Partial exchangeability (exchangeable within groups but not across them) leads directly to multilevel models. The judgment of exchangeability is, formally, the Bayesian analog of choosing a reference class. This is the mathematical reason partial pooling gives the natural answer to the reference class problem: once you commit to an exchangeability structure, the hierarchical model is the natural way to combine information across groups.
Discovering the hierarchy from data: An objection to partial pooling is that it requires pre-specifying the grouping structure. Nonparametric Bayesian methods — Dirichlet process mixtures, hierarchical Dirichlet processes (Teh et al., 2005) — can relax how much of the latent structure must be specified in advance. They do not eliminate the substantive question of which features and invariances matter, but they show that the hierarchical Bayesian framework has resources beyond the basic partial-pooling formula.
The bias-variance decomposition and the specific shrinkage weights shown here assume squared-error loss; under other loss functions the tradeoff takes a different form, but the resolution is the same — a continuous mixture across levels rather than a discrete selection of one.
The same logic extends to crossed, non-nested groupings — an individual may belong to reference classes that don’t nest, like aircraft manufacturer and airline operator — and multilevel models handle these naturally through crossed random effects. It also applies when the “reference class” is continuous rather than discrete: in kernel regression, smoothing splines, or Gaussian processes, the effective reference class becomes a neighborhood in the feature space, and the bandwidth or length scale determines how much information is borrowed from nearby cases.