What Is Probability?
Probability is one of the most successful branches of mathematics. It underlies quantum mechanics, statistical mechanics, machine learning, actuarial science, genetics, signal processing, gambling, weather prediction, and the theory of computation. Its formalism is clean, its theorems are powerful, and its applications are endless.
And yet nobody agrees on what it means.
This is not a pedagogical gap—it is not the case that the experts understand probability and the students are confused. The experts themselves disagree, sharply and irreconcilably, about what a probability is. Is it a frequency? A degree of belief? A property of a physical system? A feature of our ignorance? The answers are mutually exclusive, and each has devastating objections. The mathematics works regardless, which is either reassuring or deeply suspicious.
This essay traces the problem from the axioms through the interpretations into the physics and then into the philosophy. The thesis is that probability, information, and computation are the same problem viewed from different angles, and that the problem is not solved.
I. The Mathematics
Kolmogorov’s Axioms
In 1933, Andrey Kolmogorov ended decades of confusion about the foundations of probability by reducing it to measure theory. His axioms are brief and elegant:
A probability space is a triple $(\Omega, \mathcal{F}, P)$ where:
• $\Omega$ is a sample space (set of all outcomes)
• $\mathcal{F}$ is a $\sigma$-algebra on $\Omega$ (collection of events, closed under complement and countable union)
• $P : \mathcal{F} \to [0, 1]$ is a probability measure satisfying:
$$1. \; P(\Omega) = 1$$ $$2. \; P(A) \geq 0 \text{ for all } A \in \mathcal{F}$$ $$3. \; \text{If } A_1, A_2, \ldots \text{ are pairwise disjoint, then } P\!\left(\bigcup A_i\right) = \sum P(A_i)$$ That is all. Everything else—conditional probability, independence, random variables, the law of large numbers, the central limit theorem—is derived from these three axioms.
Notice what the axioms do not say. They do not say what $P(A)$ means. They do not say where $\Omega$ comes from. They do not say how to assign $P$ to any particular event. They give you a calculus of probability—rules for manipulating the symbol $P$—without ever telling you what $P$ refers to. This is by design. Kolmogorov, in the introduction to his 1933 book, explicitly declined to interpret the axioms. He gave mathematics. He left philosophy to others.
The axioms are isomorphic to the axioms of measure theory. A probability measure is simply a measure on a space with total measure $1$. This is mathematically powerful (you inherit the entire apparatus of Lebesgue integration, Radon-Nikodym derivatives, martingale theory) but philosophically empty. Saying “probability is a measure” tells you no more about what probability means than saying “length is a measure” tells you what length is. Length, at least, has an obvious physical referent. Probability’s referent is the question.
What the Formalism Gets You
Regardless of interpretation, the formalism yields theorems. The important ones:
Bayes’ Theorem
P(A|B) = P(B|A) P(A) / P(B). This is not controversial. It is a theorem—a logical consequence of the axioms and the definition of conditional probability P(A|B) = P(A ∩ B) / P(B). What is controversial is how to use it, because it requires P(A)—the prior—and the axioms do not tell you where the prior comes from.
The Law of Large Numbers
If $X_1, X_2, \ldots$ are i.i.d. with mean $\mu$, then the sample mean converges to $\mu$ (in probability, and almost surely). This is the theorem that connects probability to frequency. It says that if you have a well-defined probability distribution, then the frequencies will converge. It does not say that frequencies define probabilities. The distinction matters.
The Central Limit Theorem
The sum of many independent random variables, suitably normalized, converges to a Gaussian. This is why the bell curve appears everywhere—not because the world is intrinsically Gaussian, but because sums of independent factors produce Gaussian distributions regardless of the factors’ individual distributions. It is a theorem about convergence in distribution, and it is one of the most powerful results in all of mathematics.
None of these theorems require you to know what probability is. They are relationships between probabilities. Given a measure P satisfying the axioms, these theorems follow. The interpretation of P is a separate question.
II. The Three Interpretations
Frequentism
The frequentist says: P(A) is the limiting relative frequency of A in an infinite sequence of identical, independent trials. Flip a coin infinitely many times; the fraction of heads converges to P(heads). This is what probability is.
Frequentism has the appeal of seeming objective and empirical. It grounds probability in observable quantities. It dominated 20th-century statistics and still dominates most science education. The p-value, the confidence interval, and the hypothesis test are all frequentist constructs.
The problems are severe:
- No infinite sequences exist. You cannot flip a coin infinitely many times. The definition refers to a limit that cannot be realized. This makes the frequentist definition either an idealization (fine in mathematics, but then it’s no longer empirical) or a claim about what would happen (a counterfactual, which is a philosophical concept, not an empirical one).
- Unique events have no frequency. What is the probability that the earth’s average temperature exceeds 2°C above pre-industrial levels by 2100? There is no ensemble of earths to take a frequency over. The frequentist must either refuse to assign a probability (which is practically useless) or smuggle in some notion of “similar events” (which requires judgment about what counts as “similar”—which is a Bayesian move in disguise).
- The reference class problem. Even for repeatable events, the frequency depends on which class of trials you group the event into. A 45-year-old male with high blood pressure and a family history of heart disease has a different mortality probability depending on whether you reference “all humans,” “all 45-year-old males,” or “all 45-year-old males with high blood pressure.” The frequency is not unique. It depends on a choice, and the axioms do not make the choice for you.
Bayesianism
The Bayesian says: P(A) is a degree of belief. It quantifies the subjective uncertainty of an agent about whether A is true. Different agents with different information may assign different probabilities to the same event, and this is not a defect—it is the point. Probability is always relative to a state of knowledge.
Bayesianism is grounded in Dutch book arguments and Cox’s theorem. The Dutch book argument shows that if your degrees of belief do not satisfy the probability axioms, then a clever bookie can construct a set of bets that guarantee you lose money. Cox’s theorem shows that any system of plausible reasoning satisfying certain desiderata (consistency, universal applicability, agreement with Boolean logic) must be isomorphic to probability theory. Both arguments say: if you want to reason consistently under uncertainty, you must use probability theory. It is not a choice. It is the unique consistent framework.
The Bayesian approach handles unique events naturally. “The probability that the temperature exceeds 2°C by 2100” is simply your degree of belief given your evidence. It changes as you get more evidence (via Bayes’ theorem). There is no need for infinite sequences or reference classes.
The problems:
- The prior. Bayes’ theorem updates a prior into a posterior. But where does the prior come from? If two Bayesians start with different priors, they may reach different conclusions from the same data. In the limit of infinite data, priors wash out (Bayesian consistency theorems guarantee this under mild conditions). But you never have infinite data. In practice, the prior encodes assumptions, and those assumptions are often doing the real work.
- Subjectivity. If probability is a degree of belief, then it is subjective. Many scientists find this deeply uncomfortable. They want probability to be a property of the world, not of their minds. The Bayesian response is that objectivity comes from the data, not the prior—and that the pretense of objectivity in frequentist methods is exactly that, a pretense. The debate is unresolved.
- Whose beliefs? When a physicist says the electron has a 30% chance of being spin-up, is this a statement about the physicist’s beliefs, or about the electron? The Bayesian answer is: the physicist’s beliefs. The physicist finds this unsatisfying. (We return to this in the quantum mechanics section.)
The Measure-Theoretic / Formalist View
The formalist says: probability is a mathematical structure defined by axioms. It is no more in need of philosophical interpretation than the real numbers or topology. You set up a measure space, prove theorems, and apply the results. Questions about “what probability really is” are extramathematical and not the mathematician’s problem.
This view is coherent. It is also evasive. The moment you apply probability theory to anything—physics, statistics, medicine, engineering, gambling—you must interpret the measure P as representing something in the world. The formalist framework does not tell you how to do this. It gives you a hammer and refuses to discuss nails.
The uncomfortable summary
Frequentism tells you what probability means but can’t handle the cases that matter most. Bayesianism handles all cases but requires you to accept subjectivity and priors. Formalism avoids the question. No interpretation is satisfactory. The mathematics works anyway.
III. The Philosophical Abyss
Is Probability Epistemic or Ontic?
This is the central philosophical question. Is probability a feature of the world (ontic) or a feature of our knowledge about the world (epistemic)?
If I say “the probability that this coin lands heads is 50%,” am I describing a property of the coin (it is built in a way that produces heads half the time) or a property of my ignorance (I don’t know enough about the initial conditions, air currents, and table surface to predict the outcome)?
The classical answer, from Laplace, is firmly epistemic:
“Probability is relative, in part to this ignorance, in part to our knowledge.”
For Laplace, a sufficiently powerful intellect—Laplace’s Demon—who knew the position and velocity of every particle in the universe could predict everything with certainty. There would be no need for probability. Probability exists only because we are not that intellect. It is a measure of our ignorance, not of the world’s randomness.
This view held for a century. Then quantum mechanics arrived.
IV. Quantum Mechanics and the Dice of God
The Measurement Problem
In quantum mechanics, the state of a system is described by a wave function $\psi$, which evolves deterministically according to the Schrödinger equation:
Probability enters through measurement. When you measure an observable on a system in state $\psi$, you get a random outcome. The probability of getting eigenvalue $\lambda_n$ is $|\langle\phi_n|\psi\rangle|^2$, where $\phi_n$ is the corresponding eigenstate. After measurement, the state “collapses” to $\phi_n$. This is the Born rule.
The Born rule is not derived from the Schrödinger equation. It is an additional postulate. The Schrödinger equation says the wave function evolves smoothly and deterministically. The Born rule says that upon measurement, the wave function collapses randomly. These two rules are in tension—one is deterministic, the other stochastic; one is smooth, the other discontinuous. Where, exactly, does the “measurement” happen? What counts as a measurement? The Schrödinger equation does not contain the word “measurement.” This is the measurement problem, and it has been open since 1926.
Does God Play Dice?
Einstein believed the probabilities in quantum mechanics were epistemic—that quantum mechanics was incomplete, and that some deeper theory with hidden variables would eventually restore determinism. “God does not play dice,” he wrote to Max Born in 1926. The quantum randomness, he insisted, must reflect our ignorance of some underlying reality, just as the randomness in coin flips reflects our ignorance of classical mechanics.
In 1964, John Bell proved that this is (almost certainly) wrong.
Bell’s Theorem (1964)
Any theory that is (1) locally realistic—outcomes are determined by pre-existing local properties, and no influence travels faster than light—must satisfy certain statistical inequalities (Bell inequalities) in the correlations between measurements on entangled particles. Quantum mechanics predicts violations of these inequalities. Experiments (Aspect 1982, Gisin 1998, Hensen 2015, and many others) confirm the violations with overwhelming statistical significance.
Therefore: either locality fails (influences travel faster than light), or realism fails (outcomes are not determined by pre-existing properties), or both. There is no local hidden variable theory that reproduces quantum predictions.
Bell’s theorem does not strictly prove that the universe is random. It proves that any deterministic explanation of quantum mechanics must be non-local—it must involve faster-than-light correlations. This is possible (Bohmian mechanics does exactly this), but it comes at a steep price. The result is that the “obvious” view—particles have definite properties, we just don’t know them, and measuring reveals them—is ruled out by experiment.
The Interpretations
The measurement problem has produced a remarkable collection of interpretations, each of which gives a different answer to the question “is probability fundamental?”:
| Interpretation | Is the universe random? | What is $\psi$? | What is measurement? |
|---|---|---|---|
| Copenhagen | Yes. Measurement outcomes are irreducibly random. | A tool for computing probabilities. Not physically real. | A primitive notion. Don’t ask what it is. |
| Many-Worlds (Everett) | No. Every outcome occurs in some branch. There is no collapse. | Physically real. The universal wave function is all there is. | Decoherence causes branching. All branches are real. |
| Bohmian Mechanics | No. Particle positions are determined by a guiding equation. The randomness comes from ignorance of initial conditions. | A real pilot wave guiding real particles. | An interaction between particle and apparatus governed by the guiding equation. |
| QBism | The question is ill-posed. Probability is an agent’s degree of belief. | An agent’s personal tool for decision-making. Not a physical object. | An action an agent takes on the world. |
| Objective Collapse (GRW) | Yes. Collapse is a real physical process (spontaneous localization). | Physically real. Subject to rare, random collapses. | Not special. The same collapse process, amplified by the apparatus. |
These interpretations all predict the same experimental outcomes (so far). They differ only in their metaphysics. This means the question “is the universe fundamentally random?” may be empirically undecidable—a question that experiment cannot answer, because the competing answers make identical predictions.
Note the extraordinary situation. In Many-Worlds, there is no randomness at all—the Schrödinger equation evolves deterministically and the appearance of probability comes from the observer’s perspective within a branch. In Bohmian mechanics, there is no randomness either—particles follow deterministic trajectories and the apparent randomness comes from ignorance of initial conditions (exactly Laplace’s view, transplanted to quantum mechanics). In Copenhagen and GRW, the randomness is fundamental. In QBism, the question is deflected—probability is about the agent, not the world.
The mathematical formalism of quantum mechanics—Hilbert spaces, unitary evolution, Born rule—is the same in every interpretation. The physics is the same. Only the philosophy differs. But the philosophy determines the answer to the question that started this section.
“I, at any rate, am convinced that He does not throw dice.”
“Einstein, stop telling God what to do.”
The honest answer, a century later: we do not know whether God plays dice. We know that if He doesn’t, the deterministic mechanism must be non-local. We know that the mathematics of probability works either way. And we know that the question may be unanswerable by experiment.
V. Information
Shannon’s Definition
In 1948, Claude Shannon did for information what Kolmogorov did for probability: he axiomatized it. His definition:
Shannon Entropy
$$H(X) = -\sum p(x) \log_2 p(x)$$The information content (entropy) of a random variable $X$ with distribution $p$ is the expected number of bits needed to specify an outcome. It is maximized when all outcomes are equally likely (maximum uncertainty) and zero when one outcome is certain (no uncertainty).
Notice that Shannon’s entropy is defined in terms of probability. If probability is already on shaky philosophical ground, then so is information. Shannon was aware of this. He deliberately avoided giving information a physical or semantic interpretation. “The fundamental problem of communication,” he wrote, “is that of reproducing at one point either exactly or approximately a message selected at another point. Frequently the messages have meaning; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem.”
Shannon’s information is syntactic, not semantic. It measures surprise, not meaning. A string of random bits has maximum entropy (maximum information in Shannon’s sense) despite being meaningless. A Shakespeare sonnet has lower entropy (it’s highly structured and compressible) despite being profoundly meaningful. This is either a limitation of Shannon’s definition or a feature, depending on your purposes.
The Deep Connection to Probability
Information and probability are duals. Consider:
- The information content of an event with probability $p$ is $-\log_2(p)$ bits. A fair coin flip ($p = 1/2$) carries 1 bit. An event with probability $1/1024$ carries 10 bits. Rare events carry more information. This makes intuitive sense: learning that something unlikely happened tells you more than learning that something likely happened.
- Entropy is expected information content. It is the average surprise. It measures how uncertain you are before you observe the outcome.
- Bayes’ theorem is an information-processing rule. The posterior encodes the information in the prior plus the information in the data. Bayesian updating is literally the process of incorporating new information.
- KL divergence $D(P \| Q) = \sum p(x) \log(p(x)/q(x))$ measures the information cost of using the wrong distribution $Q$ when the true distribution is $P$. It is the number of extra bits wasted by a code optimized for $Q$ when the data comes from $P$.
Jaynes formalized the connection in his Maximum Entropy Principle: given some constraints (known expected values, symmetry properties), the probability distribution that makes the fewest additional assumptions is the one that maximizes Shannon entropy subject to those constraints. This gives you a principled way to derive priors. If you know nothing, you should be maximally uncertain—maximum entropy. If you know the mean and variance, the maximum entropy distribution is the Gaussian. If you know only that the variable is positive with a given mean, it’s the exponential.
Jaynes argued that probability theory is the logic of science—the unique consistent framework for reasoning under uncertainty—and that entropy is the bridge between information and probability. Probability measures what you don’t know. Information is what reduces that uncertainty. They are two faces of the same concept.
Is Information Physical?
Here the question takes a turn from mathematics into physics.
Rolf Landauer proposed in 1961 that erasing a bit of information has an unavoidable thermodynamic cost: it must dissipate at least $kT \ln 2$ of energy as heat (about $3 \times 10^{-21}$ joules at room temperature). This is Landauer’s principle. It has been experimentally confirmed.
The connection between Shannon entropy and thermodynamic entropy is not a metaphor. It is a mathematical identity (up to a constant factor). Boltzmann’s entropy $S = k \ln \Omega$, where $\Omega$ is the number of accessible microstates, is exactly $k \ln 2$ times the Shannon entropy of the uniform distribution over those microstates. The thermodynamic entropy of a gas is the amount of information you would need (in bits) to specify the exact microstate, multiplied by $k \ln 2$.
If erasing information costs energy, then information is physical—it is not merely an abstraction but something that has measurable thermodynamic consequences. This was Landauer’s claim: “Information is physical.” It cannot be created or destroyed without corresponding changes in the physical state of the universe.
Conservation of Information
In classical mechanics, Liouville’s theorem states that the volume of phase space is preserved under Hamiltonian evolution. If you start with a set of possible initial conditions occupying some volume in phase space, that volume is conserved as the system evolves. The shape may distort, stretch, and fold, but the volume remains constant.
This is a conservation law for information. The phase-space volume is (exponentially related to) the entropy, and Liouville’s theorem says it is conserved. The information content of the system—how many bits you would need to specify which microstate you’re in—is constant. Information is neither created nor destroyed by the laws of classical mechanics.
In quantum mechanics, the analog is unitarity: the Schrödinger equation preserves the inner product between states. No information is lost in unitary evolution. Two initially distinct states remain distinct forever. This is not an optional feature of quantum mechanics—it is built into the mathematical structure.
The black hole information paradox arises precisely because black holes appear to destroy information. Matter falls in, the black hole evaporates via Hawking radiation, and the radiation appears to be thermal—it carries no information about what fell in. If this is correct, information is not conserved, and unitarity is violated. This would be a crisis for quantum mechanics. Most physicists believe information is preserved (the “information is somehow encoded in the radiation” view), but proving this requires a theory of quantum gravity that we do not yet have. The debate has driven some of the most important developments in theoretical physics over the past 50 years, including the holographic principle and the AdS/CFT correspondence.
That arguments about the nature of information have driven fundamental physics should tell you something about the depth of the concept.
VI. Computation and Irreducibility
What Does Probability Have to Do with Computation?
Return to Laplace’s Demon. In principle, said Laplace, a sufficiently powerful intellect could predict the future with certainty. If the universe is deterministic (as classical mechanics claims), then probability is merely a measure of computational inability—we use it because we can’t compute the exact trajectory, not because the trajectory is indeterminate.
But what if the computation itself is the bottleneck? What if there is no shortcut—no way to predict the future of a system without literally simulating it step by step?
This is Stephen Wolfram’s concept of computational irreducibility. Some systems (like elementary cellular automata Rule 30) have simple, deterministic rules but produce behavior that cannot be predicted by any method faster than running the system itself. There is no closed-form solution, no shortcut, no compression. The only way to know what the system does at step $10^9$ is to simulate all $10^9$ steps.
If a system is computationally irreducible, then even Laplace’s Demon, with perfect knowledge of initial conditions and laws, cannot predict the future without spending at least as much computational work as the universe itself spends in evolving. Prediction requires simulation, and simulation takes time. In this sense, the future is “random” not because it is indeterminate but because it is incompressible—there is no shorter description of the outcome than the outcome itself.
Kolmogorov Complexity and Randomness
A string is algorithmically random (in the sense of Kolmogorov, Solomonoff, and Chaitin) if the shortest program that produces it is approximately as long as the string itself. It is incompressible. This is the formal definition of randomness in the theory of computation.
Notice that this definition makes no reference to probability distributions, frequencies, or degrees of belief. A string is random if and only if it has no pattern, no structure, no shorter description. Randomness is the absence of compressibility.
This connects back to Shannon: an incompressible string has maximum Shannon entropy. And it connects to probability: if you cannot compress the string, you cannot predict the next bit better than chance—which is what it means for the bits to be “random” in the probability sense.
Computational irreducibility suggests a view of probability that is neither purely epistemic nor purely ontic. The system is deterministic (the rules are known, the initial conditions are fixed), but the outcome is effectively random because no agent embedded in the universe can compute it in advance. The randomness is real for any finite observer, even though the system itself is deterministic. Is this “true” randomness? The question may not have a meaningful answer.
VII. The Synthesis (Or Lack Thereof)
We have arrived at the following picture:
- Probability is a mathematical formalism (Kolmogorov) that can be interpreted as a frequency (frequentist), a degree of belief (Bayesian), or left uninterpreted (formalist). No interpretation is fully satisfactory.
- Information is measured by entropy (Shannon), is connected to probability by a mathematical identity, and appears to be physical (Landauer). It is conserved by the fundamental laws of physics (Liouville, unitarity), and its possible non-conservation (black holes) is a crisis for physics.
- Quantum mechanics introduces probability at the fundamental level (Born rule), but whether this probability is ontic (irreducible randomness) or epistemic (hidden variables, many worlds, Bayesian agents) depends on the interpretation, and the interpretations are empirically indistinguishable.
- Computation provides a framework (Kolmogorov complexity, computational irreducibility) in which deterministic systems can produce effectively random behavior because no finite agent can predict the outcomes faster than the system itself generates them.
These four threads are not independent. They are the same problem.
Probability is what you use when you cannot predict. Information is what you gain when you observe. Entropy measures both your uncertainty (probability) and the disorder of the system (thermodynamics). Computation determines whether prediction is possible in principle. And quantum mechanics either tells you that prediction is fundamentally impossible (Copenhagen) or that the appearance of unpredictability is an artifact of your position within a deterministic but branching universe (Many-Worlds) or a deterministic but non-local one (Bohm).
The core question, restated:
Is there a fact of the matter about whether the universe is random, or is “random” a relationship between an observer and a system?
If randomness is a relationship between observer and system, then probability, information, and computation are all aspects of the same thing: the limits of what a finite agent embedded in the universe can know and predict. This is close to the Bayesian/information-theoretic view, and it is the view I find most coherent. But it requires you to give up the idea that probability is a property of the world independent of any observer. Many physicists find this unacceptable.
If randomness is a property of the world itself (ontic randomness), then the universe contains a source of genuine novelty—outcomes that are not determined by any prior state. This is the Copenhagen/GRW view, and it fits naturally with the Born rule. But it raises the question: where does the randomness come from? What is the mechanism? If there is no mechanism—if “it just is random”—then we have traded one mystery (what is probability?) for another (what is the source of fundamental randomness?).
I do not think this question has been answered. I am not sure it is answerable. What I am sure of is that probability, information, and computation are not three separate subjects. They are one subject, viewed from the perspectives of mathematics, physics, and computer science, respectively. And the subject is: the relationship between an observer and the world.
“Information is not a disembodied abstract entity; it is always tied to a physical representation. It is represented by engraving on a stone tablet, a spin, a charge, a hole in a punched card, a mark on paper, or some other equivalent. This ties the handling of information to all the possibilities and restrictions of our real physical word, its laws of physics and its storehouse of available parts.”
“Probability theory is nothing but common sense reduced to calculation.”
Perhaps Laplace was right, and probability is common sense. Perhaps common sense is all we have, and the question of whether the universe “really” is random is the kind of question that dissolves rather than gets answered. The mathematics works. The physics works. The engineering works. The philosophy remains open. Maybe that is the answer.