Bayes Factor
Which hypothesis the data prefer, and by how much. A Bayes factor is the evidence ratio between two hypotheses — a Bayesian answer where the p-value gives a one-sided one.
- Term
- Bayes factor
- Is
- Evidence ratio between two hypotheses
- Above 1 favors
- The first hypothesis
- Unlike p-value
- Can support the null too
Parts of speech & senses
- A Bayes factor is the ratio of the marginal likelihood of the observed data under one hypothesis to its marginal likelihood under another, quantifying which hypothesis the data support and by how much. "A Bayes factor of 12 gave strong evidence for the new design."
What a Bayes factor is
A Bayes factor is a number that says which of two hypotheses the data prefer, and by how much. Formally, it is the ratio of the probability of the observed data under one hypothesis to the probability of that same data under a rival hypothesis — each hypothesis's marginal likelihood. If the data are, say, twelve times more probable under hypothesis A than under hypothesis B, the Bayes factor comparing A to B is twelve. A value above one favors the first hypothesis, below one favors the second, and near one means the data barely distinguish them. Because it is a ratio of how well each hypothesis predicted what actually happened, a Bayes factor is a direct measure of relative evidence, and it is the quantity that turns your prior odds between two hypotheses into your posterior odds after seeing the data.
To read the magnitude, researchers lean on interpretive scales such as the one Harold Jeffreys proposed. Roughly, a Bayes factor between one and three is only anecdotal evidence, three to ten is moderate, ten to thirty is strong, thirty to one hundred is very strong, and above one hundred is extreme — with the reciprocals applying to evidence for the other hypothesis. These labels are conventions, not laws: a Bayes factor of 2.9 and one of 3.1 are not meaningfully different even though they fall either side of a boundary, so you report the actual number, not just the category. The practical appeal is that the Bayes factor speaks in the intuitive language of odds — this hypothesis is eight times better supported than that one — rather than in the language of tail probabilities.
Bayes factor versus p-value
The Bayes factor is most useful to contrast with the p-value, because they answer different questions and are constantly confused. A p-value is the probability of data at least as extreme as what you saw, assuming the null hypothesis is true. It only ever addresses the null, and it cannot provide evidence for the null — a large p-value means not enough evidence against, which is not the same as evidence for. A Bayes factor is symmetric: it compares two hypotheses directly and can come out in favor of either, so it can actually support the null that nothing is going on, not merely fail to reject it. That distinction matters whenever a null result is scientifically interesting, such as showing a treatment has no meaningful effect.
There are further differences worth keeping straight. A p-value says nothing about how probable a hypothesis is; it is a statement about data under an assumption. A Bayes factor, combined with prior odds, updates the probability of the hypotheses themselves. The Bayes factor also depends on how each hypothesis is specified — in particular on the prior distribution assumed for effects under the alternative — so two analysts can get different Bayes factors from the same data by choosing different priors, and that modeling choice must be stated and defensible. Neither number is a magic verdict. But where a p-value answers whether the data are surprising under the null, a Bayes factor answers which hypothesis the data favor, and by how much — usually the question people actually wanted answered.
Using Bayes factors well
Use a Bayes factor when you want to weigh two specific hypotheses against each other and are willing to state them precisely — including the prior for effects under the alternative, since the Bayes factor depends on it. Report the actual value, not just its Jeffreys-scale label, and be honest that a factor near one means the data are inconclusive rather than that the null is confirmed. Its great strength is the ability to gather evidence for a null, which lets you distinguish we found no effect from we could not tell, a distinction p-values cannot make. Run a sensitivity check across reasonable priors so a conclusion does not hang on one arbitrary choice, and remember that a Bayes factor is relative evidence between the two hypotheses you named, not a verdict on any hypothesis you did not.
The failures mirror those of significance testing. Treating the Jeffreys categories as hard cutoffs recreates the same false dichotomy that mindless use of the 0.05 threshold produces. Hiding the prior, or picking one that quietly favors the answer you want, makes the Bayes factor look objective while it is doing the opposite. Reading a Bayes factor as the probability that a hypothesis is true confuses the evidence ratio with the posterior it only helps produce. And comparing a well-chosen hypothesis against a straw-man alternative can manufacture strong-looking evidence. The discipline is to specify both hypotheses honestly, state and stress-test the priors, report the number and its uncertainty, and treat the Bayes factor as a transparent measure of relative evidence rather than a certificate of truth.
Synonyms & antonyms
Synonyms
Antonyms
Origin & history
A Bayes factor is the ratio of the evidence two hypotheses give the observed data, named for Thomas Bayes and popularized by Harold Jeffreys.
Etymology: source.
Usage trends
Search interest for this term over the last five years:
Common questions
- What is a Bayes factor?
- A Bayes factor is the ratio of how well two competing hypotheses predict the observed data. Above one favors the first hypothesis, below one the second, and it turns your prior odds between them into posterior odds after seeing the data.
- How is a Bayes factor different from a p-value?
- A p-value only measures how surprising data are under the null and can never support the null. A Bayes factor compares two hypotheses directly and can favor either one, so it can actually provide evidence that no effect exists.
- How do you interpret the size of a Bayes factor?
- Roughly, using the Jeffreys scale, one to three is anecdotal, three to ten moderate, ten to thirty strong, and above one hundred extreme, with reciprocals for the other hypothesis. These are conventions, so report the exact value, not just the label.
Resources & people to follow
- referenceRGM analysis — definitions, senses, and usage verified per term
Curated, non-competitor resources verified per term.
Related training
Disciplines
Areas of marketing where bayes factor is a core concern: