Difference between revisions of "Bayes' theorem"

From Conservapedia
Jump to navigation Jump to search
Line 1: Line 1:
 +
'''Bayes' Theorem''', also known as Bayes' Rule, is used in [[statistics]] and [[probability]] theory to relate marginal probabilities and conditional probabilities.  In the context of [[Bayesian probability theory]], it is used to update degrees of belief (probabilities) given new information. The theorem is attributed to the Reverend [[Thomas Bayes]] (1702-1761) an English nonconformist minister, but it is likely to have been known to earlier mathematicians such as [[Pierre Laplace]].
  
  
−
'''Bayes' theorem''' is a result in [[probability theory]] that shows how to invert a [[conditional probability]]. An extended form of the theorem provides a rigorous [[mathematics|mathematical]] way of learning from uncertain data. This latter form is the basis of [[Bayesian statistics]].
+
== Theorem ==
  
−
Let A and B be two uncertain events. We denote the probability of A by p(A) and probability of B by p(B). We denote the event that A and B are both true by AB, and the probability of AB by p(AB). In general, the fact that A has already occurred will affect the chances that B will occur, and vice versa. We denote the probability that B will occur, given that A has occurred, by p(B|A). We call this the probability of B, conditional on A. The event AB can occur in two ways. A can occur first, then B occurs, or B occurs first and then A occurs. The rules of probability require that the probability of two independent events both occurring is the product of the probabilities. Thus
+
Bayes' theorem is stated mathematically as
  
 +
<math> P(X|Y,I)=\frac{P(Y|X,I)P(X|I)}{P(Y|I)} </math>,
  
−
                    p(AB) = p(A)p(B|A) = p(B)p(A|B)
+
where X and Y are independent statements, I is available background information, and
  
−
This can be rearranged to be
+
<math>P(X|Y,I)</math> is the [[posterior probability]] for ''X'' given ''Y'' and ''I'',
  
−
                            p(B|A) = p(B)p(A|B)/p(A)
+
<math>P(Y|X,I)</math> is the [[likelihood]] for ''Y'' given ''X'' and ''I'',
  
−
This is Bayes' theorem. The theorem is attributed to the Reverend [[Thomas Bayes]] (1702-1761) an English nonconformist minister, but it is likely to have been known to earlier mathematicians such as [[Pierre Laplace]].
+
<math>P(X|I)</math> is the [[prior probability]] for ''X'' given only ''I'', and
  
−
The theorem has many applications.
+
<math>P(Y|I)</math> is sometimes called the [[evidence (statistics)|evidence]] or probability for ''Y'' given only ''I''.
  
−
Example 1.
+
In a scientific context, ''X'' may be a [[hypothesis]], and ''Y'' may be experimental [[data]]. The theorem can then be used to determine the degree of belief in the hypothesis by using the experimental data.
 +
 
 +
 
 +
== Applications and Examples ==
 +
 
 +
One popular example of the use of Bayes' theorem is the [[Monty Hall problem]], inspired by the [[television]] show ''Let's Make a Deal''.
 +
 
 +
Another example of how Bayes' theorem would be used is:
  
 
Suppose a particular disease afflicts 1% of the population. Suppose that a test for the disease is 95% accurate. Suppose that someone tests positive for the disease but there is no other evidence that they have the disease. What is the probability that they have the disease?
 
Suppose a particular disease afflicts 1% of the population. Suppose that a test for the disease is 95% accurate. Suppose that someone tests positive for the disease but there is no other evidence that they have the disease. What is the probability that they have the disease?
  
−
Let A be the event that the test result is positive. Let B be the event that the person actually has the disease.
+
Let X be the event that the test result is positive. Let Y be the event that the person actually has the disease.
  
−
Before the test result is known our probability for the person having the disease is p(B) = 1% = 0.01. The probability that the person has a positive test result, giving that they have the disease, is 95% = 0.95. The denominator term, p(A) is a little more complex since A can occur in two different ways. If the person has the disease, they will test positive with probability 0.95. If the person does not have the disease they will test positive with probability 5%. We denote the probability that B is not true by B'. The laws of probability require that B' = 1 - B. Thus  
+
Before the test result is known our probability for the person having the disease is p(Y) = 1% = 0.01. The probability that the person has a positive test result, giving that they have the disease, is 95% = 0.95. The denominator term, p(X) is a little more complex since X can occur in two different ways. If the person has the disease, they will test positive with probability 0.95. If the person does not have the disease they will test positive with probability 5%. We denote the probability that Y is not true by Y'. The laws of probability require that Y' = 1 - Y. Thus  
−
 
+
<math> P(X|Y,I)=\frac{P(Y|X,I)P(X|I)}{P(Y|I)} = (0.01)(0.95) + (0.99)(0.05) = 0.059 </math>
−
p(A) = p(B)p(A|B) + p(B')p(A|B') = 0.01 x 0.95 + 0.99 x 0.05 = 0.059  
 
  
 
Hence  
 
Hence  
  
−
p(B|A) 0.01 x 0.95/0.059 = 0.16
+
<math>\frac{P(X|Y)(0.01)(0.95)}{0.059} = 0.16</math>
  
 
In other words, there is only a 16% chance that a person testing positive actually has the disease.
 
In other words, there is only a 16% chance that a person testing positive actually has the disease.
  
−
An extended form of Bayes's theorem is obtained by noting that it applies to [[probability distribution]]s as well as to events. Let ''y'' be a (vector valued) observable quantity that we want to use to estimate some unknown, unobservable (vector valued) quantity <math>\theta</math>. Prior to seeing the data ''y'', we summarise our knowledge about <math>\theta</math> by a probability distribution <math>p(\theta)</math>. Assume that we have a model of the relationship between ''y'' and <math>\theta</math>. Call this <math>p(y|\theta)</math>. We can use Bayes' theorem to update our knowledge of <math>\theta</math> by incorporating the information contained in the observed data ''y''.
+
An extended form of Bayes's theorem is obtained by noting that it applies to [[probability distribution]]s as well as to events. Let ''y'' be a (vector valued) observable quantity that we want to use to estimate some unknown, unobservable (vector valued) quantity <math>\theta</math>. Prior to seeing the data ''y'', we summarize our knowledge about <math>\theta</math> by a probability distribution <math>p(\theta)</math>. Assume that we have a model of the relationship between ''y'' and <math>\theta</math>. Call this <math>p(y|\theta)</math>. We can use Bayes' theorem to update our knowledge of <math>\theta</math> by incorporating the information contained in the observed data ''y''.
  
 
We have                         
 
We have                         
  
−
               <math>p(\theta|y) = p(\theta)p(y|\theta)/p(y)</math>
+
               <math>p(\theta|y) = \frac{p(\theta)p(y|\theta)}{p(y)}</math>
 +
 
 +
 
 
[[category:Probability and Statistics]]
 
[[category:Probability and Statistics]]

Revision as of 19:21, February 8, 2010

Bayes' Theorem, also known as Bayes' Rule, is used in statistics and probability theory to relate marginal probabilities and conditional probabilities. In the context of Bayesian probability theory, it is used to update degrees of belief (probabilities) given new information. The theorem is attributed to the Reverend Thomas Bayes (1702-1761) an English nonconformist minister, but it is likely to have been known to earlier mathematicians such as Pierre Laplace.


Theorem

Bayes' theorem is stated mathematically as

<math> P(X|Y,I)=\frac{P(Y|X,I)P(X|I)}{P(Y|I)} </math>,

where X and Y are independent statements, I is available background information, and

<math>P(X|Y,I)</math> is the posterior probability for X given Y and I,

<math>P(Y|X,I)</math> is the likelihood for Y given X and I,

<math>P(X|I)</math> is the prior probability for X given only I, and

<math>P(Y|I)</math> is sometimes called the evidence or probability for Y given only I.

In a scientific context, X may be a hypothesis, and Y may be experimental data. The theorem can then be used to determine the degree of belief in the hypothesis by using the experimental data.


Applications and Examples

One popular example of the use of Bayes' theorem is the Monty Hall problem, inspired by the television show Let's Make a Deal.

Another example of how Bayes' theorem would be used is:

Suppose a particular disease afflicts 1% of the population. Suppose that a test for the disease is 95% accurate. Suppose that someone tests positive for the disease but there is no other evidence that they have the disease. What is the probability that they have the disease?

Let X be the event that the test result is positive. Let Y be the event that the person actually has the disease.

Before the test result is known our probability for the person having the disease is p(Y) = 1% = 0.01. The probability that the person has a positive test result, giving that they have the disease, is 95% = 0.95. The denominator term, p(X) is a little more complex since X can occur in two different ways. If the person has the disease, they will test positive with probability 0.95. If the person does not have the disease they will test positive with probability 5%. We denote the probability that Y is not true by Y'. The laws of probability require that Y' = 1 - Y. Thus <math> P(X|Y,I)=\frac{P(Y|X,I)P(X|I)}{P(Y|I)} = (0.01)(0.95) + (0.99)(0.05) = 0.059 </math>

Hence

<math>\frac{P(X|Y)(0.01)(0.95)}{0.059} = 0.16</math>

In other words, there is only a 16% chance that a person testing positive actually has the disease.

An extended form of Bayes's theorem is obtained by noting that it applies to probability distributions as well as to events. Let y be a (vector valued) observable quantity that we want to use to estimate some unknown, unobservable (vector valued) quantity <math>\theta</math>. Prior to seeing the data y, we summarize our knowledge about <math>\theta</math> by a probability distribution <math>p(\theta)</math>. Assume that we have a model of the relationship between y and <math>\theta</math>. Call this <math>p(y|\theta)</math>. We can use Bayes' theorem to update our knowledge of <math>\theta</math> by incorporating the information contained in the observed data y.

We have

             <math>p(\theta|y) = \frac{p(\theta)p(y|\theta)}{p(y)}</math>