Recall the Beta-Binomial model,
Further recall that the Beta-Binomial model is from a conjugate family.
Now, we will learn about the Gamma-Poisson, another conjugate family.
Suppose we are now interested in modeling the number of spam calls we receive.
We take a guess and say that the value of \lambda that is most likely is around 5,
Why can’t we use the Beta distribution as our prior distribution?
Why can’t we use the binomial distribution as our data distribution?
Y | \lambda \sim \text{Pois}(\lambda),
f(y|\lambda) = \frac{\lambda^y e^{-\lambda}}{y!}, \ \ \ y \in \{0,1, 2, ... \}
\lambda \sim \text{Gamma}(s, r)
f(\lambda) = \frac{r^s}{\Gamma(s)} \lambda^{s-1} e^{-r\lambda}
Let’s now tune our prior assuming \lambda \approx 5.
We know the mean of the gamma distribution,
E(\lambda) = \frac{s}{r} \approx 5 \to 5r \approx s
plot_gamma() function to figure out what value of s and r we need.Y_i|\lambda \sim \text{Pois}(\lambda)
f(y_i|\lambda) = \frac{\lambda^{y_i} e^{-\lambda}}{y_i!}
\begin{align*} f\left(\overset{\to}{y_i}|\lambda\right) &= \prod_{i=1}^n f(y_i|\lambda) \\ &= f(y_1|\lambda) \times f(y_2|\lambda) \times ... \times f(y_n|\lambda) \\ &= \frac{\lambda^{y_1}e^{-\lambda}}{y_1!} \times \frac{\lambda^{y_2}e^{-\lambda}}{y_2!} \times ... \times \frac{\lambda^{y_n}e^{-\lambda}}{y_n!} \\ &= \frac{\left( \lambda^{y_1} \lambda^{y_2} \cdot \cdot \cdot \ \lambda^{y_n} \right) \left( e^{-\lambda} e^{-\lambda} \cdot \cdot \cdot e^{-\lambda}\right)}{y_1! y_2! \cdot \cdot \cdot y_n!} \\ &= \frac{\lambda^{\sum y_i}e^{-n\lambda}}{\prod_{i=1}^n y_i !} \end{align*}
f\left(\overset{\to}{y_i}|\lambda\right) = \frac{\lambda^{\sum y_i}e^{-n\lambda}}{\prod_{i=1}^n y_i !}
\begin{align*} L\left(\lambda|\overset{\to}{y_i}\right) &= \frac{\lambda^{\sum y_i}e^{-n\lambda}}{\prod_{i=1}^n y_i !} \\ & \propto \lambda^{\sum y_i} e^{-n\lambda} \end{align*}
Let \lambda > 0 be an unknown rate parameter and (Y_1, Y_2, ... , Y_n) be an independent sample from the Poisson distribution.
The Gamma-Poisson Bayesian model is as follows:
\begin{align*} Y_i | \lambda &\overset{\text{ind}}\sim \text{Poisson}(\lambda) \\ \lambda &\sim \text{Gamma}(s, r) \\ \lambda | \overset{\to}y &\sim \text{Gamma}\left( s + \sum y_i, r + n \right) \end{align*}
Suppose we use Gamma(10, 2) as the prior for \lambda, the daily rate of calls.
On four separate days in the second week of August (i.e., independent days), we received \overset{\to}y = (6, 2, 2, 1) calls.
We know our prior distribution is Gamma(10, 2) and the data distribution is \text{Poi}(\lambda), with observed sample mean \bar{y} = 2.75.
Thus, the posterior is as follows,
\begin{align*} \lambda | \overset{\to}y &\sim \text{Gamma}\left( s + \sum y_i, r + n \right) \\ &\sim \text{Gamma}\left(10 + 11, 2 + 4 \right) \\ &\sim \text{Gamma}\left(21, 6 \right) \end{align*}
plot_gamma_poisson() function:summarize_gamma_poisson() function:Your turn! What is different if we had used one of the other priors?
Recall, we considered
| model | shape | rate | mean | mode | var | sd |
|---|---|---|---|---|---|---|
| prior | 5 | 1 | 5.0 | 4 | 5.00 | 2.236068 |
| posterior | 16 | 5 | 3.2 | 3 | 0.64 | 0.800000 |
| model | shape | rate | mean | mode | var | sd |
|---|---|---|---|---|---|---|
| prior | 10 | 2 | 5.0 | 4.500000 | 2.5000000 | 1.5811388 |
| posterior | 21 | 6 | 3.5 | 3.333333 | 0.5833333 | 0.7637626 |
| model | shape | rate | mean | mode | var | sd |
|---|---|---|---|---|---|---|
| prior | 15 | 3 | 5.000000 | 4.666667 | 1.6666667 | 1.2909944 |
| posterior | 26 | 7 | 3.714286 | 3.571429 | 0.5306122 | 0.7284314 |
| model | shape | rate | mean | mode | var | sd |
|---|---|---|---|---|---|---|
| prior | 20 | 4 | 5.000 | 4.75 | 1.250000 | 1.1180340 |
| posterior | 31 | 8 | 3.875 | 3.75 | 0.484375 | 0.6959705 |
\begin{equation*} \begin{aligned} Y|\lambda &\sim \text{Poi}(\lambda) \\ \lambda &\sim \text{Gamma}(s, r) & \end{aligned} \Rightarrow \begin{aligned} && \lambda | y &\sim \text{Gamma}(s + \sum y_i, r + n) \\ \end{aligned} \end{equation*}
The prior model, f(\lambda), is given by Gamma(s, r).
The data model, f(Y|\lambda), is given by Poi(\lambda).
The posterior model is a Gamma distribution with updated parameters s+\sum y_i and r + n.
As scientists learn more about brain health, the dangers of concussions are gaining greater attention.
We are interested in \mu, the average volume (cm3) of a specific part of the brain: the hippocampus.
Wikipedia tells us that among the general population of human adults, each half of the hippocampus has volume between 3.0 and 3.5 cm3.
We will take a sample of n=25 participants and update our belief.
f(y) = \frac{1}{\sqrt{2\pi\sigma^2}} \exp \left\{ \frac{-(y-\mu)^2}{2\sigma^2} \right\}
Y_i | \mu \sim N(\mu, \sigma^2)
f(\overset{\to}y | \mu) = \prod_{i=1}^n f(y_i | \mu) = \prod_{i=1}^n \frac{1}{\sqrt{2 \pi \sigma^2}} \exp \left\{ \frac{-(y_i-\mu)^2}{2\sigma^2} \right\}
L(\mu|\overset{\to}y) \propto \prod_{i=1}^n \frac{1}{\sqrt{2 \pi \sigma^2}} \exp \left\{ \frac{-(y_i-\mu)^2}{2\sigma^2} \right\} = \exp \left\{ \frac{- \sum_{i=1}^n(y_i-\mu)^2}{2\sigma^2} \right\}
Y_i | \mu \sim N(\mu, \sigma^2)
\mu \sim N(\theta, \tau^2)
f(\mu) = \frac{1}{\sqrt{2 \pi \tau^2}} \exp \left\{ \frac{-(\mu - \theta)^2}{2 \tau^2} \right\}
We can tune the hyperparameters \theta and \tau to reflect our understanding and uncertainty about the average hippocampal volume (\mu) among people with a history of concussions.
Wikipedia showed us that hippocampal volumes tend to be between 6 and 7 cm3 \to \theta=6.5.
When we set the standard deviation we can check the plausible range of values of \mu:
\theta \pm 2 \times \tau
(6.5 \pm 2 \times 0.4) = (5.7, 7.3)
Let \mu \in \mathbb{R} be an unknown mean parameter and (Y_1, Y_2, ..., Y_n) be an independent N(\mu, \sigma^2) sample where \sigma is assumed to be known.
The Normal-Normal Bayesian model is as follows:
\begin{align*} Y_i | \mu &\overset{\text{iid}} \sim N(\mu, \sigma^2) \\ \mu &\sim N(\theta, \tau^2) \\ \mu | \overset{\to}y &\sim N\left( \theta \frac{\sigma^2}{n\tau^2 + \sigma^2} + \bar{y} \frac{n\tau^2}{n\tau^2 + \sigma^2}, \frac{\tau^2 \sigma^2}{n \tau^2 + \sigma^2} \right) \end{align*}
\mu | \overset{\to}y \sim N\left( \theta \frac{\sigma^2}{n\tau^2 + \sigma^2} + \bar{y} \frac{n\tau^2}{n\tau^2 + \sigma^2}, \frac{\tau^2 \sigma^2}{n \tau^2 + \sigma^2} \right)
\mu | \overset{\to}y \sim N\left( \theta \frac{\sigma^2}{n\tau^2 + \sigma^2} + \bar{y} \frac{n\tau^2}{n\tau^2 + \sigma^2}, \frac{\tau^2 \sigma^2}{n \tau^2 + \sigma^2} \right)
\begin{align*} \frac{\sigma^2}{n\tau^2 + \sigma^2} &\to 0 \\ \frac{n\tau^2}{n\tau^2 + \sigma^2} &\to 1 \\ \frac{\tau^2 \sigma^2}{n \tau^2 + \sigma^2} &\to 0 \end{align*}
\begin{align*} \frac{\sigma^2}{n\tau^2 + \sigma^2} &\to 0 \\ \frac{n\tau^2}{n\tau^2 + \sigma^2} &\to 1 \\ \frac{\tau^2 \sigma^2}{n \tau^2 + \sigma^2} &\to 0 \end{align*}
The posterior mean places less weight on the prior mean and more weight on the sample mean \bar{y}.
The posterior certainty about \mu increases and becomes more in sync with the data.
plot_normal_normal() function:summarize_normal_normal() function:Let us now apply this to our example.
We have our prior model, \mu \sim N(6.5, 0.4^2).
Let’s look at the football dataset in the bayesrules package.
Let us now apply this to our example.
We have our prior model, \mu \sim N(6.5, 0.4^2).
Let’s look at the football dataset in the bayesrules package.
L(\mu|\overset{\to}y) \propto \exp \left\{ \frac{-(5.735 - \mu)^2}{2(0.5^2/25)} \right\}
\mu | \overset{\to}y \sim N\left( \theta \frac{\sigma^2}{n\tau^2 + \sigma^2} + \bar{y} \frac{n\tau^2}{n\tau^2 + \sigma^2}, \frac{\tau^2 \sigma^2}{n \tau^2 + \sigma^2} \right)
\begin{align*} \mu | \overset{\to}y &\sim N\left( \theta \frac{\sigma^2}{n\tau^2 + \sigma^2} + \bar{y} \frac{n\tau^2}{n\tau^2 + \sigma^2}, \frac{\tau^2 \sigma^2}{n \tau^2 + \sigma^2} \right) \\ &\sim N\left( 6.5 \frac{0.5^2}{25 \cdot 0.4^2 + 0.5^2} + 5.735 \frac{25 \cdot 0.4^2}{25 \cdot 0.4^2 + 0.5^2}, \frac{0.4^2 \cdot 0.5^2}{25 \cdot 0.4^2 + 0.5^2} \right) \\ &\sim N(6.5 \cdot 0.0588 + 5.735 \cdot 0.9412, 0.09^2) \\ &\sim N(5.78, 0.09^2) \end{align*}
summarize_normal_normal() function to summarize the distribution,\begin{equation*} \begin{aligned} Y_i | \mu &\overset{\text{iid}} \sim N(\mu, \sigma^2) \\ \mu &\sim N(\theta, \tau^2) & \end{aligned} \Rightarrow \begin{aligned} && \mu | \overset{\to}y &\sim N\left( \theta \frac{\sigma^2}{n\tau^2 + \sigma^2} + \bar{y} \frac{n\tau^2}{n\tau^2 + \sigma^2}, \frac{\tau^2 \sigma^2}{n \tau^2 + \sigma^2} \right) \\ \end{aligned} \end{equation*}
The prior model, f(\mu), is given by N(\theta,\tau^2).
The data model, f(Y|\mu), is given by N(\mu, \sigma^2).
The posterior model is a Normal distribution with updated parameters
STA6349 · Applied Bayesian Analysis · Fall 2026