Thinking Like a Bayesian
(Part 1)

Introduction

  • On Monday:
    • Refresher on probability theory
    • What does each distribution do?
      • beta \to outcomes limited to [0, 1]
      • binomial \to binary outcomes
      • gamma \to continuous & positive outcomes; skewed right
      • normal \to continuous outcomes; mound-shaped & symmetric distribution
      • Poisson \to count outcomes; skewed right
      • uniform \to each outcome has equal probability; rectangular distribution
  • Today: building up Bayesian analysis concepts

R Set Up

  • To follow today’s lecture, please load the following packages:
library(tidyverse)
library(janitor)
library(bayesrules)
  • If you need to install any, you can do so with the install.packages() function.
install.packages("bayesrules")

Thinking Like a Bayesian

  • Bayesian analysis involves updating beliefs based on observed data.

Thinking Like a Bayesian

  • Bayesian analysis involves updating beliefs based on observed data.

Thinking Like a Bayesian

  • Bayesian analysis involves updating beliefs based on observed data.

Thinking Like a Bayesian

  • Bayesian analysis involves updating beliefs based on observed data.

Thinking Like a Bayesian

  • This is the natural Bayesian knowledge-building process of:
    • acknowledging your preconceptions (prior distribution),

Thinking Like a Bayesian

  • This is the natural Bayesian knowledge-building process of:
    • acknowledging your preconceptions (prior distribution),
    • using data (data distribution) to update your knowledge (posterior distribution),

Thinking Like a Bayesian

  • This is the natural Bayesian knowledge-building process of:
    • acknowledging your preconceptions (prior distribution),
    • using data (data distribution) to update your knowledge (posterior distribution), and
    • repeating (posterior distribution \to new prior distribution)

Thinking Like a Bayesian

  • Bayesian and frequentist analyses share a common goal: to learn from data about the world around us.
    • Both Bayesian and frequentist analyses use data to fit models, make predictions, and evaluate hypotheses.
    • When working with the same data, they will typically produce a similar set of conclusions.
  • Statisticians typically identify as either a “Bayesian” or “frequentist” …
    • We are not going to “take sides.”
    • We will see these as tools in our toolbox.

Thinking Like a Bayesian

  • Bayesian probability: the relative plausibility of an event. This considers prior belief.

Thinking Like a Bayesian

  • Frequentist probability: the long-run relative frequency of a repeatable event. This does not consider prior belief.

Thinking Like a Bayesian

  • The Bayesian framework depends upon prior information, data, and the balance between them.

    • The balance between the prior information and data is determined by the relative strength of each.
  • When we have little data, our posterior can rely more on prior knowledge.

  • As we collect more data, the prior can lose its influence.

Thinking Like a Bayesian

  • We can also use this approach to combine analysis results.

Thinking Like a Bayesian

  • We will use an example to work through Bayesian logic.

  • The Collins Dictionary named “fake news” the 2017 term of the year.

    • Fake, misleading, and biased news has proliferated along with online news and social media platforms which allow users to post articles with little quality control.
  • We want to flag articles as “real” or “fake.”

  • We’ll examine a sample of 150 articles which were posted on Facebook and fact checked by five BuzzFeed journalists (Shu et al. 2017).

Thinking Like a Bayesian

  • Information about each article is stored in the fake_news dataset in the bayesrules package.
fake_news <- bayesrules::fake_news
print(colnames(fake_news))
 [1] "title"                   "text"                   
 [3] "url"                     "authors"                
 [5] "type"                    "title_words"            
 [7] "text_words"              "title_char"             
 [9] "text_char"               "title_caps"             
[11] "text_caps"               "title_caps_percent"     
[13] "text_caps_percent"       "title_excl"             
[15] "text_excl"               "title_excl_percent"     
[17] "text_excl_percent"       "title_has_excl"         
[19] "anger"                   "anticipation"           
[21] "disgust"                 "fear"                   
[23] "joy"                     "sadness"                
[25] "surprise"                "trust"                  
[27] "negative"                "positive"               
[29] "text_syllables"          "text_syllables_per_word"

Thinking Like a Bayesian

  • Suppose that the most recent article posted to a social media platform is titled: The president has a funny secret!
    • Some features of this title probably set off some red flags.
    • For example, the usage of an exclamation point might seem like an odd choice for a real news article.
  • In the dataset, what is the split of real and fake articles?
fake_news %>% tabyl(type)

Thinking Like a Bayesian

  • In the dataset, what is the split of real and fake articles?
fake_news %>% tabyl(type)
type n percent
fake 60 0.4
real 90 0.6

Thinking Like a Bayesian

  • What does the data say?
fake_news %>% tabyl(title_has_excl, type) %>%
  adorn_percentages("col") %>%
  adorn_pct_formatting(digits = 2) %>%
  adorn_ns()

Thinking Like a Bayesian

  • What does the data say?
fake_news %>% tabyl(title_has_excl, type) %>%
  adorn_percentages("col") %>%
  adorn_pct_formatting(digits = 2) %>%
  adorn_ns()
title_has_excl fake real
FALSE 73.33% (44) 97.78% (88)
TRUE 26.67% (16) 2.22% (2)
  • 26.67% of fake news titles use ! vs 2.22% of real news titles use !

Thinking Like a Bayesian

  • We now have two pieces of contradictory information.
    • Our prior information suggested that incoming articles are most likely real.
    • However, the exclamation point data is more consistent with fake news.

Building a Bayesian Model

  • Thinking like Bayesians, we know that balancing both pieces of information is important in developing a posterior understanding of whether the article is fake.

  • Our fake news analysis studies two variables:

    • an article’s fake vs real status and
    • its use of exclamation points.
  • We can represent the randomness in these variables using probability models.

Building a Bayesian Model

  • Let’s now formalize our prior understanding of whether the new article is fake.

  • Based on our data, we saw that 40% of articles are fake and 60% are real.

    • Before reading the new article, there’s a 0.4 prior probability that it’s fake.

P\left[B\right] = 0.40 \text{ and } P\left[B^c\right] = 0.60

  • Remember that a valid probability model must:

    1. account for all possible events
    2. assign prior probabilities to each event
    3. have probabilities that sum to 1

Building a Bayesian Model

  • Revisiting the contingency table for conditional probabilities,
title_has_excl fake real
FALSE 73.3% (44) 97.8% (88)
TRUE 26.7% (16) 2.2% (2)
  • If an article is fake (B), what is the probability that it uses exclamation points in the title?



  • If an article is real (B^c), what is the probability that it uses exclamation points in the title?

Building a Bayesian Model

  • We have the following conditional probabilities:
    • If an article is fake, then there is about a 27% chance it uses exclamation points in the title.
      • P[A|B]=0.2667
    • If an article is real, then there is only a 2% chance it uses exclamation points in the title.
      • P[A|B^c]=0.0222
  • Exclamation point usage is much more likely among fake news than real news.
    • We have evidence that the article is fake.

Building a Bayesian Model

  • Note that we know that the incoming article used exclamation points (A), but we do not actually know if the article is fake (B or B^c).

  • In this case, we compared P[A|B] and P[A|B^c] to ascertain the relative likelihood of observing A under different scenarios. That is,

    • L\left[B|A\right] = P\left[A|B\right] \to the likelihood of fake given !
    • L\left[B^c|A\right] = P\left[A|B^c\right] \to the likelihood of real given !

Building a Bayesian Model

Event B Bc Total
Prior Probability 0.4 0.6 1.0
Likelihood 0.2667 0.0222 0.2889

L\left[B|A\right] = P\left[A|B\right] \text{ and } L\left[B^c|A\right] = P\left[A|B^c\right]

  • It is important for us to note that the likelihood function is not a probability function.
    • This is a framework to determine the relative compatibility of our exclamation point data with B and B^c.

Building a Bayesian Model

Event B (fake) Bc (real) Total
Prior Probability 0.4 0.6 1.0
Likelihood 0.2667 0.0222 0.2889
  • The prior evidence suggested the article is most likely real,

P[B] = 0.4 < P[B^c] = 0.6

  • The data, however, is more consistent with the article being fake,

L[B|A] = 0.2667 > L[B^c|A] = 0.0222

Building a Bayesian Model

  • We can summarize our probabilities in a table,
B B^c Total
A
A^c
Total 0.4 0.6 1
  • As found earlier, P[A|B] = 0.2667 and P[A|B^c]=0.0222.

  • Find P[A \cap B].

Building a Bayesian Model

  • We can summarize our probabilities in a table,
B B^c Total
A
A^c
Total 0.4 0.6 1
  • As found earlier, P[A|B] = 0.2667 and P[A|B^c]=0.0222.

\begin{align*} P[A \cap B] &= P[A|B] \times P[B] \\ &= 0.2667 \times 0.4 \\ &= 0.1067 \end{align*}

Building a Bayesian Model

  • Here’s what we know,
B B^c Total
A 0.1067
A^c
Total 0.4 0.6 1
  • As found earlier, P[A|B] = 0.2667 and P[A|B^c]=0.0222.

  • Find P[A^c \cap B].

Building a Bayesian Model

  • Here’s what we know,
B B^c Total
A 0.1067
A^c
Total 0.4 0.6 1
  • As found earlier, P[A|B] = 0.2667 and P[A|B^c]=0.0222.

\begin{align*} P[A^c \cap B] &= P[A^c|B] \times P[B] \\ &= \left(1-P[A|B]\right) \times P[B] \\ &= (1-0.2667) \times 0.4 \\ &= 0.2933 \end{align*}

Building a Bayesian Model

  • Here’s what we know,
B B^c Total
A 0.1067
A^c 0.2933
Total 0.4 0.6 1
  • As found earlier, P[A|B] = 0.2667 and P[A|B^c]=0.0222.

  • Find P[A \cap B^c].

Building a Bayesian Model

  • Here’s what we know,
B B^c Total
A 0.1067
A^c 0.2933
Total 0.4 0.6 1
  • As found earlier, P[A|B] = 0.2667 and P[A|B^c]=0.0222.

\begin{align*} P[A \cap B^c] &= P[A|B^c] \times P[B^c] \\ &= 0.0222 \times 0.6 \\ &= 0.0133 \end{align*}

Building a Bayesian Model

  • Here’s what we know,
B B^c Total
A 0.1067 0.0133
A^c 0.2933
Total 0.4 0.6 1
  • As found earlier, P[A|B] = 0.2667 and P[A|B^c]=0.0222.

  • Find P[A^c \cap B^c].

Building a Bayesian Model

  • Here’s what we know,
B B^c Total
A 0.1067 0.0133
A^c 0.2933
Total 0.4 0.6 1
  • As found earlier, P[A|B] = 0.2667 and P[A|B^c]=0.0222.

\begin{align*} P[A^c \cap B^c] &= P[A^c|B^c] \times P[B^c] \\ &= 0.9778 \times 0.6 \\ &= 0.5867 \end{align*}

Building a Bayesian Model

  • Here’s what we know,
B B^c Total
A 0.1067 0.0133
A^c 0.2933 0.5867
Total 0.4 0.6 1
  • Finally,

\begin{align*} &P[A] = 0.1067 + 0.0133 = 0.12 \\ &P[A^c] = 0.2933 + 0.5867 = 0.88 \end{align*}

Building a Bayesian Model

  • Using rules of probability, we have completed the table.
B B^c Total
A 0.1067 0.0133 0.12
A^c 0.2933 0.5867 0.88
Total 0.4 0.6 1

Building a Bayesian Model

  • With more information, we can finally find the probability that the latest article is fake.

  • We will use the posterior probability, P[B|A], which is found using Bayes’ Rule.

  • Bayes’ Rule: For events A and B,

P[B|A] = \frac{P[A \cap B]}{P[A]} = \frac{P[B] \times L[B|A]}{P[A]}

  • But really, we can think about it like this,

\text{posterior} = \frac{\text{prior} \times \text{likelihood}}{\text{normalizing constant}}

Building a Bayesian Model

  • Let’s now apply Bayes’ Rule to our fake news example.
B B^c Total
A 0.1067 0.0133 0.12
A^c 0.2933 0.5867 0.88
Total 0.4 0.6 1
  • Find P[B|A].

Building a Bayesian Model

  • Applying Bayes’ Rule to our fake news example,

P[B|A] = \frac{P[A \cap B]}{P[A]} = \frac{0.1067}{0.12} = 0.889

  • There is an 88.9% posterior probability that the article is fake given that it uses an exclamation point in the title.

Building a Bayesian Model

  • Comparing our prior and posterior probabilities,
    • Prior: P[B] = 0.40, i.e., before seeing the title, 40% chance the article is fake.
    • Posterior: P[B|A] = 0.889, i.e., after seeing the exclamation point in the title, 88.9% chance the article is fake.
  • The exclamation point adds strong evidence to the article being fake.
    • Our posterior is much higher than our prior.

Wrap Up

  • Today we have introduced thinking like a Bayesian.

  • The biggest take away: we now consider prior data/beliefs, rather than just the collected data.

  • Next week: formalizing the Bayesian model with conjugate families.

    • Beta-Binomial

    • Gamma-Poisson

    • Normal-Normal