AE 08: Scraping articles from the Cornell Review

Application exercise
Modified

September 29, 2026

Packages

We will use the following packages in this application exercise.

  • {tidyverse}: For data import, wrangling, and visualization.
  • {rvest}: For scraping HTML files.
  • {robotstxt}: For verifying if we can scrape a website.

Data scraping

This will be done in the scrape-cornell-review.R R script. Save the resulting data frame in the data folder.

# load packages
library(tidyverse)
library(rvest)
library(robotstxt)

# check that we can scrape data from the cornell review
paths_allowed("https://www.thecornellreview.org/")

# read the first page
page <- read_html("https://www.thecornellreview.org/")
# page <- read_html("data/cornell-review-raw.html") # use this if we break the website

# extract desired components
titles <- html_elements(x = page, css = "______") |>
  html_text2()

# text field cannot separate dates from authors - have to do it manually
dates_authors <- html_elements(x = page, css = "______") |>
  html_text2()

topics <- html_elements(x = page, css = "______") |>
  html_text2()

post_urls <- html_elements(x = page, css = "______") |>
  html_attr(name = "href")

# extract dates and authors as separate vectors
dates <- str_split_i(string = dates_authors, pattern = "______", i = 1)

authors <- str_split_i(string = dates_authors, pattern = "______", i = 2) |>
  # remove leading and trailing whitespace
  str_trim()

# create a tibble with this data
review <- tibble(______)

# save to disk
write_csv(x = review, file = "data/cornell-review.csv")