Lab 5

Due: end-of-lab Friday October 2

Our data today describe the demand of two different types of hotels. Each observation represents a hotel booking between July 1, 2015 and August 31, 2017. Some bookings were cancelled and others were kept (ie. the guest checked in).

hotels <- read_csv("data/hotels.csv")

The variables in the dataset and their descriptions are as follows:

Below is an exploratory summary of the variables in the dataset we will use in our analysis.

Cancellation status Number of bookings
0 - Not canceled 37598
1 - Canceled 22097

 

Number of adults Number of bookings
0 187
1 11622
2 44757
3 3089
4 31
5 1
6 1
20 1
26 2
27 2
40 1
50 1

Setup

  1. Log-in to your container;
  2. Double-check that you have the sta101-f26-files project loaded in the upper-right corner of RStudio (this should always be true);
  3. Go to the “Git” tab in the upper-right panel of RStudio and click the Pull button (arrow pointing down). Now the new assignment template and dataset should appear in the lab folder of your Files;
  4. At the top of the new .qmd, in between ---, you have the settings for the document (the so-called YAML, but don’t worry about that). Modify the authors so it lists yourself and your teammates.

Part 1 - data prep

Question 1

  1. One of the outcome variables you’ll use is is_canceled. Transform this variable to the appropriate data class so that the levels are ordered such that when we fit a logistic regression model to predict this outcome, success is defined as “canceled” (what we’re predicting).

  2. Based on the exploratory summaries above, you should address a few data quality issues before moving forward with the analysis. To do so, filter the dataset to remove

    • any bookings with average daily rate greater than $1,000 and
    • any bookings with number of adults greater than or equal to 5.
  3. Split the data into a training set (75%) and a testing set (25%), setting the random seed to 1117 for reproducibility.

Part 2 - Predicting cancellations

Question 2

Using these data, one of our goals is to explore the following question:

Are reservations earlier in the month or later in the month more likely to be cancelled?

  1. In a single pipeline, calculate the mean arrival dates (arrival_date_day_of_month) for reservations that were cancelled and reservations that were not cancelled.

Think carefully about which dataset you should use: hotels, the training subset, or the testing subset?

  1. In your own words, explain why we can not use a linear model to model the relationship between if a hotel reservation was cancelled and the day of month for the booking.

  2. Fit the appropriate model to predict whether a reservation was cancelled from arrival_date_day_of_month and display a tidy summary of the model output. Then, interpret the slope coefficient in context of the data and the research question.

The slope interpretation will have the following format:

The model predicts that, for each day the booking is ___ (later / earlier) in the month, the ___ (chance / probability / odds) of a hotel cancellation is ___ (lower / higher) by a factor of ___, on average.

  1. Calculate the probability that the hotel reservation is cancelled if it the arrival date is on the 17th of the month. Based on this probability, would you predict this booking would be cancelled or not cancelled. Explain your reasoning for your classification.

Question 3

  1. Fit another model to predict whether a reservation was cancelled from arrival_date_day_of_month and hotel type (Resort or City Hotel), allowing the relationship between arrival_date_day_of_month and is_canceled to vary based on hotel type. Display a tidy output of the model.

  2. Interpret the intercept in context of the data.

Part 3 - Predicting daily rates

Question 4

The dataset also contains information about the average daily rate (adr) for each reservation. The following model predicts adr from adults and hotel type.

# A tibble: 3 × 5
  term              estimate std.error statistic   p.value
  <chr>                <dbl>     <dbl>     <dbl>     <dbl>
1 (Intercept)           51.7     0.853      60.5 0        
2 adults                29.0     0.439      66.1 0        
3 hotelResort Hotel    -10.8     0.455     -23.7 6.53e-124
  1. Which of the following is the best interpretation of the slope coefficient for adults?

    For each additional adult in the booking, the average daily rate is predicted to be higher by $29.00…

    1. on average, holding hotel type constant.
    2. for Resort Hotels compared to City Hotels, on average.
    3. for City Hotels compared to Resort Hotels, on average.
    4. on average, not holding any other variables constant.
  2. Which of the following is the correct interpretation of the slope coefficient for hotel?

    1. For each additional Resort Hotel booking, the predicted average daily rate is $10.80 lower, holding number of adults constant.
    2. For each additional adult in the booking, the average daily rate is predicted to be lower by $10.80 for resort hotels compared to City Hotels, on average.
    3. Resort Hotels bookings are predicted to have an average daily rate that is $10.80 lower than City Hotels, on average, holding number of adults constant.
    4. Resort Hotels bookings are predicted to have an average daily rate that is $10.80 higher than City Hotels, on average, holding number of adults constant.
  3. Which of the following is the correct interpretation of the intercept?

    1. The predicted average daily rate for a bookings with 0 adults at a Resort Hotel is $51.70, on average.
    2. The predicted average daily rate for a bookings with 0 adults at a City Hotel is $51.70, on average.
    3. For each additional adult and Resort Hotel in the booking, the average daily rate is predicted to be $51.70 higher, on average.
    4. For each additional adult and City Hotel in the booking, the average daily rate is predicted to be $51.70 higher, on average.
  4. Which of the following (Plot I or Plot II) is the correct visual representation of this model?

Part 4 - study

Congratulations! You have just completed all of the work you have to submit for today’s lab. If any time remains in the session, I encourage you to stick around and work on the midterm study guide with your teammates.

Submission

You collaborated with your team, but now everyone submits individually:

  1. Hit the blue Render button to generate your final PDF;
  2. Give your work a final look over to double-check a few things:
    • that your code is stylish;
    • that none of your code or pictures runs off the page. We cannot grade what we cannot read;
    • that all of your plots are well-labeled and human-readable. In other words “Flipper length (mm)” instead of flipper_length_mm;
    • If your work is lacking on any of these items, fix them and re-render as needed;
  3. Download the PDF from your container;
  4. Upload it to Gradescope;
  5. Don’t forget to mark your pages.