Homework 4

Due: Tuesday October 6 at 3 PM

On this homework you’ll get some practice with our last two regression topics: variable selection and logistic regression.

Setup

  1. Log-in to your container;
  2. Double-check that you have the sta101-f26-files project loaded in the upper-right corner of RStudio (this should always be true);
  3. Go to the “Git” tab in the upper-right panel of RStudio and click the Pull button (arrow pointing down). Now the new assignment template and dataset should appear in the hw folder of your Files;
  4. At the top of the new .qmd, in between ---, you have the settings for the document (the so-called YAML, but don’t worry about that). Modify the author so that it lists lil’ ol’ you.

Exercise 0

Recommend some music for us to listen to while we grade this.

Part 1: model selection

Exercise 1

Let’s revisit the weight lifting data from Homework 2:

ipf <- read_csv("data/ipf.csv")
  1. Don’t worry, I’m not going to make you convert to pounds. However, we do need to remove some rows:
    • Only keep the rows for which best3deadlift_kg, best3squat_kg, and bodyweight_kg are all positive.
    • Today we will focus specifically on the results for male lifters (it’s just the only way I could get the numbers to work out);
    • Make sure you perform those filters and then store the result in a data frame with a new name.
  2. We are considering predicting best3deadlift_kg using best3squat_kg, bodyweight_kg, and age. So we have \(p=3\) candidate predictors, and we want to know which of these variables to include in an additive model. How many additive models could we construct from this set of predictors, and which model is best? Explain your reasoning;
  3. In this example, explain step-by-step how a forward stepwise procedure would search for the best model. Which model would be selected?
  4. In this example, explain step-by-step how a backward elimination procedure would search for the best model. Which model would be selected?
  5. Did you reach the same model in each of the last three parts? If not, why didn’t the selection procedures agree? If so, explain within the context of this application why the final model choice makes sense.

Part 2: logistic regression

The General Social Survey (GSS) gathers data on contemporary American society in order to monitor and explain trends and constants in attitudes, behaviours, and attributes. Hundreds of trends have been tracked since 1972. In addition, since the GSS adopted questions from earlier surveys, trends can be followed for up to 70 years. The GSS contains a standard core of demographic, behavioural, and attitudinal questions, plus topics of special interest. Among the topics covered are civil liberties, crime and violence, intergroup tolerance, morality, national spending priorities, psychological well-being, social mobility, and stress and traumatic events.

In this part you will work with the 2022 General Social Survey:

gss22 <- read_csv("data/gss22.csv")

We will focus on these variables:

  • advfront: does the respondent agree or not with “Even if it brings no immediate benefits, scientific research that advances the frontiers of knowledge is necessary and should be supported by the federal government;”
  • educ: years of education;
  • polviews: is the respondent “Conservative,” “Moderate,” or “Liberal.”

Exercise 2 (data prep)

  1. After reading in the data, make the categorical variables advfront and polviews factors and make the order of their levels “Agree” - “Not Agree” and “Conservative” - “Moderate” - “Liberal,” respectively;
  2. Split the data into a training set (75%) and a testing set (25%), and use the random number seed 8675309 for reproducibility.

Exercise 3 (a simple model)

  1. Fit a logistic regression model that predicts advfront from educ. Report the tidy output of the model;
  2. Write out the estimated model in proper notation;
  3. Interpret all coefficient estimates;
  4. Using your estimated model, predict the probability that someone with 7 years of education agrees with the statement.

Exercise 4 (model selection)

  1. Fit a second model that adds the additional explanatory variable of polviews to your model from the previous question. Report the tidy output of the model;
  2. Write out the estimated model in proper notation;
  3. You have two logistic regression models that are both attempting to predict the same response: advfront. Which one is best? Explain your reasoning and provide evidence for your conclusion.

Part 3: IMS exercises

These exercises from the textbook do not require code, but you should type your responses into the same Quarto file you’ve been using. Make sure to answer the questions in full sentences.

Exercise 5

IMS - Chapter 8 exercises, #12: Palmer penguins, predicting body mass.

Exercise 6

IMS - Chapter 8 exercises, #14: Palmer penguins, backwards elimination.

Submission

  1. Hit the blue Render button to generate your final PDF;
  2. Give your work a final look over to double-check a few things:
    • that your code is stylish;
    • that none of your code or pictures runs off the page. We cannot grade what we cannot read;
    • that all of your plots are well-labeled and human-readable. In other words “Flipper length (mm)” instead of flipper_length_mm;
    • If your work is lacking on any of these items, fix them and re-render as needed;
  3. Download the PDF from your container;
  4. Upload it to Gradescope;
  5. Don’t forget to mark your pages.