Lecture 0
Duke University
STA 101 Fall 2026
2026-08-25
Campaign manager: What share of the popular vote do we think our candidate will receive?
(A flurry of analysis takes place.)
Data scientist: Our best guess is 54%.
Campaign manager: How reliable is that estimate? How confident are we in that? What’s the margin of error?
Parallel Universe 1
Data scientist: It’s 54% give or take 3%.
Parallel Universe 2
Data scientist: It’s 54% give or take 20%.
It’s all about decision-making under uncertainty
The manager is going to make wildly different decisions about campaign spending and strategy depending on how uncertain the environment is. Therefore, it pays to know what you’re dealing with. Statistics can help!
Data analysis
Statistical inference
Quantifying our uncertainty about that knowledge.
Back in Spring 2014…





This is much closer to what contemporary practitioners actually do.
That’s okay!
I can’t guarantee where exactly on this grid the actual letter grades fall, but wherever you land, there is no judgement.
Your final course grade will be calculated as follows:
| Category | Percentage |
|---|---|
| Labs | 10% |
| Homeworks | 20% |
| Project | 20% |
| Midterm Exam | 25% |
| Final exam | 25% |
The final letter grade is based on the usual thresholds.
What you see is what you get
No curve, no rounding, no extra credit.
Hands-on practice with coding and data analysis;
Completed in-person, in lab, in randomly-assigned teams;
Developed collaboratively, but turned in individually by the end of the lab session;
10 throughout semester; 2 lowest scores dropped;
No late work accepted.
Traditional, in-class, no-tech exams:
You are only allowed both sides of one 8.5” x 11” note sheet created by you.
If you seek testing accommodations…
Make sure I get an SDAO letter, and make your appointments in the Testing Center now.
If you wish to ask questions in writing…
Post on Ed: about general course policies and content;
Email JZ directly: personal matters.
You should not really be emailing the TAs directly for any reason.
You are enthusiastically encouraged to work together on labs and homeworks. You will learn a lot from each other! Two policies:
Violation of the second policy is plagiarism. Sharers and recipients alike are referred to the conduct office and receive zeros.
Check out the discussion in the syllabus.
tl; dr:
They’re similar (101 and 199 use the same textbook), but:
Also, 198/199 counts as an elective for the stats major. 101 does not. 101 can only count towards the minor.
Ethel Merman

| Born | January 16, 1908 |
| Died | February 15, 1984 |
| Age | 76 |
| Claim to fame | JZ’s favorite singer |
Megan Pete

| Born | February 15, 1995 |
| Age | 31 |
| Claim to fame | Rapper |
봉준호

| Born | September 14, 1969 |
| Age | 56 |
| Claim to fame | Directed Parasite, Snowpiercer, etc |
When the picture was taken, how old was the person?
# A tibble: 82 × 6
celeb1 celeb2 celeb3 celeb4 celeb5 celeb6
<dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
1 36 63 78 79 42 58
2 18 25 72 34 26 64
3 31 35 61 83 58 68
4 26 32 64 50 40 65
5 34 56 67 52 41 71
6 40 40 84 64 42 80
7 34 45 80 40 40 70
8 35 62 64 71 53 82
9 20 35 77 45 37 45
10 45 64 86 65 48 67
# ℹ 72 more rows
Each row is a 101 student, and the columns record their guesses for each celeb.


| Born | 2/10/1987 |
| Age in pic | 36 |
| Claim to fame | Classical pianist |
What dataset do we want to visualize?
What part of it do we want to visualize?
What kind of plot do we want?
Change the horizontal axis range:
Add some human-readable labels:
Add a vertical line at the answer:
Where are your guesses concentrated? How spread out are they?
# A tibble: 1 × 2
average sdev
<dbl> <dbl>
1 31.2 6.34
Say hello to the mean and standard deviation.
|>.
Better known as Charo.

| Born | 3/13/1941? 1/15/1951? Who knows. |
| Age in pic | 58 - 68 |
| Claim to fame | Renaissance woman |

His actual birthday was not known at the time:

| Born | 2/7/1887 |
| Died | 2/12/1983 |
| Age in pic | 82 |
| Claim to fame | Composer |



| Born | 2/3/1963 |
| Age in pic | 48 |
| Claim to fame | UChicago economist |
| RBI governor |


| Born | 5/27/1911 |
| Died | 10/25/1993 |
| Age in pic | 38 - 39 |
| Claim to fame | Horror actor |


| Born | 10/21/1925 |
| Died | 7/16/2003 |
| Age in pic | 76 |
| Claim to fame | Queen of Salsa |
| Your name | celeb1 | celeb2 | celeb3 | celeb4 | celeb5 | celeb6 | sqerror |
|---|---|---|---|---|---|---|---|
| (answers) | 36 | 63 | 82 | 48 | 38.5 | 76 | 0.00000 |
| NA | 42 | 65 | 75 | 45 | 37.0 | 63 | 44.87500 |
| NA | 35 | 56 | 78 | 39 | 47.0 | 67 | 50.04167 |
| NA | 34 | 56 | 67 | 52 | 41.0 | 71 | 54.20833 |
| NA | 41 | 62 | 85 | 44 | 52.0 | 66 | 55.54167 |
| NA | 30 | 50 | 75 | 45 | 45.0 | 70 | 56.87500 |
| NA | 27 | 68 | 75 | 59 | 48.0 | 72 | 63.70833 |
(no doxxing version)


Domain knowledge and modeling assumptions: data do not speak for themselves. You need some subject-matter expertise about what you’re studying, as well as an interpretive lens;
Are you asking questions the data can actually answer?
Uncertainty has many sources, and in some cases, it may be simply irreducible, no matter how hard you try;
Data quality and data cleaning: Data are not gospel. There could be noise and mistakes. Then what?
Wisdom of crowds: aggregating many imperfect guesses can do better than any one individual guess.

sta101-f26-files