---
title: 'PSYC 193R: Homework 3'
output: html_notebook
---

Due date: 11:59pm, Friday Feb 2, 2018

- Name: 
- PID: 

Grading Rubic | Percentage
------------- | -------------
Part 1        | 10%
Part 2        | 25%
Part 3        | 25%
Part 4        | 10%
Part 5        | 20%
Code Clarity (e.g., commenting, blocking, breaking down problems) | 10%


### Notes: 
To make sure your code runs: When you are done programming, close RStudio so that everything in the Console and the Environment is cleared. Then press "Run All" from the menu bar and make sure there are no errors. 

- Save a copy of this R Notebook and rename it to psyc193r_hw03_(your pid).Rmd; e.g., psyc193r_hw03_A01234567.Rmd
- When you hit "Preview" in the menu bar, an html file will be generated
- You will have to upload this R Notebook and the html output; i.e., psyc193r_hw03_A01234567.Rmd & psyc193r_hw03_A01234567.nb.html using the submission link on the class website
- Codes are run in sequential order: the code appear earlier on may be needed for later parts. When you run some codes, make sure the gray boxes above it are also executed
- Unless otherwise specified, use alpha = 0.05 (two-tailed test) for hypothesis testing and confidence interval construction
- list out the information you have whenever needed to make sure they have not been overwritten in the boxes above

### Goals: 
The objective of this assignment is to practice doing Hypothesis Testing (again), and calculate other statistics that are informative. 


## Background of the project: 
A group of UC San Diego economists (hereafter "the economists") is interested in financial literacy among teenagers. They are worried that targeted advertisements have become increasing successful through social media, which in turn decrease teenagers' ability to manage their personal finance when they grow up.    
The economists devised a financial education program that would hopefully improve people's financial management skills. A 5-day program was run in the UCSD Preuss School in 2010. Senior high school students took classes on money concepts and budget management. The senior class had 120 students and all students participated in the program. At the time of the program, those students did not have a credit score.  
In 2018, the economists contacted the former students, and solicited their [FICO credit scores](https://www.myfico.com/credit-education/credit-scores/). FICO is a standardized scale and the scores range from 300 - 850.  
The mean of the scale is 715 and the standard deviation is 140.  
Visualizing the distribution:  
```{r}
# FICO distribution graph: Do not change anything in this box

# Distribution parameters
fico = seq(300, 850, 5)
fico_mean = 715
fico_sd = 140

# Probability function
prob = dnorm(fico, fico_mean, fico_sd)

# Plotting
plot(x = fico, y = prob, main = "FICO Credit Score Distribution", xlab = "FICO credit score", ylab = "Probability", type = "l", lwd = 3, xaxt="n")
axis(1, at = round(seq(fico_mean - 4*fico_sd, fico_mean + 4*fico_sd, by = fico_sd),0), las=1)

# Shading
# Lowest 25%
fico1 = 300
fico2 = qnorm(.25, mean = fico_mean, sd = fico_sd)
polygon(
  c( fico1, fico[fico>=fico1 & fico<=fico2], fico2, fico2, fico1 ),
  c( dnorm(fico1, fico_mean, fico_sd), prob[fico>=fico1 & fico<=fico2], dnorm(fico2, fico_mean, fico_sd), 0, 0 ),
  col=rgb(.9, .1, .1), lty = 0)

# 25% - 50%
fico1 = qnorm(.25, mean = fico_mean, sd = fico_sd)
fico2 = qnorm(.50, mean = fico_mean, sd = fico_sd)
polygon(
  c( fico1, fico[fico>=fico1 & fico<=fico2], fico2, fico2, fico1 ),
  c( dnorm(fico1, fico_mean, fico_sd), prob[fico>=fico1 & fico<=fico2], dnorm(fico2, fico_mean, fico_sd), 0, 0 ),
  col=rgb(.9, .5, .5), lty = 0)

# 50% - 75%
fico1 = qnorm(.50, mean = fico_mean, sd = fico_sd)
fico2 = qnorm(.75, mean = fico_mean, sd = fico_sd)
polygon(
  c( fico1, fico[fico>=fico1 & fico<=fico2], fico2, fico2, fico1 ),
  c( dnorm(fico1, fico_mean, fico_sd), prob[fico>=fico1 & fico<=fico2], dnorm(fico2, fico_mean, fico_sd), 0, 0 ),
  col=rgb(.5, .9, .5), lty = 0)

# highest 25%
fico1 = qnorm(.75, mean = fico_mean, sd = fico_sd)
fico2 = 850
polygon(
  c( fico1, fico[fico>=fico1 & fico<=fico2], fico2, fico2, fico1 ),
  c( dnorm(fico1, fico_mean, fico_sd), prob[fico>=fico1 & fico<=fico2], dnorm(fico2, fico_mean, fico_sd), 0, 0 ),
  col=rgb(.1, .9, .1), lty = 0)

# Annotate
text(500, .0005, "Lowest \n  25%", cex=1.2, pos=4, col="white")
text(625, .0005, " 26 -\n 50%", cex=1.2, pos=4, col="white")
text(720, .0005, " 51 -\n 75%", cex=1.2, pos=4, col="white")
text(805, .0005, "76+\n %", cex=1.2, pos=4, col="white")

# Remove all variables
rm(list=ls(all=TRUE))
```

Prior to the commencement of the program, the economists estimated that the program would improve people's score by an average of 55 points.  

Reference:  
[How to help teenagers manage their money](https://www.moneyadviceservice.org.uk/en/articles/how-to-help-teenagers-manage-their-money)

## Part 1: Design of the study  
(SKip a line and start your answer with a ">", so that your answer appears in a "block")  

1.1. Was the study an experiment, a quasi-experiment, or an observational study? Why?  

> 

1.2 What was the research question?  

> 

1.3 What was the target population?   

> 

1.4 What were the IV(s) and DV(s)? (if any)

> IV(s):  
DV(s): 

1.5 How were the IV(s) and DV(s) operationalized? (if any; also, include the levels of the variables and scales of measurement)  

> IV(s):   
DV(s):   

1.6 Was it a between- / within-subjects design?  

> 

1.7 What are the expected results?  

> If the manipulation does not work:  
If the manipulation works:  


## Part 2: Statistical Power  
Assume an alpha level of 0.05 is common in the economics literature, just like that in the psychology literature.  

2.1 If everyone reports his / her FICO credit score, what is the probability that the economists can show an effect of their program if it is actually effective?    
Hint: make use of the functions you wrote in class / in previous assignments, paste them in the following box (optional):  
```{r}

```

List out all the information you have. Calculate the probability:  
```{r}

```

2.2 Clearly, hoping that all former students are willing to share their FICO scores is unrealistic. What is the minimal number of participants needed for the economists to achieve a power of at least 80%, assuming those who respond will be representative of the whole sample?  
Write a for loop and estimate the number needed. Hints:  

- list out all the information the economists have again
- they only have a certain number of students in the senior class (don't loop over to infinity)
```{r}

```

2.3 Calculate the estimated effect size of the study, from the expected results of the economists
```{r}

```

Summarize your analyses above (2.1, 2.2, & 2.3) in a "block" below:   

> 


## Part 3: Hypothesis Testing
An exciting time finally arrived as a number of former students got in touch with the economists and reported their FICO credit scores. The scores are stored in the following variable:  
```{r}
# Run this box, do not change anything
score = c(734,930,736,782,761,850,775,738,819,767,875,710,832,718,634,855,845,754,537,755,674,785,571,571,561,1226,502,827,850,909,880,620,929,714,706,710,845,741,824,787,565,888,665,735,776,881,766,793,627,634,557,565,639,930,830,581,740,997,782,793,781,651,643,432,722,780,740,828,850,717,950,662,771,714,671,865,734,771,728,767)
```

Perform an appropriate statistical test for the effectiveness of the program, and report the results.  
Hint: make use of the function(s) you wrote in class / in previous assignments, paste them in the following box (optional):  
```{r}

```

Use the four-step approach to conduct the statistical test, including step 0 and check all the assumptions. Comment all steps that does not require calculations.  
```{r}
# Step 0: 

# Step 1: 

# Step 2: 

# Step 3: 

# Step 4: 

```


## Part 4: Type I error, Effect Size, Confidence Interval
#### Type I error
4.1 Calculate / estimate the type I error of the test you performed above and explain your answer
```{r}

```

#### Effect Size of the Study
4.2 Calculate the effect size of the study, and explain what they mean. Comment your explanation in the box
```{r}

```

#### Confidence Interval
4.3 Calculate / estimate a 95% confidence interval of the test above, and explain what it means. Comment your explanation in the box
```{r}

```


## Part 5: Scientific Communications (no coding needed)
An important part of doing science is to communicate our ideas well within and beyond the scientific community. We do science because we believe what we do has an impact on the society (somehow, maybe, hopefully positive). Based on all the analysis you have (Part 2 to 4), explain to the public what the study is about, and the major findings. Include all the test statistics, but make sure to explain them in terms that are understandable by the funding agency, Preuss School board of directors, the participants, and the general public. It would be helpful to imagine that you are explaining the results to the participants, who might not go on to college after graduating from Preuss.  
Hint: Use the last two slides of Lecture 5 to guide your thinking.  
(answer this question in a "block")

> 

#### End of Homework 3