---
title: 'PSYC 193R: Homework 2'
output: html_notebook
---

# Homework Assignment 2  
Due date: 11:59pm, Friday Jan 26, 2018

- Name: 
- PID: 

Grading Rubic | Percentage
------------- | -------------
Part 1        | 10%
Part 2        | 60%
Part 3        | 20%
Code Clarity (e.g., commenting, breaking down problems) | 10%


### Notes: 
To make sure your code runs: When you are done programming, close RStudio so that everything in the Console and the Environment is cleared. Then press "Run All" from the menu bar and make sure there are no errors. 

- Save a copy of this R Notebook and rename it to psyc193r_hw02_(your pid).Rmd; e.g., psyc193r_hw02_A01234567.Rmd
- When you hit "Preview" in the menu bar, an html file will be generated
- You will have to send in (silau -et- ucsd.edu) this R Notebook and the html output; i.e., psyc193r_hw02_A01234567.Rmd & psyc193r_hw02_A01234567.nb.html
- Codes are run in sequential order: the code appear earlier on may be needed for later parts. When you run some codes, make sure the gray boxes above it are also executed
- Unless otherwise specified, use alpha = 0.05 (two-tailed test) for hypothesis testing and confidence interval construction

### Goals:
The objective of this assignment is to practice using the pnorm() and qnorm() functions, and to perform hypothesis testing using Z-test using the more common approach (comparing p and alpha). 

## Part 1: probability and quantile conversion
Hint:  
- Decide whether it is a percentile --> quantile or a quantile --> percentile problem  
- Always sketch the curve even when you're not required to do so in this assignment  

#### Your task: 
SAT is a standardized college entrance exam. It has a maximum score of 1600, a mean of 990 and standard deviation of 220. 

1.1 UC San Diego only consider applicants scoring at or above the 80th percentile. What is the cut-off SAT score that an applicant has to attain to be considered?
```{r}
# your code here

```

1.2 In 2017, The UCSD Preuss School had 64 students in its senior class. The class attained a mean SAT score of 1200. What is the probability of getting such a mean score or above in the population?
```{r}
# your code here

```

1.3 Jason was a senior studying at the Preuss School in 2017. He got a SAT score of 1350. How many students in the population (in percentage) score above Jason?
```{r}
# your code here

```

1.4 A researcher plans to perform a 2-tailed Z test, and set the alpha level as 0.01. What is/are the critical Z-score(s) for her to reject the null hypothesis?
```{r}
# your code here

```

The IQ scores in the general population has a mean of 100 and a standard deviation of 15.   
1.5 To be eligible to operate a vehicle, the DMV requires an individual to have an IQ score of 70 or above. What is the percentage of the population that is eligible to drive (provided that they pass the driving test)?
```{r}
# your code here

```

1.6 What is the percentage of the population that has an IQ score between 80 to 120?
```{r}
# your code here

```


## Part 2: Pre-term birth on IQ scores
Preterm birth is risky, and the newborns are known to have some health vulnerabilities. 
According to some Finland government statistics, adults in the Finland population who are born full-term have an average IQ score of 102 and a standard deviation of 15. 
A Finland study tested 100 adults who were born pre-term. The IQ scores obtained are stored in the numeric array "pretermIQ".  
Reference: [Being Born Premature May Hurt Your IQ](http://healthland.time.com/2011/12/06/being-born-premature-may-hurt-your-iq/)  

```{r}
# Load the data set
pretermIQ = c(102,86,78,100,100,87,92,89,68,82,85,108,84,74,110,102,89,103,92,100,105,94,102,103,94,100,77,116,84,80,78,91,96,103,90,79,92,93,83,85,115,96,121,101,79,83,103,87,82,123,99,109,88,98,113,83,68,82,103,91,96,102,102,100,94,104,85,90,82,110,78,109,69,97,76,84,88,97,80,90,111,72,85,109,132,105,81,102,102,93,77,123,83,86,87,98,90,103,82,106)

```

#### Your task: 
Answer the follow questions: (no coding needed in the following box, answer the question as comments)
```{r}
# Was this study a/an experimental / quasi-experimental / observational study? What is your reasoning behind?
# Your answer: 

# What was the population of interest? 
# Your answer: 
```

Calculate the best estimation of the mean of the target population, and save it to a variable "mean_pretermIQ"
```{r}
# Your code here (1 line)
mean_pretermIQ = NA
# Your code ends here
```

Gather the information, save them into appropriate variables, and perform a hypothesis test on the research question (as you would for a regular psychology research study), and conclude whether preterm babies have lower IQ scores when they grow up  

Hints:  

- Use the 4-step approach (indicate each of the 4 steps)
- save all the information you have as variables in Step 0
- Comment all the steps and other information that does not involve calculations
```{r}
# Your code here
# Step 0: 
pop_mean = NA
sample_mean = NA
pop_sd = NA
sample_size = NA

# Step 1

# Step 2

# Step 3

# Step 4

# Your code ends here
```

Calculate a 95% C.I. of the population mean for the target population  
Another way of putting it: If we repeated the study over and over with the same condition, what is the range of sample means that you would get for 95% of the time  

Hints:  

- List out all the information you need for the caluclation, save them as variables (or copy from above)
- store the alpha level as "alpha"
- Use the information above for the calculation
```{r}
# Your code here (multiple lines)

# Your code ends here
```

## Part 3: Functions
R does not have a built in function for Z-tests. But you can always create one for your needs!

Let's break down the problem into two parts: 

3.1 A function that converts a test mean (test_mean) into a sample Z score.  
You'll need the following as inputs:  

- sample mean (sample_mean)
- sample size (sample_n)
- the baseline population mean (pop_mean)
- the population standard deviation (pop_sd)

Let's name the function sample_mean_to_z()
```{r}
sample_mean_to_z = function(sample_mean, sample_n, pop_mean, pop_sd){
  # Calculate the standard error -- 
  # that's the dispersion of the sampling distribution (1 line)
  
  # Convert the sample mean into a sample z score (z_sample) 
  # using pop_mean, standard error, and test mean (1 line)
  
  # Return the sample z score (1 line)
  return (NA)
}
```

- Implementation check: (Do not change anything in this box)
```{r}
# Some test parameters
pop_mean = 15
pop_sd = 5
sample_mean = 10
sample_n = 25

# To test your function
# Expected output: -5
sample_mean_to_z(sample_mean = sample_mean, sample_n = sample_n, pop_mean = pop_mean, pop_sd = pop_sd)
```


3.2 Create a function that would first convert the sample mean into a sample Z score (using the function above), saving it as z_sample.  
Then it would turn z_sample into a p-value. The p-value denotes the probability of getting such a sample Z score, or a more extreme one (two-tailed).  

The inputs are:  

- the test mean (sample_mean)
- sample size (sample_n)
- the baseline pooulation mean (pop_mean)
- the population standard deviation (pop_sd)

The output is an array.  
- First element of the array is the sample Z score
- Second element of the array is the p-value. 

Hints:  
- The function should work for sample means that lie either on the left or right hand side of the population mean
- You can built the function that works for sample_mean < pop_mean first, then eventually it should work for both sample_mean > pop_mean & sample_mean < pop_mean
- You can use the function abs() to get the absolute value of a number.  

Let's name our function z.test()
```{r}
z.test = function(sample_mean, sample_n, pop_mean, pop_sd){
  # Convert the test mean into a Z score (1 line)
  sample_z = NA
  
  # Take the absolute value of the Z score (1 line), saving it to sample_z
  # Note: that way you make sure the Z score is 
  #       on the right-hand side of the mean (i.e., 0)
  sample_z_new = NA
  
  # Make the Z score negative, but retain the same value (1 line)
  # Note: that way you make sure the Z score is now 
  #       on the left-hand side of the mean
  sample_z_new = NA
  
  # Convert the Z score into 1-tailed probability (1 line)
  # That's the probability for getting a -ve Z score above, 
  # or a score less than the value
  p_value = NA
  
  # Multiple the probability you got above by a factor 
  # to produce a two-tailed p-value (1 line)
  p_value = NA
  
  # Create a new array with 2 scores, 
  # the sample Z score and the p-value (1 line)
  answer = NA
  
  # Return the array
  return (NA)
}
```

- Implementation check: (Do not change anything in this box)
```{r}
# Some test parameters
pop_mean = 15
pop_sd = 5
sample_mean = 10
sample_n = 25
alpha = 0.05

# To test your function
# Expected output: (-5.000000e+00  5.733031e-07)
z.test(sample_mean, sample_n, pop_mean, pop_sd)

# To test your function
sample_mean = 16
# Expected output: (1.0000000 0.3173105)
z.test(sample_mean, sample_n, pop_mean, pop_sd)
```

- Use your function to check your answer in Part 2
```{r}
# Make sure you list out all the variables again (multiple lines)

# Run the test with your z.test() function
# Expected result: Check against your Part 2 result (1 line)
z.test(NA, NA, NA, NA)
```

#### End of Homework Assignment 2
