---
title: 'PSYC 193R: Homework 5'
output: html_notebook
---

Due date: 11:59pm, Friday Feb 16, 2018

- Name: 
- PID: 

Grading Rubic | Percentage
------------- | -------------
Part 1        | 10%
Part 2        | 10%
Part 3        | 25%
Part 4        | 5%
Part 5        | 30%
Part 6        | 10%
Code Clarity (e.g., commenting, blocking, breaking down problems) | 10%


### Notes: 
To make sure your code runs: When you are done programming, close RStudio so that everything in the Console and the Environment is cleared. Then press "Run All" from the menu bar and make sure there are no errors. 

- Save a copy of this R Notebook and rename it to psyc193r_hw05_(your pid).Rmd; e.g., psyc193r_hw05_A01234565.Rmd
- When you hit "Preview" in the menu bar, an html file will be generated
- You will have to upload this R Notebook and the html output; i.e., psyc193r_hw05_A01234567.Rmd & psyc193r_hw05_A01234567.nb.html using the submission link on the class website
- Codes are run in sequential order: the code appear earlier on may be needed for later parts. When you run some codes, make sure the gray boxes above it are also executed
- Unless otherwise specified, use alpha = 0.05 (two-tailed test) for hypothesis testing and confidence interval construction
- List out the information you have whenever needed to make sure they have not been overwritten in the boxes above

### Goals: 
The objective of this assignment is to practice doing Hypothesis Testing (again), and calculate other statistics that are informative. 


## Background of the project: 
There are claims that [random noise can improve productivity](https://www.noisli.com/). As a [scientist](https://en.wikipedia.org/wiki/Scientist), Jonas wanted to test the claim. He recruited a number of undergraduate students through the UC San Diego psychology experiment website.  
Each participant in the study went through two learning phases. In each phase, they were asked to memorize 15 difficult vocabularies randomly selected from a GRE practice exam for 10 minutes. They were then shown 5 music videos as distraction, before being asked to reproduce the definitions when given the newly learned words. 
Around half of the participants studied the first list with noise cancelling headphones, so they had a completely silent environment. They moved on to study the second list with a "prodictivity" noise from the website listed above. The noise was set to 70dB for all participants.  
The other group went through the same training and test in the reversed order. 

You can access the data file at: 
https://psyc193r.ucsd.edu/data/psyc193r_wordList_data.csv

Import the data from the URL, saving it to a variable "wordListData". 
```{r}
wordListData = NA
```

The data frame has 5 variables:  

- subjId: participant number  
- order: the order of the silence / noise condition (0 == silence first)  
- gender: gender of the participant (0 == female)  
- silence: number of correct definitions recalled in the silence condition  
- noise: number of correct definitions recalled in the noise condition  

Your task is to perform an appropriate statistical test on the data set. 


## Part 1: Design of the study  
In Parts 1 - 4, we are *only interested* in the difference between silence and noise conditions.  

(Skip a line and start your answer with a ">", so that your answer appears in a "block")  

1.1. Was the study an experiment, a quasi-experiment, or an observational study? Why?  

> 

1.2 What was the research question?  

> 

1.3 What was the target population?   

> 

1.4 What were the IV(s) and DV(s)? (if any)

> IV(s):   
DV(s): 

1.5 How were the IV(s) and DV(s) operationalized? (if any; also, include the levels of the variables and scales of measurement)  

> IV(s):   
DV(s): 

1.6 Was it a between- / within-subjects design?  

> 

1.7 What are the expected results?  

> If the manipulation does not work:  
If the manipulation works:  


## Part 2: Statistical Power  

Not being convinced that the manipulation would work, Jonas expected a small to medium effect size of 0.35.  

2.1 With the sample size they had, what was the power Jonas could expect?  
```{r}
# load the package

# number of subjects in the data set
sample_size = NA

# power calculation

```

2.2 With the expected effect size, what was the minimal sample size needed to get a significant result with 80% probability if there was a true effect?  
```{r}

```

Briefly explain what the calculation (Parts 2.1 & 2.2) above means.

> 


## Part 3: Data Cleaning / Processing + Hypothesis Testing
3.1 Make sure the variables are in the correct "classes"
```{r}
wordListData$subjId = NA
```

Implementation check:  
"subjId" should be a factor, "silence" & "noise" should be numeric
```{r}
# Run this box, do not change anything
class(wordListData$subjId)
class(wordListData$silence)
class(wordListData$noise)
```


3.2 Create a new column the denotes the difference between conditions
```{r}
wordListData$diff = NA
```

Implementation check: 
```{r}
head(wordListData)
```

3.3 Check the assumptions / distribution(s) of your data
```{r}
# create histogram(s) of the distribution(s)


# other assumptions (if any)

```

3.4 Perform an appropriate statistical test for the effectiveness of having noise while studying, and report the results.  

Instructions:  

- Use the 4-step approach  
- Use the appropriate function  
- You do not need to report power, effect size, or confidence intervals in this section  
- Comment all steps that do not require calculations.  
```{r}
# Step 0: 


# Step 1: 


# Step 2: 


# Step 3: 


# Step 4: 


```


## Part 4: Type I error, Effect Size, Confidence Interval
#### Type I error
Calculate / estimate the type I error of the test you performed above

> 

#### Effect Size
Calculate / estimate the effect size of the study (from data)
```{r}
# load the package


# calculation

```

Explain what the calculations above means.  

> 

#### Confidence Interval
State the 95% confidence interval of the test above. You can obtain the answer from the t.test() output. Explain what the C.I. means in this context.  

> 

## Part 5: Gender difference in vocabulary learning
Jonas also suspected that female may be better at memorizing vocabularies than male. He believed that the difference may be small (effect size of 0.3) but reliable.  
He planned to use the total number of definitions recalled (combining silence and noise conditions) as a measure.  

5.1 Statistical test  
State the appropriate statistical test that you plan to perform  

> 

5.2 Power Analysis  
Perform a power analysis to estimate the number of female and male participants needed to show a gender difference if the two populations are actually different.  
```{r}

```

5.3 Data processing  
Make changes to the data frame for the analysis
```{r}
# change gender to a factor, add the labels, check your implementations


# create a total column

```


5.4 Assumptions / distribution(s) checking  
Check all the assumptions required for the test
```{r}
# histogram(s)


# equal variance assumptions


# other assumptions

```

Briefly explain the assumptions you checked and report what is found.  

> 

5.5 Obtain basic sample statistics
```{r}
# group means


# group sd


# number of participants in each group

```


5.6 Perform an appropriate statistical test, and report the results.  

Instructions:  

- Use the 4-step approach  
- Use the appropriate function  
- You do not need to report power, effect size, or confidence intervals in this section  
- Comment all steps that do not require calculations.  
```{r}
# Step 0: 


# Step 1: 


# Step 2: 


# Step 3: 


# Step 4: 


```

5.7 Effect Size  
Calculate / obtain the effect size from the data, and explain what it means. 
```{r}

```

5.8 Confidence Interval
Calculate a 95% confidence interval, or obtain it from some output above, and explain what it means. 

>

## Part 6: Scientific Communications (no coding needed)
An important part of doing science is to communicate our ideas well within and beyond the scientific community. We do science because we believe what we do has an impact on the society (somehow, maybe, hopefully positive).  
Based on all the analyses you have (Parts 1 to 5), explain to the public what the study is about, and the major findings of the project. Include all the test statistics and analyses (including effect size, power, etc), but make sure to explain them in terms that are understandable by the general public.  
(answer this question in a "block")

Part 1 - Part 4

> 

Part 5

> 

#### End of Homework 5