---
title: 'PSYC 193R: Homework 8'
output: html_notebook
---

Due date: 3:00pm, Tuesday March 20, 2018

- Name: 
- PID: 

Grading Rubic            | Percentage
------------------------ | -------------
Part 1                   | 12%
Part 2                   | 8%
Part 3                   | 35%
Part 4                   | 9%
Part 5                   | 1%
Part 6                   | 5%
Part 7                   | 10%
Scientific Communication | 10%
Code Clarity (e.g., commenting, blocking, breaking down problems) | 10%


### Notes: 
To make sure your code runs: When you are done programming, close RStudio so that everything in the Console and the Environment is cleared. Then press "Run All" from the menu bar and make sure there are no errors. 

- Save a copy of this R Notebook and rename it to psyc193r_hw08_(your pid).Rmd; e.g., psyc193r_hw08_A01234565.Rmd
- When you hit "Preview" in the menu bar, an html file will be generated
- You will have to upload this R Notebook and the html output; i.e., psyc193r_hw08_A01234567.Rmd & psyc193r_hw08_A01234567.nb.html using the submission link on the class website
- Codes are run in sequential order: the code appear earlier on may be needed for later parts. When you run some codes, make sure the gray boxes above it are also executed
- Unless otherwise specified, use alpha = 0.05 (two-tailed test) for hypothesis testing and confidence interval construction
- List out the information you have whenever needed to make sure they have not been overwritten in the boxes above

### Goals: 
The objective of this assignment is to practice doing Hypothesis Testing (again), and calculate other statistics that are informative. 


## Background of the project: 
This study was actually performed in the PSYC 193R class at UC San Diego. However, the number of students in the class was small and it did not provide a meaningful enough conclusion. 

Jonas was not a regular coffee drinker, so he would like to study the effect of coffee drinking and decided whether coffee drinking is good for him.   
To study the effect of coffee drinking and productivity, Jonas recruited some undergraduate students in the Price Center for a study on a Wednesday afternoon. He recruited only students who did not drink coffee on that day. He provided participants free starb**ks coffee. Roughly half of the participants got a cup of decaffeinated coffee, and the remaining got regular coffee.  
Fifteen minutes after people got their coffee, they were given a comprehension test. The test contained two passages. The two parts were graded separately. They also indicated whether they were regular coffee drinkers at the end of the test.  

You can access the data file at: 
https://psyc193r.ucsd.edu/data/psyc193r_coffee_data.csv

Use the data to examine whether coffee-drinking habits and caffeine affect productivity.  

Import the data from the URL, saving it to a variable "coffeeData". 
```{r}
coffeeData = NA
```

The data frame has 7 variables:  

- subjId: participant number  
- group: coffee they got / regular coffee drinker  
  (0 -- decaf/nondrinker, 1 -- decaf/drinker  
   2 -- regular/nondrinker, 3 -- regular/drinker)  
- age: age of the participant
- type: type of coffee (0 -- decaf, 1 -- regular)
- gender: 0 -- female, 1 -- male
- people: 0 -- nondrinker, 1 -- drinker
- passageA: score for first comprehension test (max 20 points)
- passageB: score for second comprehension test (max 15 points)

Check whether the data structure is ideal for the analysis you planned to perform. 


## Part 1: Design of the study  
(Skip a line and start your answer with a ">", so that your answer appears in a "block")  
1.1. Was the study an experiment, a quasi-experiment, or an observational study?  
Explain your decision. 

> 

1.2 What was the research question?  

> 

1.3 What was the target population? What was the mean and standard deviation of the age of the sample? 

```{r}
# some calculations, if needed

```

> 

1.4 What were the factor(s) and DV(s)? (if any)

> Factors(s):   
DV(s): 

1.5 How were the factor(s) and DV(s) operationalized?  
(if any; also, include the levels of the variables and scales of measurement)  

> Factors(s):   
DV(s): 

1.6 Was it a between- / within-subjects design? What kind of design is it? (E.g., one-way ANOVA, 3 x 2 ANOVA, etc)

```{r}

```

> 

1.7 State the expected results in words (not formulas).  

> For factor 1 (state factor 1):  
> For factor 2 (state factor 2):  


## Part 2: Statistical Power  

Assuming that Jonas believes drinking regular coffee vs decaffined coffee can explain 15% of the variance in the comprehension scores. What is the minimal sample size that he needed in the study to achieve a 85% power?
```{r}
# load the package (if needed)


# Calculations


# Brief explanation (not more than one sentence)

```


## Part 3: Data Cleaning / Processing + Assumptions checking, Hypothesis Testing
3.1 Make sure the variables are in the correct "classes"
```{r}

```

Implementation check:  
```{r}
# check the classes of the variables

```

3.2 Combine the results of the two comprehension test and form a single measure. Save the values as "score" in the data frame.
```{r}

```

3.3 State the assumptions, and perform calculations if needed
```{r}
# assumption 1: (state your assumption)


# assumption 2: (state your assumption)


# assumption 3: (state your assumption)


# Other assumptions: (if any)

```

Describe the assumptions you checked, and report the results and statistics when appropriate. 

> 

3.4 Summary statistics, saving them into some variables that you can use later
```{r}
# mean of the conditions
coffeeSummary = NA 

# print out the means


# sd of the conditions
coffeeSummary.sd = NA 

# print out the sd


# number of participants in the conditions
coffeeSummary.n = NA

# print out the numbers

```

3.5 Perform an appropriate statistical test.  
Before building the model, sketch the graph you expect on a paper. 
```{r}
# build an aov object

# read out the summary table

# check the degrees of freedom to make sure you build the correct model

```

State the results. Describe the main effect(s) and the interaction. 

> 


## Part 4: Effect Size  
Calculate / report the effect size of the main effect(s) and interaction
```{r}

```

Explain what the calculations above mean.  

> 

## Part 5: Post-hoc test
5.1 Briefly explain whether / why a post-hoc test is needed. 

> 

5.2 Performed the post-hoc test if needed. 
```{r}

```


## Part 6: Visualization
Make a data frame that contains information needed for the graph.
Construct the data frame by adding columns to the coffeeSummary variable you created earlier. 
```{r}
# data frame for the graph

```

Make a graph to depict the results. 
```{r}
# load the relevant package


# make a bar graph with error bars


# make a line graph with error bars

```

State what the error bars denote. 

> 

## Part 7. Scientific Communications (no coding needed)
Summarize what you found in the study.  

- Explain the background in 2 - 3 sentences
- State the research question(s)
- Justify the number of participants used in the study (when appropriate)
- Explain the results of the hypothesis testing, and include all relevant test staitsics in a format that you would see in a scientific journal
- Include relevant summary statistics and effect size of the test.  
- Make a suggestion to Jonas and other non-coffee drinkers as to whether they should start drinking coffee, assuming your audience are college students without a lot of statistics training. 

You can assume that all the graphs are included, and you should have most of the information ready from previous sections.  
(answer this question in a "block")

> 

(disclaimer: the large-scale study was never conducted.)
#### End of Homework 8