---
title: "R Notebook"
output: html_notebook
---

# Lecture 10: Sleep Research

## Import Data / Data Cleaning
Tell R where you data file is:  
- Session --> Set Working Directory --> Choose Directory
- Save that line "setwd()" below so that you can load your working directory next time without clicking

## Format checking:  
Check the comma separated values (csv) file with MS Excel to ensure correct format

## Import file:  
the file into R as a data frame, saving it to "sleepData" variable

```{r}
# you have to do setwd() and read.csv() in the same chunk
setwd("...")
sleepData = read.csv("psyc193r_sleep_data.csv")
```

Or, if your data set is online, you can import it directly:
```{r}
# data_link = "https://psyc193r.ucsd.edu/data/psyc193r_sleep_data.csv"
# sleepData = read.csv(data_link)
```

## Check the structure of the data with head()
The data have a "long" format:  
- Each row represents an entry from a participant
- Each participant can have multiple rows
```{r}
head(sleepData)
```

## Data cleaning
Factorize the nominal variables (subjId, group)
label the two levels of "group"
```{r}
sleepData$subjId = factor(NA)
sleepData$group = factor(NA, labels = c("NA","NA"))

# check if the variables are in the correct classes
class(sleepData$subjId)
class(sleepData$group)

# check the levels of "group"
levels(sleepData$group)
```

Coerce the sleepQuality into a numeric variables
```{r}
sleepData$sleepQuality = as.numeric(NA)

# check if the variable has the correct class
class(sleepData$sleepQuality)
```

Aggregate each participant's 7-day sleep quality scores
Save the new data frame into "sleep"
```{r}
sleep = aggregate(NA ~ NA * NA, data = NA, FUN = NA)
```

## Descriptive statistics for each group
```{r}
# means of the two groups
aggregate(NA ~ NA, data = NA, FUN = NA)

# standard deviation of the groups
aggregate(NA ~ NA, data = NA, FUN = NA)

# number of participants in each group
aggregate(NA ~ NA, data = NA, FUN = NA)
```


## Check assumptions
Make a histogram of sleepQuality, separated by groups
```{r}
# Placebo group
hist(NA)

# Z drugs group
hist(NA)
```

Equal variance assumptions
```{r}
# var.test() should not be significant
var.test(NA, NA)
```

## Hypothesis Testing 
```{r}
# Step 0: gather information
mu_baseline_diff = NA
# Other information

# Step 1: the hypotheses
# H0: ... == 0
# H1: ... != 0

# Step 2: find the sample t-score ("t_sample")
# obtain from t.test()
# t.test(): a two-samples t-test
t.test(NA)

# Step 3: find the p-value ("p_value")
# obtain from t.test()
# p-value = ...

# Step 4: Decision & Conclusion, Report all the statistics

```

Effect Size
```{r}
library("effsize")
cohen.d(NA ~ NA, data = NA, pooled = TRUE, paired = FALSE)
```

```{r}
library("pwr")
# pwr.t.test(n = , d = , sig.level = , power = , type = c("two.sample", "one.sample", "paired"))
pwr.t.test(n = NULL, d = NULL, sig.level = 0.05, power = 0.8, type = "two.sample")
```
