---
title: "R Notebook"
output: html_notebook
---

# Lecture 12: Fitting Room Tricks -- Posthoc tests & Visualization

## Import Data / Data Cleaning
Tell R where you data file is:  
- Session --> Set Working Directory --> Choose Directory
- Save that line "setwd()" below so that you can load your working directory next time without clicking

## Format checking:  
Check the comma separated values (csv) file with MS Excel to ensure correct format

## Import file:  
the file into R as a data frame, saving it to "lightingData" variable

```{r}
# you have to do setwd() and read.csv() in the same chunk
setwd("...")
lightingData = read.csv("...")
```

Or, if your data set is online, you can import it directly:
```{r}
# data_link = "https://psyc193r.ucsd.edu/data/psyc193r_lighting_data.csv"
# lightingData = read.csv(data_link)
```

## Check the structure of the data with head()
The data have a "long" format:  
- Each row represents an entry from a participant
```{r}
head(lightingData)
```

## Data cleaning
Factorize the nominal variables (shopperId, condition)
label the three levels of "condition"
```{r}
lightingData$shopperId = factor(NA)
lightingData$condition = factor(NA, labels = c(NA, NA, NA))

# check if the variables are in the correct classes
class(lightingData$shopperId)
class(lightingData$condition)

# check the levels of "group"
levels(lightingData$condition)
```

Coerce the sleepQuality into a numeric variables
```{r}
lightingData$money = as.numeric(NA)

# check if the variable has the correct class
class(lightingData$money)
```


## Descriptive statistics for each group
```{r}
# means of the three conditions
aggregate(NA ~ NA, data = NA, FUN = NA)

# standard deviation of the conditions
aggregate(NA ~ NA, data = NA, FUN = NA)

# number of participants in each condition
aggregate(NA ~ NA, data = NA, FUN = NA)
```


## Check assumptions
Make a histogram of sleepQuality, separated by groups
```{r}
# Regular lighting group
hist(NA)

# Harsh lighting group
hist(NA)

# Soft lighting group
hist(NA)
```

Equal variance assumptions
```{r}
# bartlett.test() should not be significant
bartlett.test(NA ~ NA, data = NA)
```

## Hypothesis Testing 
```{r}
# Step 0: gather information
# normality assumptions

# Step 1: the hypotheses
# H0: ... == ... = ...
# H1: ...

# Step 2: Build an aov object 
lighting.aov = NA

# Step 3: obtain from summary()
# obtain the sample F score using summary()
# obtain the p-value ("p_value")
summary()

# Step 4: Decision & Conclusion, Report all the statistics

```

## New from Lecture 11
Post-hoc test
```{r}
TukeyHSD(NA, "NA")
```

Power
```{r}
# import the pwr package

# calculate power based on pairwise tests

```

Visualization: part 1 -- create summary data frame
```{r}
# data frame for different conditions
# find the means
lightingSummary = aggregate(NA ~ NA, data = lightingData, FUN = NA)
names(lightingSummary)[2] = "sales"

# find the estimated sd
lightingSummary.sd = aggregate(NA ~ NA, data = lightingData, FUN = NA)
names(lightingSummary.sd)[2] = "NA"

# find the number of participants in each condition
lightingSummary.n = aggregate(NA ~ NA, data = lightingData, FUN = NA)
names(lightingSummary.n)[2] = "NA"

# combine the data frames (by column)
lightingSummary = cbind(lightingSummary, NA, NA)
names(lightingSummary)[3] = "NA"
names(lightingSummary)[4] = "NA"

# remove redundant data frames
rm(lightingSummary.n, lightingSummary.sd)

# find the standard errors of each condition
lightingSummary$se = NA / NA

```

Visualization: part 2 -- ggplot
```{r}
#loading the ggplot2 package
library(ggplot2)

# basic graph
ggplot(data=lightingSummary, aes(x=NA, y=NA)) +
  geom_bar(stat="identity")

# with error bars
ggplot(data=lightingSummary, aes(x=NA, y=NA)) + 
  geom_bar(stat="identity") + 
  geom_errorbar(aes(ymin=NA-NA, ymax=NA+NA), width=.5)
```
