---
title: "R Notebook"
output: html_notebook
---

# Lecture 16: Second Language (L2) Acquisition

## Import Data / Data Cleaning
Tell R where you data file is:  
- Session --> Set Working Directory --> Choose Directory
- Save that line "setwd()" below so that you can load your working directory next time without clicking

## Format checking:  
Check the comma separated values (csv) file with MS Excel to ensure correct format

## Import file:  
the file into R as a data frame, saving it to "lightingData" variable

```{r}
# you have to do setwd() and read.csv() in the same chunk
setwd("...")
language = read.csv("NA")
```

Or, if your data set is online, you can import it directly:
```{r}
# data_link = "https://psyc193r.ucsd.edu/data/psyc193r_lighting_data.csv"
# lightingData = read.csv(data_link)
```

## Check the structure of the data with head()
The data have a "long" format:  
- Each row represents an entry from a participant
```{r}
head(NA)
```

## Data cleaning
Factorize the nominal variables
Label the levels of "gender" and "age"
```{r}
# Factorize variables


# check if the variables are in the correct classes


# check the levels the nominal variables
levels(NA)
```

Coerce the dv into a numeric variables
```{r}
# coerce variables into numeric type

# check if the variable has the correct class
class(NA)
```

Combine written and spoken test scores (named "score")
```{r}

```


## Descriptive statistics for each group (4 groups)
```{r}
# means
aggregate(NA ~ NA * NA, data = NA, FUN = NA)

# standard deviations
aggregate(NA ~ NA * NA, data = NA, FUN = NA)

# number of participants
aggregate(NA ~ NA * NA, data = NA, FUN = NA)
```


## Check assumptions
Number of subjects in each condition
```{r}

```

Outliers: Make a 2-by-2 histogram grid of test score, separated by conditions
```{r}
# draw a two by two plot grid
par(mfrow=c(NA,NA))

# draw a histogram for each age and gender level
for (i in levels(NA)){
  for (j in levels(NA)){
    hist(language[language$age == i & language$gender == j,]$NA, 
         xlab = "score",
         main = paste0(i, " / ", j))
  }
}

# reverse back to a 1 by 1 grid
par(mfrow=c(1,1))

```

Equal variance assumptions
```{r}
# bartlett.test() should not be significant
bartlett.test(NA ~ NA, data = NA)
```


## Build an aov() object
```{r}
# Build an aov() object 
lighting.aov = NA

# Read out summary of the aov() object
summary()

# Summarize the results

```

Effect Size
```{r}
# effect size for age


# effect size for gender


# effect size fot the age * gender interaction

```

Post-hoc test
```{r}
# perform post test only if you have 3+ levels in a factor

```

## Power  
Perform power analysis as if you test each factor independently
```{r}
# import the pwr package
library("pwr")

# follow the syntax: 
# pwr.anova.test(f = .14, k = 3, n = 30, power = NULL)

```

## Visualization
Part 1: summary data frame
```{r}
# Make a summary data frame, with these columns
# 1. age
# 2. gender
# 3. mean score of each group
# 4. sd of scores of each group
# 5. sample size n of each group
# 6. se of each group
languageSummary = NA
```

Part 2: ggplot() for Visualization
```{r}
# import the ggplot2 package
library(ggplot2)

# plotting a bar graph, separated by groups
ggplot(data = NA, aes(x = NA, y = NA, fill = NA)) + 
  geom_bar(position=position_dodge(), stat="identity") + 
  geom_errorbar(aes(ymax = NA + NA, ymin = NA - NA), 
                size = 1, width = .5, position=position_dodge(.9))
```

