---
title: "PSYC 193R: Lecture 6 - Data Processing"
output: html_notebook
---

# Lecture 6/7: data processing / one-sample t-test

## Format checking:  
Check the comma separated values (csv) file with MS Excel to ensure correct format

## Import Data / Data Cleaning
Tell R where you data file is:  
- Session --> Set Working Directory --> Choose Directory
- Save that line "setwd()" below so that you can load your working directory next time without clicking

## Import file:  
Import the file into R as a data frame, saving it to "faceReadData" variable

```{r}
# you have to do setwd() and read.csv() in the same chunk
setwd("...")
faceReadData = read.csv("psyc193r_readface_data.csv")
```

Or, if your data set is online, you can import it directly:
```{r}
# data_link = "https://psyc193r.ucsd.edu/data/psyc193r_readface_data.csv"
# faceReadData = read.csv(data_link)
```

## Check the structure of the data with head()
The data have a "long format":  
- Each column represents a variable
- Each row represents an entry (each trial)
```{r}
head(faceReadData)
```

## Data cleaning
Factorize the nominal variable (subjectId)
```{r}
factor(NA)
```

Coerce "correct" to a numeric variable
```{r}
as.numeric(NA)
```

Make a new data frame called faceRead with two columns, aggregating each subject's accuracy
```{r}
# aggregate the trials to form an accuracy score for each participant
aggregate(NA)

# rename the aggregated score to "accuracy"
names(NA)
```

## Hypothesis Testing 
```{r}
# Step 0: gather information
mu_baseline = NA
s = NA
mean_face = NA
sample_size = NA

# some processing from the info
df = NA
se = NA

# check the distribution of scores
hist(NA)

# Step 1: the hypotheses


# Step 2: find the sample t-score ("t_sample")


# Step 3: find the p-value ("p_value")


# Step 4: Make a decision, provide a conclusion and the test statistics


```
