---
title: "R Notebook"
output: html_notebook
---

# Lecture 9: Blind Tasting Data Processing

## Import Data / Data Cleaning
Tell R where you data file is:  
- Session --> Set Working Directory --> Choose Directory
- Save that line "setwd()" below so that you can load your working directory next time without clicking

## Format checking:  
Check the comma separated values (csv) file with MS Excel to ensure correct format

## Import file:  
the file into R as a data frame, saving it to "blindTasteData" variable

```{r}
# you have to do setwd() and read.csv() in the same chunk
setwd(NA)
blindTasteData = read.csv(NA)
```

Or, if your data set is online, you can import it directly:
```{r}
# data_link = "https://psyc193r.ucsd.edu/data/psyc193r_blindTaste_data.csv"
# blindTasteData = read.csv(data_link)
```

## Check the structure of the data with head()
The data have a "wide format":  
- Each column represents a variable
- Each row represents multiple responses from a participant
```{r}
head(blindTasteData)
```

## Data cleaning
Factorize the nominal variables (set1, set2)
```{r}
blindTasteData$set1 = factor(NA)

# make sure these are all factors
class(blindTasteData$set1)
class(blindTasteData$set2)
```

Coerce the ratings into a numeric variables
```{r}
blindTasteData$set1_chocolate = as.numeric(NA)

# make sure these are all numeric
class(blindTasteData$set1_chocolate)
class(blindTasteData$set1_almond)
class(blindTasteData$set2_chocolate)
class(blindTasteData$set2_almond)
```

Remove invalid entries
```{r}
blindTasteData[blindTasteData$set1 != NA,]
```


Make two new variables: each represent a composite score of chocolate + almond of the set
```{r}
# calculate each participant's set 1 chocolate + set 1 almond scores
# save it to a new variable set1_score
blindTasteData$set1_score = blindTasteData$set1_chocolate + ...

# do the same for set 2
blindTasteData$set2_score = blindTasteData$set2_chocolate + ...
```

Make two new variables: one represents the scores of set A chocolate, the other for set B
```{r}
# placeholder
blindTasteData$setA = NA
blindTasteData$setB = NA

# for loop
for (i in 1:nrow(blindTasteData)){
  if (blindTasteData$set1[i] == "A"){
    blindTasteData$setA[i] = ...
  } else {
    blindTasteData$setA[i] = ...
  }
}

```

Make a histogram of the difference between setA and setB
```{r}
blindTasteData$diff = NA
hist(NA)
```

## Hypothesis Testing 
```{r}
# Step 0: gather information
mu_baseline = NA
s_diff = NA
mean_diff = NA
sample_size = NA

# some processing from the info
df = NA
se = NA

# check the distribution of scores
hist(NA)

# Step 1: the hypotheses


# Step 2: find the sample t-score ("t_sample")


# Step 3: find the p-value ("p_value")


# Step 4: Make a decision, provide a conclusion and the test statistics


```

t.test() function
```{r}
# put in difference scores as the argument: not a paired test
t.test(blindTasteData$diff, mu = 0, paired = FALSE)

# put in the two arrays, set A and set B, as the arguments, a paired test
t.test(blindTasteData$setA, blindTasteData$setB, mu = 0, paired = TRUE)
```

Effect Size
```{r}
# information needed for cohen's d
mean_diff = NA
estimated_sd = NA
baseline_pop_mean = NA

# actual calculation
cohens_d = NA
```

Power
```{r}
library("pwr")
# pwr.t.test(n = , d = , sig.level = , power = , type = c("two.sample", "one.sample", "paired"))
pwr.t.test(NA)
```

# t-test with "long" format data
In our face reading experiment, we also collected response times for each trial. Each participant looked at male and female pictures.  
We can test if participants look at male or female faces for longer.  

## Import file:  
the file into R as a data frame, saving it to "faceReadData" variable
```{r}
# Read in the data
faceReadData = read.csv("https://psyc193r.ucsd.edu/data/psyc193r_readface_data.csv")

# remove response of "myFriend"
faceReadData = faceReadData[faceReadData$chosen != NA,]
head(faceReadData)
```
The data has a "long" format. Each row represents one trial / entry.  
We will need to create a new data frame, with these columns:  

- subjectId
- pic_gender: gender of picture
- responseTime

## Coerce data into the correct classes
also changing pronoun into male & female
```{r}
faceReadData$subjectId = factor(faceReadData$subjectId)
faceReadData$pronoun = factor(faceReadData$pronoun)
faceReadData$responseTime = as.numeric(faceReadData$responseTime)

# Make a new variable for gender of the pictures
faceReadData$pic_gender = NA
faceReadData[faceReadData$pronoun == "her",]$pic_gender = NA
faceReadData[faceReadData$pronoun == NA,]$pic_gender = NA
```

## Aggregate male and female pictures for each participant
saving the new variable to faceRead
also make sure the subjectId and pic_gender is sorted, such that each partcipate has two rows
```{r}
faceRead = aggregate(responseTime ~ NA * NA, data = faceReadData, FUN = NA)

faceRead = faceRead[order(faceRead$subjectId, faceRead$pic_gender),]
```

## Check the assumptions
```{r}
faceRead_diff = faceRead[faceRead$pic_gender == "female",]$responseTime - faceRead[faceRead$pic_gender == "male",]$responseTime

hist(faceRead_diff, freq = FALSE, breaks = seq(-5000, 20000, 2500), 
     xlab = "Difference in Response Time (Female pictures - Male pictures)")
xfit = seq(-5000, 20000, 25)
yfit = dnorm(xfit, mean = mean(faceRead_diff), sd = sd(faceRead_diff))
lines(xfit, yfit, col="blue", lty = 3, lwd = 3)

# Caution in interpreting the data: sample scores are not normally distributed, and there are outliers
```

## Perform a t.test()
```{r}
t.test(responseTime ~ NA, mu = 0, data = faceRead, paired = TRUE)

# get the means for each condition
aggregate(responseTime ~ pic_gender, data = faceRead, FUN = mean)
```

## Perform effect size / power / C.I. measures as above