---
title: "PSYC 193R: Homework"
output: html_notebook
---

Due date: 11:59pm, Friday March 9, 2018

- Name: 
- PID: 

Grading Rubic                     | Percentage
--------------------------------- | -------------
Building a regression model       | 90%
Automatic model selection         | 0%
Code Clarity (e.g., commenting, blocking, breaking down problems) | 10%


### Notes: 
To make sure your code runs: When you are done programming, close RStudio so that everything in the Console and the Environment is cleared. Then press "Run All" from the menu bar and make sure there are no errors. 

- Save a copy of this R Notebook and rename it to psyc193r_hw07_(your pid).Rmd; e.g., psyc193r_hw07_A01234565.Rmd
- When you hit "Preview" in the menu bar, an html file will be generated
- You will have to upload this R Notebook and the html output; i.e., psyc193r_hw07_A01234567.Rmd & psyc193r_hw07_A01234567.nb.html using the submission link on the class website
- Codes are run in sequential order: the code appear earlier on may be needed for later parts. When you run some codes, make sure the gray boxes above it are also executed
- Unless otherwise specified, use alpha = 0.05 (two-tailed test) for hypothesis testing and confidence interval construction
- List out the information you have whenever needed to make sure they have not been overwritten in the boxes above

### Goals: 
The objective of this assignment is to practice doing multiple regression, including assumption checking, and calculate other statistics that are informative. 


## Background of the project: 
The [*National Center for Education Statistics*](https://nces.ed.gov/) conducts large scale survey on education issues every few years. In 2016, they conducted one on [*Parent and Family Involvement in Education (PFI)*](https://nces.ed.gov/nhes/dataproducts.asp#2016dp). As the name of the project suggests, they would like to single out factors related to Parent and Family Involvement that would affect a schoolchild's education. The data is available for the public.  
Jonas is deeply interested in education issues, especially at the K - 12 level (also apparently in college education, too, and that is why he is teaching the R class). In this project, hw would like to identify the factors that can promote a child's education. He believes that building a statistical model would help the policy makers and educators.  

In his analysis, he would like to identify predictors that would affect a child's school work (item segradeq in the survey). He picked 15 predictors from the list of 800+ variables that he believes have potential predictive power of a child's school work. He only includes schoolchildren in Grade 12.  

Information of the NCES survey:  
a. Background of the study & original data set: https://nces.ed.gov/nhes/dataproducts.asp#2016dp  
b. The Full questionnaire: https://nces.ed.gov/pubs2018/2018100.pdf  
c. The codebook: https://psyc193r.ucsd.edu/RNotebook/cbook_pfi_pu_jonas.pdf  
d. Data set: https://psyc193r.ucsd.edu/data/psyc193r_edu_data.csv  

You don't need the original data set for the study (it is too large). You will need "item d" to build your regression models, and items b & c for interpretation of the model. All the variables included in the data set are highlighted in the pdf file (item c) (nicely) by Jonas.  

# 1. Design of the Study
1.1 State the goal of this analysis. 

> 

1.2 What is the dependent / criterion variable of the model? State the name and the description of the variable (from the pdf documentation)

> 

## Loading data set
Load the data set, save it as "eduData"
```{r}
# load data
eduData = NA
```

## 2. Data structure and checking
Check data structure (long data frame):  

- Note where the dv is located  
- Remove participant Id ("basmid") and "grade", because they are not useful in our regression model  
- Save the reminding variables into eduData
```{r}
eduData = NA
```

Factorize / as.numeric variables
For simplicity, you can assume all variables are either ratio or interval
```{r}
# Hint/optional: use a for-loop

```

2.1 Check for outliers with histograms
```{r}
# create a 4 by 4 grid for ploting


# draw the histograms, label the x-axes of the histograms
# Hint/optional: use a for-loop


# switch back to a 1 by 1 grid
par(mfrow=c(1,1))

```

2.2 Check for outliers with scatterplots
```{r}
# create a 4 by 4 grid


# check for linearity, name the x- and y-axes
# Hint/optional: use a for-loop


# switch back to a 1 by 1 grid
par(mfrow=c(1,1))
```

2.3 Remove Outliers +/- 3sd of each column mean
```{r}
# remove the outliers
# Hint/optional: use a for-loop
for (i in 1:ncol(eduData)){
  # calculate the column mean
  colMean = NA
  
  # calculate the column sd
  colSD = NA
  
  # retain all values within +/-3sd in each column
  eduData = eduData[NA,]
}

# check for outliers again

```

## 3. Correlation Matrix and variable selection
3.1 Correlation Matrix: save it as corMatrix so that you can click and open it from the environment
```{r}
corMatrix = NA
```

3.2 Which variables(s) are you going to include / exclude (variable names only)
Explain your decision to include / exclude variables (if any)

> Include: 
Exclude: 

3.3 Remove Redundant variables, save the data to "eduReduced"
```{r}
eduReduced = NA

# build a new correlation matrixand check the you have other variables to exclude
# name it as "corMatrixReduced"
corMatrixReduced = NA
```

## 4. Building linear models
4.1 Build the intercept model and the full model
```{r}
# intercept model
edu.intercept.lm = NA

# summary of intercept model
summary(NA)

# full model
edu.full.lm = NA

# summary of full model
summary(NA)
```

4.2 Write out the equation for the full model. Describe the percentage of variance explained by the full model

> 

## 5. Model Comparison
Model comparison: Compare the full model and the intercept model, perform a statistical test and report the result. 
```{r}
anova(NA, NA)
```

> 

## 6. Post-fit checking
6.1 Residual Plot
```{r}
# predicted values (DV) from the IV's in the data
edu.full.fitted = NA
  
# calculate the residuals
edu.full.resid = NA

# plot residuals vs fitted values
plot(NA, NA)
```

6.2 Check for multicollinearity
```{r}
# import the relevant package
library(NA)

# calculate the variance inflation factors

```

6.3 Explain what the residual plot and the multicollinearity checking inform you about the model. 

> 

6.4 Overall, state the factors in the model (the name and the description), and describe whether the factors are positively or negatively correlated with the dependent variable. 

> 

## 7. Model Comparisons
After building a model with the predictor "seenjoy", Jonas would like to know see how much *more variance* can an addition predictor "hdhealth" explain. 
```{r}
# build the two models for comparison


# the more complicated models


# model comparison
anova(edu.seenjoy.lm, edu.joy.health.lm)
```

7.1 Are the two model nested within each other? Why? 

> 

7.2 Explain how much more variance the additional hdhealth predictor contributes, and whether the change is significant. 

> 

## 8. NOT GRADED: Automatic model selection
Forward model selection: Build two linear models
```{r}
# use the original eduData
# remove outliers 

# lm() for intercept model


# full model with all predictors (the "everything" model)


```

Use the step() function for forward model selection
```{r}
step(NA)
```

Verbally compare the full model you built and the model selected by the forward selection algorithm. 

> 