Show the code
# subset <- train %>%
# mutate_if(is.character, as.factor) %>%
# select(CurrentScore, everything())
#subset <- subset[, c(1, 68:88)]
#pairs(subset, panel=panel.smooth)Course MATH 425
Lexi Soelberg
Call: glm(formula = f90 ~ FinalExamCurrentScore + AnalysesCurrentScore +
SkillsQuizzesTotalFinalScore, family = binomial, data = train)
Coefficients:
(Intercept) FinalExamCurrentScore
-218.9591 0.5273
AnalysesCurrentScore SkillsQuizzesTotalFinalScore
1.2801 0.7147
Degrees of Freedom: 119 Total (i.e. Null); 116 Residual
Null Deviance: 165.8
Residual Deviance: 10.62 AIC: 18.62
| 1 |
|---|
| 0.9728 |
AnalysesCurrentScore SkillsQuizzesTotalFinalScore FinalExamCurrentScore
Min. : 9.09 Min. : 0.00 Min. : 0.00
1st Qu.: 79.55 1st Qu.: 90.66 1st Qu.: 55.00
Median : 91.97 Median :100.00 Median : 68.00
Mean : 83.71 Mean : 90.28 Mean : 62.45
3rd Qu.: 98.24 3rd Qu.:100.00 3rd Qu.: 76.00
Max. :100.00 Max. :100.00 Max. :100.00
This prediction takes a look at the chances of getting an A based on 3 specific criteria. If you got 100 on the skills quiz (which you have to do in order to get any points), at least 90% on your analyses, and get a final exam score of 68 or higher. To break that down a little further, You need to get 17/25 on the final.
The odds of this prediction are 35.76 to 1. This looks like:
Odds = 0.9728/1-0.9728 = 35.76
You are 36 times more likely to get an A if you meet the 90% analyses, 100% skills quizzes, and 68% final exam.
1
0.9728357
AnalysesCurrentScore SkillsQuizzesTotalFinalScore FinalExamCurrentScore
Min. : 9.09 Min. : 0.00 Min. : 0.00
1st Qu.: 79.55 1st Qu.: 90.66 1st Qu.: 55.00
Median : 91.97 Median :100.00 Median : 68.00
Mean : 83.71 Mean : 90.28 Mean : 62.45
3rd Qu.: 98.24 3rd Qu.:100.00 3rd Qu.: 76.00
Max. :100.00 Max. :100.00 Max. :100.00
Hosmer and Lemeshow goodness of fit (GOF) test
data: m325.glm$y, m325.glm$fit
X-squared = 0.10364, df = 2, p-value = 0.9495
keep <- sample(1:nrow(train), 90)
mytrain <- train[keep,]
mytest <- train[-keep,]
glm.test <- glm(f90 ~ FinalExamCurrentScore + AnalysesCurrentScore + SkillsQuizzesTotalFinalScore, data=train, family=binomial)
mypreds <- predict(glm.test, newdata=mytest, type="response")
mydecs <- ifelse(mypreds > 0.5, 1, 0)
cm <- table(mydecs, mytest$f90)
pcc <- (cm[1] + cm[4]) / sum(cm)The goodness of fit test is just under the table of variables. For this, we want to see a p-value greater than .05 and we definitely got that. A .95 p-value tells us that this regression is a good fit for the data. This is what we want. Now let’s take a look at this.
library(ggplot2)
library(plotly)
train$predicted_prob <- predict(m325.glm, newdata = train, type = "response")
final.ggplot <- ggplot(train, aes(x = FinalExamCurrentScore, y = f90)) +
geom_jitter(aes(color = predicted_prob), width = 1, height = 0.05, alpha = 0.6) +
geom_smooth(method = "glm", method.args = list(family = "binomial"), se = FALSE, color = "goldenrod") +
scale_color_gradient(low = "lightgoldenrod", high = "goldenrod") +
labs(title = "Final Exam Score vs. Probability of High Grade",
x = "Final Exam Score",
y = "Probability of Getting an A",
color = "Probability") +
theme_minimal()
ggplotly(final.ggplot, tooltip = c("x", "y"))Obviously the higher your exam score the better but we have a lot more variability in getting an A based on the final exam score alone. Compared to the other two variables, you potentially could get a 44 and get an A or get an 88 and not get an A. We need to add in more variables to see what predicts the overall grade best.
library(ggplot2)
library(plotly)
train$predicted_prob <- predict(m325.glm, newdata = train, type = "response")
analyses.ggplot <- ggplot(train, aes(x = AnalysesCurrentScore, y = f90)) +
geom_jitter(aes(color = predicted_prob), width = 1, height = 0.05, alpha = 0.6) +
geom_smooth(method = "glm", method.args = list(family = "binomial"), se = FALSE, color = "steelblue") +
scale_color_gradient(low = "lightsteelblue", high = "steelblue") +
labs(title = "Analyses Score vs. Probability of High Grade",
x = "Analyses Score",
y = "Probability of Getting an A",
color = "Probability") +
theme_minimal()
ggplotly(analyses.ggplot, tooltip = c("x", "y"))This variable is my favorite, the Analyses are worth 35% of your overall grade and you can resubmit until you get 15/15 on each one. If you do, your likelihood of getting an A is way higher as demonstrated in the graph. Later we will discuss the odds of your success based on these three variables.
library(ggplot2)
library(plotly)
train$predicted_prob <- predict(m325.glm, newdata = train, type = "response")
skillsquiz.ggplot <- ggplot(train, aes(x = SkillsQuizzesTotalFinalScore, y = f90)) +
geom_jitter(aes(color = predicted_prob), width = 1, height = 0.05, alpha = 0.6) +
geom_smooth(method = "glm", method.args = list(family = "binomial"), se = FALSE, color = "firebrick") +
scale_color_gradient(low = "lightcoral", high = "firebrick") +
labs(title = "Skills Quizzes Score vs. Probability of High Grade",
x = "Skills Quizzes Score",
y = "Probability of Getting an A",
color = "Probability") +
theme_minimal()
ggplotly(skillsquiz.ggplot, tooltip = c("x", "y"))It would be a very good predictor of your score if you get 100% on your skills quizzes. Just go until you can get them to 100%.
| Estimate | Std. Error | z value | Pr(>|z|) | |
|---|---|---|---|---|
| (Intercept) | -219 | 92.25 | -2.373 | 0.01762 |
| FinalExamCurrentScore | 0.5273 | 0.2177 | 2.422 | 0.01542 |
| AnalysesCurrentScore | 1.28 | 0.6176 | 2.073 | 0.03821 |
| SkillsQuizzesTotalFinalScore | 0.7147 | 0.2925 | 2.443 | 0.01455 |
(Dispersion parameter for binomial family taken to be 1 )
| Null deviance: | 165.82 on 119 degrees of freedom |
| Residual deviance: | 10.62 on 116 degrees of freedom |
[1] 18.6182
All p-values are less than .05. Each variable’s p-value contributes significantly to the outcome of the logistic regression.
The final exam score estimate indicates that for every one point increase in that test score increases the chance of getting an A by .5273. We can use this to interpret the odds of $e$0.5273 = 1.694. You have a 1.694 times greater chance of getting an A for each 1 point exam score increase.
The analyses scores estimates that $e$1.28 = 3.59. This is the biggest predictor of your overall grade. Your odds of getting an A are 3.59 times greater per 1 point increase in your analyses score.
The skills quizzes tells us that $e$0.7147 = 2.04. Your odds of getting an A on the skills quizzes alone increased by 2.04 for every 1 point increase in your skills quizzes scores.
Also I want to point out the AIC or the Akaike Information Cirteria. This shows that the lower the AIC score, the better this model is. 18 is a really low AIC so this model works really well.
P(Y_i = 1|\, FinalExamScore,AnalysesScore,SkillsQuizzes) = \frac{e^{\beta_0 + \beta_1 {FinalExamScore} + \beta_2 {AnalysesScore} + \beta_3 {SkillsQuizzes}}}{1+e^{\beta_0 + \beta_1 {FinalExamScore} + \beta_2 {AnalysesScore} + \beta_3 {SkillsQuizzes}}} = \pi_i H_0: \beta_1 = 0 \text{ (FinalExamScore has no effect)} \\ H_a: \beta_1 \neq 0 \text{ (FinalExamScore has an effect)}
H_0: \beta_2 = 0 \text{ (AnalysesScores has no effect)} \\ H_a: \beta_2 \neq 0 \text{ (AnalysesScores has an effect)}
H_0: \beta_3 = 0 \text{ (SkillsQuizzesScore has no effect)} \\ H_a: \beta_3 \neq 0 \text{ (SkillsQuizzesScore has an effect)}
set.seed(121)
n <- nrow(train)
keep <- sample(1:n, size = floor(0.7 * n)) #putSomeNumberHere that is about 60-70% of your data set's size)
mytrain <- train[keep, ]
mytest <- train[-keep, ]
train.glm <- glm(f90 ~ FinalExamCurrentScore + AnalysesCurrentScore + SkillsQuizzesTotalFinalScore,
data = mytrain, family = "binomial")
mypreds <- predict(train.glm, mytest, type = "response")
callit <- ifelse(mypreds > 0.9, 1, 0) #you can put whatever you want for the 0.9 value
cm <- table(mytest$f90, callit)
pcc <- (cm[1] + cm[4]) / sum(cm) #sum the correct answers you got then divide by the total number of guesses you made
print(cm) callit
0 1
0 24 0
1 0 12
[1] 1
This table tells us who actually passed the class who met the requirements, who didn’t who should have, and any false-positive and false negatives.
the 24 represent those who did get not an A based on the model.
The 12 are those who did pass the class according to the model
This model is 100% accurate because there isn’t anyone who got a 68 on the final, 100% on the skils quizzes, and 90 on the analyses who didn’t pass the class.
There also isn’t anyone who met the parameters of this test who did get an A.
As you can see there are some really big predictors for your overall grade in this course. The biggest are if you turn in your analyses and get at least 90% in them, if your Final Exam score is 68 or higher, and if your Skills Quizzes are 100% completed.
: )