Show the code
library(mosaic)
library(car)
library(pander)
library(DT)
library(plotly)
library(ggplot2)Course MATH 325
Lexi Soelberg
Many teachers and other educators are interested in understanding how to best deliver new content to students. In general, they have two choices of how to do this.
A study was performed to determine whether the Meshed or Before approaches to delivering content had any positive benefits on memory recall.
Individuals were seated at a computer and shown a list of words. Words appeared on the screen one at a time, for two seconds each, until all words had been shown (40 total). After all words were shown, they were required to perform a few two-digit mathematical additions (like 15 + 25) for 15 seconds to avoid immediate memory recall of the words. They were then asked to write down as many of the 40 words as they could remember. They were given a maximum of 5.3 minutes to recall words.
The process of showing words and recalling words was repeated four times with the same list of words each time (four chances to get it right). The presentation of the first trial was the same for all treatment conditions. However, trials 2, 3, and 4 were slightly different for each treatment condition.
The SFR group (the control group) stands for Standard Free Recall. In all four trials the same list of 40 words was presented, in a random order each time.
The Before group also used the same 40 words during each trial. However, any words that were correctly recalled in a previous trial were presented first, or before the words that were not recalled in the last trial. After all the correct words were presented in random order, the non-recalled words were presented in a random order.
The Meshed group also used the same 40 words during each trial. However, words that were correctly recalled in a previous trial were alternated with a missed word during the next presentation order.
The data records the number of correctly recalled words (out of the 40 possible) from the fourth trial. Results were obtained for 30 students, 10 in each of the three treatment groups: SFR, Before, and Meshed.
The above background gives some great detailed information on a study taken of 30 participants tasked with remembering words. These participants were placed in 3 groups of reviewed before (Before), reviewed throughout (Meshed), and not reviewed (SFR).
I am really interested to see if the Meshed approach, which I feel may have been beneficial to me during my schooling, is more effective than the Before approach. I want to see if there is a significance in the medians of these two groups. This is my 1st section of testing.
I am also curious to see if any form of test reviewing is proven more effective than leaving the participant to their own devices. So, my 2nd section of testing will be viewing a Control group (SFR) and a Test group (Before & Meshed).
:::
The 1st hypothesis is that the Meshed group will remember more than the remember more words than the Before group.
H_0: \text{difference in the medians} = 0
H_a: \text{"Before" median} \ne \text{"Meshed" median}
Significance level is
\alpha = 0.05
The 2nd hypothesis is that both the Meshed and Before tested groups together will score differently than the SFR control group.
H_0: \text{difference in the medians} = 0
H_a: \text{Control Group Median} < \text{Test Group Median}
Significance level is
\alpha = 0.05
There are two groups of 20 different participants, 10 per test condition as shown in the table below.
The box plot below depicts the 5 number summary of the two sample sizes. The numbers are also listed in the numerical summary just below.
plot_ly(FriendlyH1,
y = ~CorrectWords,
x = ~as.factor(TestCondition),
type = "box",
color = ~as.factor(TestCondition),
colors = c("royalblue4", "red"),
line = list(color = "gray21", width = 2),
marker = list(line = list(color = "gray21", width = 1))) %>%
layout(title = "Effect of Test Conditions on Memory Recall",
xaxis = list(title = "Experimental Test Groups"),
yaxis = list(title = "Correct Words (out of 40)"))We can see that the median of the Meshed group is a couple points lower, at 36.5 and the Before group’s median is 39.
It seems that the Before group has a higher success rate of words remembered than the Meshed group.
However, to see the actual difference, let’s take a look at the significance using a Wilcoxon Rank Sum Test.
FriendlyH1 %>%
group_by(TestCondition) %>%
summarise(
Min = min(CorrectWords, na.rm = TRUE),
Q1 = quantile(CorrectWords, 0.25, na.rm = TRUE),
Med = median(CorrectWords, na.rm = TRUE),
Q3 = quantile(CorrectWords, 0.75, na.rm = TRUE),
Max = max(CorrectWords, na.rm = TRUE),
Average = mean(CorrectWords, na.rm = TRUE),
StDev = sd(CorrectWords, na.rm = TRUE),
SampleSize = n()
) %>%
pander()| TestCondition | Min | Q1 | Med | Q3 | Max | Average | StDev | SampleSize |
|---|---|---|---|---|---|---|---|---|
| Before | 24 | 37.25 | 39 | 39.75 | 40 | 36.6 | 5.337 | 10 |
| Meshed | 30 | 36 | 36.5 | 38.75 | 40 | 36.6 | 3.026 | 10 |
The Wilcoxon Rank Sum (Mann-Whitney) Test to see the differences in the medians is listed below. Running a two sided test to see if the Meshed group has a lower median than the Before group.
| Test statistic | P value | Alternative hypothesis |
|---|---|---|
| 62 | 0.3831 | two.sided |
Test statistic is W = 62 with a p-value of 0.378>\alpha . The results of the Wilcoxon test are not significant. We fail to reject the null hypothesis.
A boxplot chart and numerical summary of the control group (SFR) of 10 participants and the 20 who received a test condition (either Before or Meshed).
plot_ly(FriendlyH2,
y = ~CorrectWords,
x = ~as.factor(TestCondition),
type = "box",
color = ~as.factor(TestCondition),
colors = c("olivedrab4", "mediumpurple4"),
line = list(color = "gray21", width = 2),
marker = list(line = list(color = "gray21", width = 1))) %>%
layout(title = "Memory Recall: Control vs Test",
xaxis = list(title = "Benefits of Test Conditions on Memory Scores"),
yaxis = list(title = "Correct Words (out of 40)"))We see a lot of spread in the Control (SFR) group despite the sample size being half that of the Test (Before and Meshed) group.
The correct words out of 40 for the Test group is much more successful in word recall with a median at 38 than the control group who has a median of 27. That’s 11 words less!
With this plot we can see a pretty clear difference in the medians but let’s test the actual significance of that with a Wilcoxon test.
FriendlyH2 %>%
group_by(TestCondition) %>%
summarise(
Min = min(CorrectWords, na.rm = TRUE),
Q1 = quantile(CorrectWords, 0.25, na.rm = TRUE),
Med = median(CorrectWords, na.rm = TRUE),
Q3 = quantile(CorrectWords, 0.75, na.rm = TRUE),
Max = max(CorrectWords, na.rm = TRUE),
Average = mean(CorrectWords, na.rm = TRUE),
StDev = sd(CorrectWords, na.rm = TRUE),
SampleSize = n()
) %>%
pander()| TestCondition | Min | Q1 | Med | Q3 | Max | Average | StDev | SampleSize |
|---|---|---|---|---|---|---|---|---|
| Control | 21 | 25 | 27 | 38.5 | 39 | 30.3 | 7.334 | 10 |
| Test | 24 | 36 | 38 | 39.25 | 40 | 36.6 | 4.223 | 20 |
The Wilcoxon Rank Sum (Mann-Whitney) one sided less than test to see if the Test group has a significant difference over the in the medians of the Control group.
| Test statistic | P value | Alternative hypothesis |
|---|---|---|
| 51.5 | 0.0153 * | less |
Test statistic is W = 51.5 with a p-value of 0.01644<\alpha . The results of the Wilcoxon test are significant. We reject the null hypothesis.
The Wilcoxon Test showed no significant difference in the medians of the Before and Meshed groups. The medians are both high at 39 and 36.5 words remembered. We failed to reject the null hypothesis that there would be a significance difference in the medians of the groups.
We did reject the null hypothesis when testing the difference in medians of the combined Before and Meshed “Test” group and the SFR “Control” group. The p-value was below alpha and the medians were respectively 38 and 27 words remembered.
We see no significant median difference between the test groups Before and Meshed however we see a significant difference in the medians for the control SFR and the tested groups.
We do see significant evidence to prove that any form of test review is powerful in increasing the amount of words that people will remember. If we leave people without a review there is a much wider range of words they will remember. If we provide an information review while introducing new content, the probability of a student or employee retaining and recalling information is much higher.
Therefore, if we want to see better memory retention as we do when comparing these three groups we can employ methods of review, either before starting the new content or have the old intertwined with the new.
This can be applied and tested in many different settings to see if the same or similar results can be found.
: )