I picked two cities in Argentina, Junín is a city in the Rosario Argentina mission and is one of the cities where my dad served as a missionary for The Church of Jesus Christ of Latter-day Saints. Rio Cuarto is a city near Cordoba where a senior missionary I knew in my mission served as a young Elder.
I served my mission speaking Spanish and I’ve always felt connected to Argentina because of my dad and the people I’ve met from there. I’d love to study the weather for these cities in this beautiful country. I picked these two cities to compare for the year of 1996, when my father arrived in his mission.
library(mosaic)library(tidyverse)library(pander)library(DT) library(dplyr)library(leaflet)library(htmlwidgets)library(car)library(ggplot2)library(plotly)Argentina_Weather <-data.frame(NAME =c("JUNÍN","RIO CUARTO AREA DE MATERIAL"),LAT =c(-34.546, -33.085),LON =c(-60.931, -64.261))map <-leaflet(Argentina_Weather) %>%addTiles() %>%addMarkers(lng =~LON, lat =~LAT, popup =~paste(NAME)) %>%setView(lng =-63.5, lat =-32.5, zoom =5)map
Show the code
library(GSODR) #run: install.packages("GSODR") # to get the GSODR package. You'll need this package to pull in your weather data.load(system.file("extdata", "isd_history.rda", package ="GSODR"))
Show the code
#Run this in your console to see the Country Names you can pick from:View(isd_history)#Search "United States" in the search bar of the top-right corner of the data Viewer that pops up.#Or search for any other country you are interested in.#Goal, select the STNID (station ID) for two different weather stations. #For example, Rexburg is STNID == "726818-94194"#Once you have two STNID values selected, go to the next R-chunk.load(system.file("extdata", "isd_history.rda", package ="GSODR"))View(isd_history)
Show the code
#To see what columns mean, go here: https://cran.r-project.org/web/packages/GSODR/vignettes/GSODR.html#appendicesJunin <-get_GSOD(years =1996, station ="875480-99999")RioCuarto <-get_GSOD(years =1996, station ="874530-99999")weather <-rbind(Junin, RioCuarto)
Show the code
# Relabeling within NAME columnweather <- weather %>%select(STNID, NAME, COUNTRY_NAME, MIN, MAX, LATITUDE, LONGITUDE) %>%mutate(NAME =ifelse(NAME =="RIO CUARTO AREA DE MATERIAL", "RIO CUARTO", NAME))datatable(weather)
Linear Regression Analysis
\underbrace{Y_i}_{\text{Low Temp}} = \overbrace{\beta_0 + \beta_1 \underbrace{X_{i1}}_{\text{High Temp}}}^{\text{Junín Line}} + \overbrace{\beta_2 \underbrace{X_{i2}}_{\text{1 if Rio Cuarto}} + \beta_3 \underbrace{X_{i1} X_{i2}}_{\text{Interaction}}}^{\text{Rio Cuarto Adjustments to Line}} + \epsilon_i
where \epsilon_i\sim N(0,\sigma^2) and X_{i2} = 0 when the weather is in the city of Junín and X_{i2} = 1 when the weather is in Rio Cuarto. This forced 0, 1 encoding for X_{i2} produces the following models.
The hypotheses for this analysis will be testing the interaction of the minimum and maximum yearly temperature for 1996 of JUNÍNandRIO CUARTO`. We will be looking at the y-intercepts and the slopes of the two cities.
H_0: \beta_2 = 0 \\
H_a: \beta_2 \neq 0
A significant \beta_2 suggests a vertical shift in the line for Rio Cuarto compared to Junín, meaning the low temperatures differ systematically between the two cities regardless of high temperature.
H_0: \beta_3 = 0 \\
H_a: \beta_3 \neq 0
A significant \beta_3 indicates the slope of the Rio Cuarto line differs from Junín, meaning the effect of high temperature on low temperature varies between the cities.
Parameter
Effect
\beta_0
Y-intercept of the Model.
\beta_1
Controls the slope of the JUNÍN line.
\beta_2
Controls the change in y-intercept for the RIO CUARTO line in the model as compared to the y-intercept of the JUNÍN line.
\beta_3
Called the “interaction” term. Controls the change in the slope for the RIO CUARTO line in the model as compared to the slope of the JUNÍN line.
The Model
Show the code
argentina.lm <-lm(MIN ~ MAX + NAME + MAX:NAME, data=weather)n <-coef(argentina.lm)weather.plot <-ggplot(weather, aes(y=MIN, x=MAX, color=factor(`NAME`))) +geom_point(pch=16, bg="white") +stat_function(fun =function(x) n[1]+n[2]*x, color ="skyblue") +stat_function(fun =function(x) (n[1]+n[3]) + (n[2]+n[4])*x, color="goldenrod1")+scale_color_manual(name="Argentina Cities", values=c("skyblue","goldenrod1")) +labs(title="Weather Patterns of Argentina", x ="Maximum Tempurature (°C)", y ="Minimum Tempurature (°C)") +ylim(min(weather$MIN) -5, max(weather$MAX) +5)+theme_minimal()ggplotly(weather.plot)
¡Viva Argentina!
From the graph we can see that the slopes are very similar and that the placement of the dots indicating the tempurature in Celcius is moderately tight to the line for both cities. Let’s look at the actual data to determine what is signifigant here!
Show the code
pander(summary(argentina.lm))
Estimate
Std. Error
t value
Pr(>|t|)
(Intercept)
-9.442
0.6962
-13.56
1.622e-37
MAX
0.814
0.02823
28.84
1.641e-122
NAMERIO CUARTO
1.886
0.9897
1.905
0.05714
MAX:NAMERIO CUARTO
-0.05001
0.03988
-1.254
0.2102
Fitting linear model: MIN ~ MAX + NAME + MAX:NAME
Observations
Residual Std. Error
R^2
Adjusted R^2
731
3.503
0.6849
0.6836
With a p-value of 1.622e-37 we can see that \_beta_0 and we can reject the null hypothesis and accept that the intercept is significant (not 0).
We see from the (MAX) slope that for every 1°C increase in the maximum temperature, the minimum temperature for Junín increases by 0.814. This is the \beta_1 hypothesis. This hypothesis is signifigant with a p-value of 1.641e-122 < \alpha.
The \beta_2 hypothesis tests whether or not Rio Cuarto has a different y-intercept than Junín. The values of that hypothesis are shown in the Name:Rio Cuarto of the above linear model. The estimate is that the y-intercept of Rio Cuarto is 1.886 units higher than Junín’s.The p-value almost reaches significance 0.05714 > \alpha.Rio Cuarto might have a higher minimum temperature than Junín but there is not sufficient evidence to reject the null hypothesis.
\beta_3 tests whether the relationship between maximum and minimum temperatures (the slope) is different for Río Cuarto compared to Junín. The slope for Río Cuarto is slightly smaller (by 0.05001) than Junín’s slope. However, with a p-value of 0.210, this difference is not significant. This means there’s no strong evidence that the effect of maximum temperature differs between the two cities.
Results / Conclusion
First of all, I love the R^2 = 0.6849. This shows moderate strength in our linear model, the correlation between minimum and maximum temperatures is a positive, moderate relationship. We see here that as minimum daily temperatures rise so do the maximum for that day, this is cnsistent across both cities as demonstrated in the above scatter plot.
However for our hypotheses we are just barely not reaching statistical significance. \beta_2p-value is 0.05714 > \alpha. There is not sufficient difference in the y-intercepts between Junín and Rio Cuarto.
\beta_3p-value is 0.2102 > \alpha. This \beta_3 hypothesis shows there is not sufficient evidence to reject the null hypothesis that the slopes for Junín and Rio Cuarto are different.
There are 5 associated assumptions to the linear regression. All but #4 can be checked with these above plots.
Assumption 1: Linear Relationship Between X and Y
The linearity is good as shown in the residuals/fitted plot because the red line is almost flat.
Assumption 2: Normal Distribution of Error Terms
There are no obvious outliers as shown in the Q-Q plot. The data appears to be normally distributed.
Assumption 3: Constant Variance (X values)
In the residuals vs fitted plot there does not seem to be a pattern in the data, showing constant variance.
Assumption 4: Fixed X Values
It can be assumed that this assumption has been met, as The GSODR used to measure weather patterns across the world is expected to take accurate and consistent measurements that do not change over time.
Assumption 5: Independent Error Terms
The Residuals vs Order plot shows the chaos in the data, although there is almost a W, but that is natural according to the ups and downs in temperature throughout the year. Error terms do not appear to be correlated.
There is not overwhelming evidence shown in the above plots to discredit the whole linear regression analysis.
Credits
I would like to credit Brother Saunders and his Civic vs Corolla example analysis because I was completely lost without it.
I would also like to thank ChatGPT for helping explain to me how to interpret my intercepts and slopes.