CDS 101 George Mason University Predicting House Prices R Lab
In this lab, you will learn how to use linear regression to build a predictive model of California house prices.
Predictive modeling
If you took CDS 101 in a previous semester, you should quickly review the new material added this semester about using models for prediction
Modeling for prediction (commonly referred to as machine learning) has different objectives to modeling for understanding/explanation.
In machine learning, we are most interested in whether our model will accurately predict data points in the future, and so we use a different workflow and different objectives to modeling for understanding.
About the Data
The dataset has already been randomly divided into training and test sets for you (in the train and test variables respectively).
Each row contains information about a district in California; each column contains data on the characteristics of housing in those districts.
VariableDescriptionlongitudethe E-W coordinate of the center of the districtlatitudethe N-S coordinate of the center of the districthousing_median_agethe median age of houses in the districttotal_roomsthe total number of rooms of houses in the districttotal_bedroomsthe total number of bedrooms of houses in the districtpopulationthe population of the districthouseholdsthe number of households in the districtmedian_incomemedian income of the districtmedian_house_valueresponse variable: the median value of houses in the districtocean_proximitycategorical variable that indicates if a district is inland, on the oceanfront, in a bay, etc.
Exercises
Install packages
You will need to install thelmvarpackage by running this command in the RStudio console:install.packages("lmvar")
- Take a look at the
traindataset in RStudio (remember that we should never look at our test dataset – otherwise we risk biasing our model).First, visualize the geographic distribution of the data by creating a scatter plot of thetraindata with longitude on the x-axis and latitude on the y-axis.To get an idea of the density of districts in different parts of the state, set thealphaparameter to a low-ish value (0.1 – 0.2).You should see that there are 3 rough clusters where districts are more dense:- A southern coastal region
- A northern coastal region
- A northern internal region.
- The northern internal region is the Central Valley, which contains a number of cities including Sacramento and Fresno.What are the other two regions with a large density of districts?
- Take your code from Exercise 1, and copy it into a new chunk. Add a parameter inside the
aesfunction to color the points by themedian_house_valuevariable.Where do most of the high house values seem to be located in California? - Visualize the relationships between the response variable
median_house_valueand the remaining continuous explanatory variables in thetraindataset (i.e. notlatitudeandlongitude, orocean_proximity).To do this:gatherthe remaining continuous explanatory variables.- Pipe the gathered data to ggplot and create scatter plots with the value column that you created in the
gatherfunction on the x-axis andmedian_house_valueon the y-axis. facet_wrapover the key column that you created in thegatherfunction. Use thescales="free"parameter to allow the axis limits to vary between facets.- Using your graph, answer the following questions:
- Which variable has the most obvious relationship with
median_house_value? - What value does the response variable go up to? Do you think this will cause a problem for making predictions with a linear model, and why?
- Visualize the relationship between the response variable
median_house_valueand the categorical variableocean_proximityby creating a box plot (the response variable should be on the y-axis, the categorical variable on the x-axis).Which category seems to have most of its districts distributed at low median house prices? - Using the
lmfunction, create a simple linear model wheremedian_house_valueis the response variable, andmedian_incomeis the explanatory variable. You will also need to supply the argumentsy = TRUEandx = TRUEto the model, as in this example:model_1 <- lm(… ~ …, data = …, y = TRUE, x = TRUE)
Remember to use the training dataset Calculate the k-fold cross validation error of this model using thecv.lmfunction (from thelmvarpackage):cv.lm(model_1, k =…)
Usek = 5. With a regression model, we typically calculate our error (i.e. our innaccuracy) with the root mean square error (RMSE, i.e. the square-root of the mean of the squared residuals). What is the validation RMSE of this model? - Repeat the modeling process to create a new linear model, but this time use all the explanatory variables (categorical and continuous, including latitude and longitude).Hint
As a short hand you can write.instead of writing out all the explanatory variables, i.e. `lm(y ~ .)
Calculate and report the cross-validation error as before.Which model performs best? - Using the best of the two models that you created in the previous two exercises, calculate its error at making predictions on the test dataset. (As a reminder, you should not have touched the
testdataframe before this question.)To do this we can use thermsefunction from themodelrpackage.The syntax ofrmseis:rmse(, )
Pass in the model that did best at cross validation (eithermodel_1ormodel_2[or whatever you called the model in Exercise 6]) and thetestdata.What is the root mean square error of this model on thetestdata. Is that better or worse than the error in cross validation (i.e. is your model more or less accurate on thetestdata)?