Paper Summary
Share...

Direct link:

Goodness of Fit and Model Selection for Network Models

Sat, April 18, 8:15 to 9:45am, Marriott, Floor: Fifth Level, Scottsdale

Abstract

We propose multiple methods for determining the goodness of fit of a network model. Our primary models of interest are those with conditionally independent dyads (CID Models). This class of models includes the stochastic blockmodel (SBM), mixed membership stochastic blockmodel (MMSBM), and the latent space model (LSM). This set of models allows us to account for grouping structures and transitivity within a network. In addition, all three of these models have a parameter which must be chosen beforehand - the number of blocks in the stochastic blockmodels or the dimension of the latent space in the latent space model. The similarity of these models and the choice of an unknown parameter beforehand motivates the need for model selection and goodness of fit procedures for these models.

We first use cross-validation procedures to perform model selection to determine the best model to use, as well as the optimal parameter for the given model. We use two metrics to assess goodness of fit, the root mean-squared error (RMSE) and the area under the receiver operating characteristic curve (AUC).

Second, we apply the method of posterior predictive checking to determine goodness of fit. This method allows us to compare any statistic from our observed network to the distribution of that statistic under the posterior predictive distribution. Here we consider network statistics including the degree distribution and the geodesic distance distribution. This method allows us to both assess the goodness of fit of a single network model and compare the goodness of fit of multiple models.

We apply these methods to a simulated dataset from the three models stated above to show the usefulness of these methods when the true model is known. We then examine a network of real data to show how these methods can be applied in practice.

These methods are able to successfully determine the correct generative model as well as to select the number of parameters in a model. This work will allow scientists using social network models to better select models for use and provide a way of comparing competing models. Frequently these models can be used as a way of controlling for unseen effects or for discovering effects that may have been missed. By being able to accurately perform model selection, it is possible to create better controls and discover the properties of unseen network effects.

Authors