Individuals belong to geographically based social networks that include an individual’s ties (family, friends, professional contacts, etc.) in nearby and distant places (Acedo et al., 2017). A set of mapped ties comprises a spatial distribution, i.e., a unique “fingerprint” of geolocated social contacts (e.g., two friends in Rome, six family members in New York, a co-worker in Milan, etc.). This distribution is a natural part of social life, as humans have been “traveling, wandering and friending friends and kin seemingly forever” (Hampton and Wellman, 2003, p. 284).
While studies have estimated social network size (ex. Hill and Dunbar, 2003) and structure (examples abound), relatively less is known about the spatial distribution of an individual’s social ties. For example: in how many places (e.g., cities and towns) does the average person have a social tie? Does this number expand or contract for the individual over time? What is the average distance at which an individual’s ties reside? In addition, longitudinal social network studies have examined contact strength, network benefits, individuals comprising the networks, and networks through life events such as marriage, divorce, retirement, and widowhood (Ertel et al., 2009). Yet, the shift in network geography over time is not as clear.
The answers to these questions can tell us more about an individual’s level of access to social and spatial resources, as having trusted nearby contacts or trusted contacts in particular locations-of-interest leads to the acquisition of a rich variety of resources. In short, it is not just where you live and what you can access nearby, but also what you can access through your social network ties. We situate these questions within a larger framework of the theory of relational geography, which affirms that social networks (and social capital) are spatialized, and that relationships are embedded within a larger spatial system that includes institutions and economic practices (Bathelt and Glückler, 2003). Relational geography theory also asserts that economic actors acquire distant resources across regions and countries in order to reap the benefits of innovation, diffusion, collaboration, collective learning processes and knowledge generation (Bathelt and Glückler, 2005). Whether through new tie acquisition, inter-tie relationship changes, or changes in the spatial configuration of a network, time often changes the variety, diversity, and robustness of an individual’s personal Rolodex of people and places. As an individual moves, their spatial social network will add this new cost of visiting or maintaining communication, seek local alternatives for social life, or restructure in another way.
These life events and strategies to cope with the change of moving and changing geographic social networks are common, but the formal concepts are still abstract and challenging to measure. This serves as a motivation for quantitatively discretizing the spatial distribution of many individuals’ contacts. Mapping these ties over time can help codify complex structures in order to discern the cause and effect of social network life changes and to articulate commonalities of social network geographies across the population. Making this implicit, personalized knowledge more explicit and our analysis more procedural can show how humans may have a common experience. For instance, mapping contacts provides a way to find the typical individual’s balance of nearby and distant ties, and distinguishing this balance is important because nearby and distant ties offer different types of support. Nearby ties can provide help in emergency situations and yield increased power in local governance (Aldrich and Meyer, 2014) while distant friends are helpful when an agent plans a visit or seeks advice (Stafford, 2005).
To further motivate the mapping of geographic social networks, there are also practical travel and health benefits to enumerating contacts’ distances. Mapping tie locales can reveal how far people must travel to see their ties and to calculate the cost, or cost-prohibitive nature, of accessing social support (Carrasco et al., 2009). It can help measure the extent to which an individual is in jeopardy of isolation from known social ties, which can lead to loneliness and adverse health effects (Cacioppo and Patrick, 2008). Maps of contacts can also be used to estimate whether the average person is privy to information about many places (Dabbs et al., 1998; Golledge, 2002) and has opportunities to visit an array of cultural locales (Pultar and Raubal, 2009). Geographic theory supports this notion, through an understanding that contacts share their “insider” knowledge of places.
Engaging with the questions of geographic social network configuration is becoming increasingly facile with digital data sources. In particular, social media applications that allow individuals to track the geographic locations of their contacts provide instant information about the spatial expanse of a user’s contacts (where they move to, their vacation locations, etc.). Most commonly, social media acts as a repository in which users can look up locations from friends’ profiles (on platforms such as Facebook). However, more recently, the locations of large sets of friends has been mapped in real time by apps such as Snapchat’s Snap Maps, allowing users (most often teens and young adults) to peruse a changing heat map annotated with cartoon images of their friends (Juhász and Hochmair, 2018). This mass location sharing creates a “landscape” of many contacts mapped at one time (Næss, 2018). Users most often opt to share location with friends and family, as well as nearby contacts in order to ease the logistics of meeting up (Consolvo et al., 2005). The digital collection of these data encourages better studies not only on of social network size and structure, but also on the distribution of social tie locations.
Motivated by a conceptual systems approach of relational geography and enhanced offerings of online social network data, we perform a case study to measure the geographic distribution of social ties. The objective of this work is to take advantage of digital social network data, and present a proof-of-concept study that illustrates how geographic information systems (GIS) mapping and modeling can help analyze the underlying structure and organization of an individual’s relational geography. In our approach, we map the Facebook friends of a group of 20 volunteer egos. We use two locations for each alter: the location where the alter lived when they met the ego (pre-period) and the alter’s location at the time of the study (post-period). We perform a descriptive analysis of this data set to answer the following research questions:
RQ1. In how many cities and countries do individuals have ties?
RQ2. Do individuals have ties in more locations and at farther distances over time?
RQ3. Do tie locations deviate from the expected distances at which an individual is likely to have ties?
RQ4. Do different types of alter groups (e.g., family, school friends, neighbors, etc.) tend to spread (disperse) over time?
This study illustrates our argument that an individual’s set of places they can access through ties will change over time, as will their ability to access geographic resources through ties. Instead of thinking about the individual’s geographic network as highly personalized and difficult to codify, the case study shows the technology needed to create a working model of geolocated social networks, what measurements and statistics are pertinent, and what visualizations are helpful. As a result, the network provides new insights about how social ties are spread across geographic space. This case study is a non-conventional approach to social network data analysis, as it combines the GIS analysis and measures of spatial network expanse at the individual level. By replicating the procedure described in our case study, others can implement their own networks and collect their own data for future studies.
In the following section, we review major concepts mobilized in this work, and briefly describe our case study approach. We next describe the Facebook data collection process, network measurements, GIS methods and summary statistics. We then describe the results of the 20 networks’ spatial distributions and their change over time. Finally, we contextualize our findings in a wider body of literature, discuss study limitations and conclude.
Background: concepts revisited in this work
This section explores previous research describing contacts’ expected and actual distribution over geographic space.
Estimating connectivity
The gravity model is a classic method for estimating a “baseline” of spatial interaction across locales. Interaction can represent number of phone calls, migrants, commuters, number of relationships or other flows that connect two locales. This model computes an expected value of interaction (I ij ) as the product of a source (i) and target (j) city’s respective populations, divided by the Euclidian distance between i and j raised to an exponent β (called the coefficient of friction, most frequently parameterized as 2) (Dodd, 1950; Reilly, 1953) (Eq. (1)). Distance can also be redefined as travel time or another cost factor (de Smith, 2004). This value is multiplied by a constant (K) to better reflect the actual magnitudes of the interaction values.
This equation has been used to estimate inter-city travel (Ben-Akiva and Lerman, 1985) and retail patronage (Huff, 1963) and has been reparametrized by adding demographic data such as income at a destination city (Greenwood, 1985). The gravity model provides an expectation from which to measure whether real-world inter-city connections deviate from this expected value. Recently, actual interaction data have been used to re-calibrate the coefficient of friction from the default value of two (Krings et al., 2009), and determine where sets of cities are over- or under-connecting in comparison to the gravity estimates (Dugundji et al., 2011; Takhteyev et al., 2012).
Next, survey data and large data sets describing tie distribution in geographic space have shown that distance affects the likelihood of relationships.
Relationships and distance
Distance plays a role in the likelihood of relationships and relationship maintenance. The likelihood of interaction (or friendship) tends to decrease with each increment of distance between individuals, a longstanding economic geography concept known as distance decay. Recently, large data sets from GPS traces, Location-based social networks (LBSNs), and call data records have confirmed that the probability of having a social tie in a certain location decreases exponentially with distance to that location (Blondel et al., 2008; Lambiotte et al., 2008; Leskovec and Horvitz, 2008; Liben-Nowell et al., 2005; Onnela et al., 2011; Preciado et al., 2012; Scellato et al., 2011), although rates can vary depending on data source (Spiro et al., 2016). Each different function reveals the likelihood of an individual having a contact at certain distances. For instance, in Belgium, this probability drops off significantly at 40 km from the individual’s locale (Lambiotte et al., 2008). In studies of Twitter data, it was shown that 34% of Twitter friends who follow one another live within 25 miles while only 18% of users who do not follow each other, but “mentioned” one another live within that radius (McGee et al., 2011), and that Twitter friendships are more likely to follow flight patterns than Euclidean distance (Takhteyev et al., 2012). Moreover, the average distance of an individual’s contacts has been shown to vary by hometown: a recent Facebook study showed that residents in isolated US areas (such as Eastern Kentucky) had up to 82% of Facebook friends living within 50 miles, whereas residents of in the Western US had fewer than 43% of friends living within the same radius (Bailey et al., 2018).
In addition to big data harnessed from online social networks, studies using standard survey approaches confirm that the accessibility of social support is an important descriptor of a community. For example, in his study of the “spatial dimensions of personal relations”, Fischer (1982) finds that 26% of semi-rural residents’ relatives lived within a 5 minutes’ drive, but only 15% of urbanites’ relatives lived within a 5 minutes’ drive. The General Social Survey reports that 27% of respondents live within a 15-minute drive of their mother, and 12% live 12+ hours away (Smith et al. n.d.). Similar studies of long distance relationships find that 66% of elderly parents live within 30 minutes of an adult child (Lye, 1996; Stafford, 2005). These types of studies are particularly helpful because they distinguish different types of ties.
Survey data and distance decay functions provide rules-of-thumb for where relationship probability declines. Yet, they should be paired with map-based approaches that can communicate a richer portrait of relationships. Our case study aims to dig deeper into personal distributions for individuals by listing individual cities that egos interact with and describing how these dynamics change over time.
Case study, data and methods
Case study
In total, 20 volunteers downloaded their friendship networks (totaling 8,549 friends) from http://www.facebook.com. This platform was chosen because Facebook ties have been shown to mimic real-life acquaintances (Mayer and Puller, 2008) and because the platform has been widely used in human social behavior research in messaging patterns (Golder et al., 2007), social connections (Ellison et al., 2007), and cultural preferences (Gross and Acquisti, 2005; Lampe et al., 2006; Pempek et al., 2009).
Each volunteer (ego) annotated each of their friends (alters) with an accompanying home location at two time periods (the location from when the ego and alter met, and a current location). Individuals’ ties were mapped in geographic space and analyzed both individually and as a combined group. This study is equipped to respond to our four research questions because it specifically maps individuals’ social ties, calculates distance between ties, associates ties with different cities, and measures tie dispersion over time.
Participants
Volunteers were enrolled at The Pennsylvania State University (Penn State) as graduate or undergraduate students at the time of the study. All volunteers resided in State College, PA, a small college town in the northeastern USA, and were recruited to participate in the study through a seminar course project. Participants were required to have a Facebook account with associated “friends” (i.e., alters) and be at least 18 years old. The group ranged in age between 20–40 years old, and included 10 women and 10 men. It was comprised of 17 white and three Asian respondents, and four participants were non-native English speakers. Further information about egos was concealed for privacy reasons.
Social network acquisition
Within Facebook, each volunteer used the NameGenWeb application (as described in Hogan, 2011) to download their social network as a graph of their friends and friend inter-connections (i.e., if the ego’s friends are friends with one another). NameGenWeb has been used to show that agents in dense networks influence one another’s emotional status (Lin and Qiu, 2012), to find structurally- and semantically-related groups of nodes (Cruz et al., 2013), and to identify social groups and clusters (Brooks et al., 2014).
NameGenWeb created an undirected network of friends, where an ego with k friends can render a list of (k × (k − 1))/2 possible connections (Jackson, 2008). Each data set was downloaded as an edgelist of mutual friendships and later converted to other network data structures (i.e., Graph Markup Language) based on software input requirements. Data were stored within the Neo4j graph database structure. The NameGenWeb application became unavailable toward the end of 2014, and a similar application, NetVizz (Rieder, 2013) was used for four volunteers. As of January 1, 2015, the ability to download a friendship network was no longer a feature of Facebook or its applications. Yet, independent researchers still provide these functionalities. For example, the Lost Circles team of University of Konstanz in Germany offers a free plug-in for Facebook network visualization and download (https://lostcircles.com/).
Network metrics
We calculated standard network descriptors, including diameter, density, and average clustering coefficient using the “igraph” library (Csárdi and Nepusz, 2006) in the R statistical computing environment.
Geolocation
Next, volunteers geocoded the locations of each of their friends at two time steps (using recollection and assistance from social media, such as Facebook profiles): the alter’s home location when the relationship began (pre-city) and the alter’s home location at the time of the study (post-city). A few alters met in “cyberspace” or had unknown locations and were removed because, although virtual spaces are geographic (Chen et al., 2013), we required metrics of distance change, i.e., two distinct geographic points, over time. Volunteers assigned longitude and latitude coordinates to alters by geolocating cities in Google Maps, Mapquest or Esri’s global gazetteer.
Each ego and alter were assigned an ID number to preserve anonymity. After anonymizing the data, the geographic coordinates were mapped within the ArcMap GIS software environment. These coordinates were spatially joined to existing shapefiles (i.e., spatial data files) of core-based statistical areas (CBSA) (metropolitan areas) or, if in a rural area, counties. Coordinates outside the USA were assigned to global administrative areas (GADM). As a result, each alter was assigned to a standardized urban center or a US county.
Spatial calculations
We then calculated the Euclidean distance between the ego’s current location and the locations of each alter’s pre-city and post-city, using the Great Circle method from the R “geosphere” package based on the WGS84 ellipsoid (Hijmans et al., 2017). To examine how social networks spread over time, we found the standard distance (i.e., the standard deviation of the longest linear axis of a point pattern) of each ego’s alters for both time steps. We determined the mean centers, that is, the average center of friend distributions, and reported the extent of the shift in mean centers over time in the ArcMap environment.
Ground truthing
The gravity model (as described earlier) is used in this study to predict the places where the egos are likely to have ties. We compared the locations of the alters’ cities to the gravity model for US cities only. First, we found the product of the population of State College and each alter’s locale (using 2000 US census population counts). We next computed the Euclidean distance from the town’s centroid to all other cities’ centroids.
Community detection
Facebook friendship networks tend to naturally cluster around different areas of a person’s life (Hogan et al., 2007) with significant clustering within universities (Lewis et al., 2008). To find clusters, or modules, within each network, we used the Louvain method for community detection (i.e., modularity calculation) (Blondel et al., 2008) within the Gephi environment (Bastian et al., 2009). To avoid double-counting alters, we removed 44 “overlapping” alters, i.e., nodes that are friends with more than one ego.
For each modularity group, egos identified an institution that the modularity group represents and the year of the group’s inception. We provided a crowdsourced set of labels, including university, secondary school, place-based cultural groups, and non-residential gatherings, such as conferences and vacation travel. Volunteers were invited to label groups using the aforementioned examples but were also able to choose their own labels. A modularity group required three members to be considered, and 90% of egos annotated their groups with a classification. We assume that each unique cluster within a Facebook network represents a specific social context (Brooks et al., 2014).
Results
Network description
All 20 egos had a total of 8,549 alters ranging from 123 to 772 per ego (Table 1, Figure 1). The number of edges connecting the nodes ranged from 784 to 12,667. The average degree of friends ranged from 6.9 to 37.7, indicating that some alters had many shared friends in the same network while others did not. Density is defined as the number of edges divided by the total possible edges in a network and is computed as the number of actual edges (e) that exist over the number of possible edges (k × (k − 1))/2). Average density ranged from 0.024 to 0.172 (Table 1, Figure 1). Diameter, defined as the longest shortest-path that connects two nodes, ranged from 6 to 14. However, not all nodes were reachable, due to some isolates and disconnected components. Lower diameter values imply more “friends of friends” who are also friends (e.g., triads), whereas higher values imply more “chains” of friends who do not have common friends (Jackson, 2008). The average clustering coefficient, ranging from 0.16 to 0.74, represents the probability of a node being part of a network triangle. Dense networks with high clustering coefficients and low diameters can be considered tight-knit.
Table 1.
Summary statistics and geographic distribution of alters for each ego’s network.
| Network characteristics | Geographic characteristics | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Agent | Nodes | Edges | Average degree | Diameter | Density | Clustering coefficient | Modules; modularity | Average distance to friends (post-period) | Standard distance pre-period | Standard distance post-period | Change in standard distance | Change of mean center |
| A | 376 | 2679 | 12.3 | 10 | 0.038 | 0.36 | 4; 0.61 | 431 | 1468 | 2392 | 924 | 192 |
| B | 336 | 2533 | 14.5 | 9 | 0.045 | 0.53 | 13; 0.67 | 7827 | 5570 | 5691 | 121 | 667 |
| C | 601 | 7392 | 23.5 | 11 | 0.041 | 0.49 | 15; 0.62 | 277 | 1592 | 2395 | 803 | 383 |
| D | 370 | 3686 | 19.4 | 9 | 0.054 | 0.57 | 11; 0.71 | 2567 | 4377 | 3686 | -691 | 597 |
| E | 573 | 9833 | 33.8 | 10 | 0.06 | 0.47 | 12; 0.51 | 1409 | 2754 | 3727 | 973 | 1873 |
| F | 232 | 2706 | 22.3 | 5 | 0.10 | 0.53 | 12; 0.37 | 1174 | 2917 | 3903 | 986 | 455 |
| G | 361 | 2989 | 14.8 | 12 | 0.04 | 0.54 | 12; 0.71 | 215 | 887 | 2283 | 1396 | 368 |
| H | 437 | 6002 | 27.2 | 8 | 0.06 | 0.46 | 10; 0.51 | 1185 | 627 | 1267 | 640 | 188 |
| I | 403 | 5022 | 24.7 | 8 | 0.062 | 0.43 | 9; 0.52 | 687 | 3087 | 2383 | –704 | 518 |
| J | 437 | 2763 | 11.5 | 10 | 0.029 | 0.52 | 13; 0.78 | 1619 | 3726 | 4157 | 431 | 409 |
| K | 639 | 9784 | 30.5 | 7 | 0.048 | 0.33 | 10; 0.46 | 712 | 2534 | 2893 | 359 | 575 |
| L | 188 | 1002 | 10.4 | 10 | 0.057 | 0.64 | 6; 0.74 | 697 | 3047 | 2605 | –442 | 129 |
| M | 201 | 784 | 6.9 | 10 | 0.039 | 0.16 | 16; 0.46 | 8522 | 5727 | 5610 | –117 | 325 |
| N | 772 | 7143 | 18.2 | 11 | 0.024 | 0.44 | 16; 0.76 | 529 | 3274 | 3357 | 83 | 292 |
| O | 123 | 1028 | 15.8 | 9 | 0.137 | 0.74 | 9; 0.40 | 6701 | 3588 | 5564 | 1976 | 448 |
| P | 659 | 11,708 | 34.9 | 8 | 0.054 | 0.50 | 14; 0.52 | 1508 | 3726 | 3552 | –174 | 1150 |
| Q | 727 | 12,667 | 34.5 | 9 | 0.048 | 0.53 | 10; 0.69 | 492 | 1792 | 2951 | 1159 | 520 |
| R | 227 | 4412 | 37.7 | 6 | 0.172 | 0.55 | 10; 0.23 | 176 | 49 | 1658 | 1609 | 196 |
| S | 408 | 7639 | 37.0 | 7 | 0.092 | 0.52 | 9; 0.52 | 626 | 1693 | 2929 | 1236 | 441 |
| T | 566 | 4157 | 14.3 | 14 | 0.026 | 0.53 | 15; 0.75 | 1148 | 1156 | 2167 | 1011 | 560 |
| Pre-city | Alters | Egos | Post-city | Alters | Egos |
|---|---|---|---|---|---|
| State College, PA | 1384 | 17 | State College, PA | 797 | 18 |
| Washington–Arlington–Alexandria, DC–VA–MD–WV | 803 | 12 | Washington–Arlington–Alexandria, DC–VA–MD–WV | 721 | 18 |
| Seattle–Tacoma–Bellevue, WA | 310 | 9 | New York–Northern New Jersey–Long Island, NY–NJ–PA | 346 | 17 |
| Poughkeepsie–Newburgh–Middletown, NY | 306 | 3 | Boston–Cambridge–Quincy, MA–NH | 276 | 17 |
| Chicago–Naperville–Joliet, IL–IN–WI | 288 | 10 | Charlotte–Gastonia–Concord, NC–SC | 239 | 8 |
| Boston–Cambridge–Quincy, MA–NH | 245 | 13 | Seattle–Tacoma–Bellevue, WA | 231 | 17 |
| Knoxville, TN | 222 | 5 | Chicago–Naperville–Joliet, IL–IN–WI | 224 | 16 |
| Cedar Rapids, IA | 205 | 2 | Philadelphia–Camden–Wilmington, PA–NJ–DE–MD | 207 | 15 |
| Charlotte–Gastonia–Concord, NC–SC | 199 | 4 | Atlanta–Sandy Springs–Marietta, GA | 193 | 12 |
| Atlanta–Sandy Springs–Marietta, GA | 196 | 7 | San Francisco–Oakland–Fremont, CA | 183 | 14 |
| San Francisco–Oakland–Fremont, CA | 186 | 8 | Los Angeles–Long Beach–Santa Ana, CA | 133 | 17 |
| Williamsport, PA | 180 | 3 | Dallas–Fort Worth–Arlington, TX | 132 | 15 |
| Austin–Round Rock, TX | 178 | 5 | Albany–Schenectady–Troy, NY | 129 | 6 |
| Philadelphia–Camden–Wilmington, PA–NJ–DE–MD | 173 | 9 | Knoxville, TN | 126 | 5 |
| Dallas–Fort Worth–Arlington, TX | 159 | 6 | Austin–Round Rock, TX | 120 | 15 |
| Appleton, WI | 125 | 2 | Williamsport, PA | 107 | 4 |
| Riverside–San Bernardino–Ontario, CA | 104 | 5 | Cedar Rapids, IA | 92 | 2 |
| New York–Northern New Jersey–Long Island, NY–NJ–PA | 63 | 16 | Pittsburgh, PA | 91 | 15 |
| Harrisburg–Carlisle, PA | 61 | 3 | Augusta–Richmond County, GA–SC | 82 | 5 |
| San Diego–Carlsbad–San Marcos, CA | 45 | 7 | Riverside–San Bernardino–Ontario, CA | 76 | 7 |
| Group type | Modules | Avg. standard distance before; after (km) | Difference in standard distance (km) | Avg. years since inception |
|---|---|---|---|---|
| University | 36 | 2298; 3810 | 1512 | 8.5 |
| Professional | 21 | 2468; 3927 | 1460 | 6.6 |
| Secondary education | 20 | 527; 2441 | 1914 | 14.2 |
| Place-based cultural group | 16 | 1437; 2224 | 788 | 10.7 |
| Family | 15 | 1678; 2459 | 781 | 22.2 |
| Non-residential gathering | 14 | 2898; 4865 | 1967 | 6.5 |
| Non-place-based cultural group | 8 | 2463; 4267 | 1805 | 9.5 |
| Other | 3 | 4429; 5410 | 981 | 5.5 |





