13
Political Ideology and Racial Preferences in Online Dating Ashton Anderson, a Sharad Goel, b Gregory Huber, c Neil Malhotra, a Duncan J. Watts b a) Stanford University; b) Microsoft Research; c) Yale University Abstract: What explains the relative persistence of same-race romantic relationships? One possible explanation is structural–this phenomenon could reflect the fact that social interactions are already stratified along racial lines–while another attributes these patterns to individual-level preferences. We present novel evidence from an online dating community involving more than 250,000 people in the United States about the frequency with which individuals both express a preference for same-race romantic partners and act to choose same-race partners. Prior work suggests that political ideology is an important correlate of conservative attitudes about race in the United States, and we find that conservatives, including both men and women and blacks and whites, are much more likely than liberals to state a preference for same-race partners. Further, conservatives are not simply more selective in general; they are specifically selective with regard to race. Do these stated preferences predict real behaviors? In general, we find that stated preferences are a strong predictor of a behavioral preference for same-race partners, and that this pattern persists across ideological groups. At the same time, both men and women of all political persuasions act as if they prefer same-race relationships even when they claim not to. As a result, the gap between conservatives and liberals in revealed same-race preferences, while still substantial, is not as pronounced as their stated attitudes would suggest. We conclude by discussing some implications of our findings for the broader issues of racial homogamy and segregation. Keywords: political ideology; racial preferences; homogamy; homophily; online dating; computational social science Editor(s): Jesper Sørensen, Stephen L. Morgan; Received: September 18, 2013; Accepted: October 10, 2013; Published: February 18, 2014 Citation: Anderson, Ashton, Sharad Goel, Gregory Huber, Neil Malhotra, and Duncan J. Watts. 2014. “Political Ideology and Racial Preferences in Online Dating.” Sociological Science 1: 28-40. DOI: 10.15195/v1.a3 Copyright: c 2014 Anderson et al.. This open-access article has been published and distributed under a Creative Commons Attribution License, which allows unrestricted use, distribution and reproduction, in any form, as long as the original author and source have been credited. A lthough interracial marriages have been stea- dily increasing in number over time (Fu and Heaton 2008), racial homogamy—the dispropor- tionate prevalence of same-race romantic partners (Fu and Heaton 2008; Schoen and Wooldredge 1989; Blackwell and Lichter 2004)—is a persis- tent phenomenon. Among all newlyweds in 2008, for example, only 9 percent of whites and 16 percent of blacks married someone whose race was different than their own (Passel, Wang, and Taylor 2010). Such racial homogamy is conse- quential both sociologically and economically. To the extent that information, resources, and op- portunities are structured by one’s social network (Coleman 1988; Portes 1998), the homogeneity of marital and family ties is likely to affect both individual-level outcomes, such as educational achievement, occupation, and income (Campbell, Marsden, and Hurlbert 1986; Grodsky and Pager 2001), and collective phenomena, such as racial inequality, segregation, and polarization (Baldas- sarri and Bearman 2007). Population-level statistics indicate the extent of racial homogamy in society. They do not, how- ever, reveal its underlying causes. In particular, there are at least two possible—and qualitatively different—contributing factors. First, relation- ship partners may be selected from a pool of racially similar candidates because of the preex- isting homogeneity of an individual’s social envi- ronment (Feld 1981), including the individual’s educational institution, profession, and friends. Second, individuals may simply prefer same-race relationships for reasons as diverse as religious beliefs, social or cultural expectations, a sense of shared identity, or race-related physical at- tributes. Although these two mechanisms, one structural and the other preference based, are the- oretically distinct, differentiating between them empirically can be problematic. As has been pre- viously pointed out (McPherson, Smith-Lovin, and Cook 2001), cross-sectional network data are equally consistent with either mechanism; and although recent work utilizing longitudinal net- sociological science | www.sociologicalscience.com 28 February 2014 | Volume 1

PoliticalIdeologyandRacialPreferencesinOnlineDating 1/february... · We present novel evidence from an online ... and the user’s preferences for a potential part- ... normalized

Embed Size (px)

Citation preview

Political Ideology and Racial Preferences in Online DatingAshton Anderson,a Sharad Goel,b Gregory Huber,c Neil Malhotra,a Duncan J. Wattsba) Stanford University; b) Microsoft Research; c) Yale University

Abstract: What explains the relative persistence of same-race romantic relationships? One possible explanation isstructural–this phenomenon could reflect the fact that social interactions are already stratified along racial lines–whileanother attributes these patterns to individual-level preferences. We present novel evidence from an online datingcommunity involving more than 250,000 people in the United States about the frequency with which individuals bothexpress a preference for same-race romantic partners and act to choose same-race partners. Prior work suggests thatpolitical ideology is an important correlate of conservative attitudes about race in the United States, and we find thatconservatives, including both men and women and blacks and whites, are much more likely than liberals to state apreference for same-race partners. Further, conservatives are not simply more selective in general; they are specificallyselective with regard to race. Do these stated preferences predict real behaviors? In general, we find that stated preferencesare a strong predictor of a behavioral preference for same-race partners, and that this pattern persists across ideologicalgroups. At the same time, both men and women of all political persuasions act as if they prefer same-race relationshipseven when they claim not to. As a result, the gap between conservatives and liberals in revealed same-race preferences,while still substantial, is not as pronounced as their stated attitudes would suggest. We conclude by discussing someimplications of our findings for the broader issues of racial homogamy and segregation.

Keywords: political ideology; racial preferences; homogamy; homophily; online dating; computational social scienceEditor(s): Jesper Sørensen, Stephen L. Morgan; Received: September 18, 2013; Accepted: October 10, 2013; Published: February 18, 2014

Citation: Anderson, Ashton, Sharad Goel, Gregory Huber, Neil Malhotra, and Duncan J. Watts. 2014. “Political Ideology and Racial Preferences in OnlineDating.” Sociological Science 1: 28-40. DOI: 10.15195/v1.a3

Copyright: c© 2014 Anderson et al.. This open-access article has been published and distributed under a Creative Commons Attribution License, whichallows unrestricted use, distribution and reproduction, in any form, as long as the original author and source have been credited.

Although interracial marriages have been stea-dily increasing in number over time (Fu and

Heaton 2008), racial homogamy—the dispropor-tionate prevalence of same-race romantic partners(Fu and Heaton 2008; Schoen and Wooldredge1989; Blackwell and Lichter 2004)—is a persis-tent phenomenon. Among all newlyweds in 2008,for example, only 9 percent of whites and 16percent of blacks married someone whose racewas different than their own (Passel, Wang, andTaylor 2010). Such racial homogamy is conse-quential both sociologically and economically. Tothe extent that information, resources, and op-portunities are structured by one’s social network(Coleman 1988; Portes 1998), the homogeneityof marital and family ties is likely to affect bothindividual-level outcomes, such as educationalachievement, occupation, and income (Campbell,Marsden, and Hurlbert 1986; Grodsky and Pager2001), and collective phenomena, such as racialinequality, segregation, and polarization (Baldas-sarri and Bearman 2007).

Population-level statistics indicate the extentof racial homogamy in society. They do not, how-ever, reveal its underlying causes. In particular,there are at least two possible—and qualitativelydifferent—contributing factors. First, relation-ship partners may be selected from a pool ofracially similar candidates because of the preex-isting homogeneity of an individual’s social envi-ronment (Feld 1981), including the individual’seducational institution, profession, and friends.Second, individuals may simply prefer same-racerelationships for reasons as diverse as religiousbeliefs, social or cultural expectations, a senseof shared identity, or race-related physical at-tributes. Although these two mechanisms, onestructural and the other preference based, are the-oretically distinct, differentiating between themempirically can be problematic. As has been pre-viously pointed out (McPherson, Smith-Lovin,and Cook 2001), cross-sectional network data areequally consistent with either mechanism; andalthough recent work utilizing longitudinal net-

sociological science | www.sociologicalscience.com 28 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

work data has found that observed homophilyon both racial (Wimmer and Lewis 2010) andnonracial attributes (Kossinets and Watts 2009)is likely due to a combination of structural andpsychological forces, these studies were not de-signed to measure individual preferences directly.When used to elicit attitudes about race, more-over, traditional survey tools are thought to besusceptible to social desirability bias (Krosnick1999; Crowne and Marlowe 1998); that is, respon-dents seeking not to appear racist to interviewersand researchers may not be honest about theirracial preferences and attitudes. Accordingly, es-timates of racial preferences may be biased down-ward. A second bias, potentially compoundingthe first, is that individuals often have inaccuratebeliefs about their own preferences (Gilbert 2006;Bernard et al. 1984; Nisbett and Wilson 1977).Thus even survey tools that are designed to cor-rect for social desirability bias may underestimatepreferences for same-race partners for the simplereason that respondents believe themselves to bemore race-blind than they actually are.

We take a novel approach to measuring same-race preferences for romantic relationships, lever-aging a unique data set compiled from an onlinedating website. Although limited in some re-spects, online data are increasingly being used toshed light on social scientific questions in general(Lazer et al. 2009) and offer several advantages foraddressing this topic in particular. First, our dataset is considerably larger and more diverse thanthe data sets of previous related studies, com-prising more than 250,000 individuals of widelyvarying demographic and socioeconomic status,from hundreds of U.S. cities and all regions ofthe country. Second, in contrast to traditionalsurveys, the data were collected in a natural set-ting where individuals were less susceptible tosocial pressures to appeal to an interviewer; hencestated preferences are more likely to reflect actualattitudes. Third, we can account for the entirepool of available online romantic partners in ageographic area and thereby control for the possi-bility that homogamy arises because of differencesin available dating pools. Fourth, because individ-uals on the site provide a substantial amount ofinformation about themselves, we can investigatehow same-race preferences vary with other fac-tors, such as income and education, and therebyaccount for many possible confounding variables.

Finally, because we observe which other per-sonal profiles individuals select to view, we canaugment stated attitudes with a behavioral mea-sure of same-race preference, thus allowing us tomitigate biases in self-reported preferences. Im-portantly, our data allow us to assess these pref-erences at one of the earliest stages of selection:when a user decides whether to view a candidate’sfull profile after seeing the candidate’s photo andbrief biographical information. We can thereforeunderstand how race affects initial screening de-cisions in the dating environment, the point atwhich individuals rule out many potential datingpartners from further consideration. Prior work,by contrast, has focused on later-stage selectioneffects—examining who individuals choose to con-tact from among those whose full profiles theyview—and therefore potentially misses the effectof race and other factors during the initial win-nowing of the dating pool. Hitsch et al. (2010),for example, find that at this later stage, men’sobserved behavior is in line with their stated pref-erences, in sharp contrast to our own finding thateven those who do not state a racial preferencedisplay a strong tendency to prefer same-racecandidates early in the selection process.

We focus our attention on three particulardemographic attributes: sex, race, and politicalideology. Given that the outcome variable ofinterest is a preference for same-race romanticpartners of the opposite sex, our focus on sex andrace is self-explanatory. Our focus on politicalideology, meanwhile, is motivated by a significantbody of research that shows political conservatismis correlated with a host of attitudes that mayreflect low desire to form personal relationshipswith people of different races: explicitly statedtraditional and symbolic racism, implicit preju-dice, affect, and xenophobia (Sidanius, Pratto,and Bobo 1996; Federico and Sidanius 2002; Feld-man and Huddy 2005; Nail, Harton, and Decker2003; Whitley 1999). Relatively little work, how-ever, has directly assessed how preferences forsame-race relationships vary by political orienta-tion and whether those differences in expressedpreferences predict real behavior.

DataOur data were assembled from user activity logsfor a popular online dating website in which users

sociological science | www.sociologicalscience.com 29 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

could view personal profiles and send messages toother members of the site. To protect the privacyof individuals, all data were anonymized prior toanalysis. We collected a complete snapshot ofactivity on the site during a two-month period(October–November 2009). Member profiles con-sisted of a picture, a short piece of freeform textin which the member could describe himself orherself, and answers to various multiple-choicequestions about both the user’s characteristicsand the user’s preferences for a potential part-ner. For example, for the question “What is yourethnicity?” users could respond with “white,”“black,” “Asian,” “Hispanic,” or “other.” For eachsuch multiple-choice question, users could alsoindicate a subset of answers they would preferfrom a potential mate and the strength of thatpreference. For example, they could state thatthey would prefer potential partners to have an-swered the ethnicity question with either “white”or “Asian” and could list this as either a “nice-to-have" preference or a “must-have" preference.Users could also specify that any answer to thequestion is acceptable. Finally, users were freeto answer as few or as many questions as theywished. Political ideology was asked on a 5-pointresponse scale ranging from 1 (very liberal) to5 (very conservative). We restrict our analysisto users with relatively complete demographicprofiles—those reporting age, sex, location, eth-nicity, education, income, political ideology, mar-ital status, religion, height, body type, drinkinghabits, smoking habits, presence of children, anddesire for more children—and who also explic-itly express a preference, or lack of a preference,for a potential partner’s race. We also restrictour attention to whites and blacks because His-panics and Asians are sufficiently heterogeneouscategories that “same-race” preference may havelittle meaning. Finally, we limit our sample toheterosexuals. After these restrictions, our dataset consists of 251,701 users for whom we haveboth profile data and a record of which profilesthey chose to view in full.

As shown in Figure 1, the sample of users westudy comprises a diverse set of individuals interms of age, education, income, geography, andpolitical ideology. Although we make no claimthat our sample is representative of the generalU.S. dating population (which itself differs sys-tematically from the overall U.S. population), it

does exhibit significant mass over a broad rangeof relevant demographics, including, for example,both younger (18–29) and older (60+) users, ed-ucation levels ranging from “some high school”to “postgraduate,” annual income ranging fromless than $25,000 to more than $150,000, substan-tial populations from all regions of the country,and a variety of political affiliations, where mostusers describe themselves as “middle of the road.”One respect in which our sample is clearly notrepresentative of the general dating population,however, is that men are highly overrepresented1

(75 percent)—a disparity that has been noted inother, smaller samples of online dating communi-ties from the same era (Hitsch et al. 2010).

Results

Stated Preferences

We begin by examining explicitly stated same-race preferences, where we classify a user as ex-pressing such a preference only if the user’s de-clared partner race set matches the user’s ownself-declared ethnicity (i.e., the only race the userprefers is the user’s own). For the reasons out-lined in the introduction, we are mainly interestedin three key demographic attributes associatedwith differences in same-race preferences and be-haviors: sex, race, and political ideology. Figure 2shows the stated same-race preference distribu-tion jointly over these attributes. Specifically,Figure 2 (top) shows the observed fraction of in-dividuals of different gender and race who expressat least a “nice-to-have” preference (solid lines)and separately a “must-have” preference (dottedlines).

Although these raw figures have the benefitof being easy to interpret, they are potentiallyconfounded by other variables, such as incomeand education, that are correlated with race andideology. To correct for these potential confounds,we estimate the likelihood a user qi states a “nice-to-have” or “must-have” same-race preference viatwo separate logistic regression models. Specifi-

1As we discuss later, this disparity very likely con-tributes to greater overall selectivity by women relativeto men; however, it should not affect our other results,which control for gender.

sociological science | www.sociologicalscience.com 30 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

Sex Age Race Education Income Region Political Idelogy

0%

25%

50%

75%

100%

Male

Fem

ale

18−2

9

30−3

9

40−4

9

50−5

960

+W

hite

Black

Some

High S

choo

l

High S

choo

l Gra

d

Some

Colleg

e

Colleg

e Gra

d

Post−

Gradu

ate

Less

Tha

n $2

4,99

9

$25,

000

To $

34,9

99

$35,

000

To $

49,9

99

$50,

000

To $

74,9

99

$75,

000

To $

99,9

99

$100

,000

To $

149,

999

Mor

e Tha

n $1

50,0

00

North

east

Midw

est

South

Wes

t

Very L

ibera

l

Liber

al

Midd

le Of T

he R

oad

Conse

rvat

ive

Very C

onse

rvat

ive

Figure 1: Demographic composition of individuals in study sample.

cally, we fit models of the form

Pr [qi states “nice-to-have” preference] =logit−1(βnice ·Xi)

Pr [qi states “must-have” preference] =logit−1(βmust ·Xi),

where Xi is a vector of user qi’s attributes, βniceand βmust are vectors of corresponding regressioncoefficients, and logit−1(x) = ex/(1 + ex). Thesemodels adjust for every demographic attributeusers specify: age, sex, height, ethnicity, educa-tion, income, geography, political affiliation, mar-ital status, religion, body type, drinking habits,smoking habits, presence of children, and desirefor more children.

The majority of these attributes are categori-cal, in which case we use indicator variables foreach category to allow for the greatest amountof model flexibility. Two exceptions are age andheight, which are modeled by including age andage squared and height and height squared inthe attribute vector X. Age and height are alsonormalized to have mean 0 and standard devi-ation 1. To adjust for geography, we includethe population density of the user’s declared zipcode, an indicator variable specifying whether theuser lives in an urban area (defined as having atleast 1,000 people per square mile), a categoricalvariable for geographic region (Northwest, West,

South, and Northeast), and the fraction of peoplein the user’s zip code who are the same race asthe user. We also include two separate continu-ous variables specifying the number of (nonrace)nice-to-have and must-have preferences. Theselatter two variables capture the user’s general se-lectivity, aside from any race preferences. Finally,given the substantial differences in the numberof men and women active on these sites and thatheterosexual dating sites are two-sided marketsstratified by gender, all of these attributes areinteracted with sex.

In sum, the structural form of the “nice-to-have” stated preference model (omitting the indi-vidual subscript qi for clarity) is

Pr [qi states “nice-to-have” race preference] =

logit−1(βrace×sex + βpolitical×sex

+βeducation×sex + βage×sex + βage2×sex

+βheight×sex + βheight2×sex + . . .).

(1)An analogous model is used for “must-have” pref-erences.

Table A1 in the supplement lists fitted coeffi-cient values for key variables of interest. Giventhe complexity of these models and the largenumber of interactions, the regression coefficientscan be difficult to interpret on their own. We

sociological science | www.sociologicalscience.com 31 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

Male Female

0%

25%

50%

75%

100%

0%

25%

50%

75%

100%

White

Black

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Like

lihoo

d to

sta

te s

ame−

race

pre

fere

nce

At least nice−to−have

Must−have

Male Female

● ●●

●●

● ●●

●●

● ● ● ● ●

● ● ● ● ●

●● ●

●● ●

●● ●

● ● ●

0%

25%

50%

75%

100%

0%

25%

50%

75%

100%

White

Black

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Like

lihoo

d to

sta

te s

ame−

race

pre

fere

nce

At least nice−to−have

Must−have

Figure 2: Estimated probability of stating a same-race preference by sex, race, and political ideology.Top: Unadjusted sample proportions. Bottom: Estimates derived from a model that controls forall other available demographic attributes. The size of the dots in the top panel corresponds tothe number of individuals for each data point, while in the bottom panel, the bars are 95 percentconfidence intervals.

sociological science | www.sociologicalscience.com 32 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

therefore use our fitted models to estimate thelikelihood of stating nice-to-have and must-havepreferences across various demographic groups,holding other factors constant. In particular, af-ter constructing a “typical individual”—based onthe median or modal value of the empirical distri-bution for each attribute—we then vary sex, race,and political affiliation, allowing us to isolate theeffects of each of these factors.2 Figure 2 (bot-tom) shows these model-adjusted estimates. Thesimilarity between the raw and model-adjustedestimates indicates that the patterns we observeare indeed reflective of race, gender, and politicalideology and not simply driven by the correlationbetween racial preferences and other demographiccharacteristics.

Perhaps most strikingly, Figure 2 illustratesthat women are substantially more likely thanmen to express both weak and strong same-racepreferences. Specifically, more than half (52 per-cent) of white, politically moderate women ex-press at least a “nice-to-have” same-race prefer-ence, with 27 percent explicitly stating that asame-race partner is a “must-have”; by compar-ison, 21 percent of white, politically moderatemen state having a “nice-to-have” same-race pref-erence, and 10 percent report having a “must-have” preference.3 Similar differences betweenwomen and men are apparent among blacks.

Figure 2 further indicates a strong associationbetween political ideology and stated same-racepreferences. Though the effect is apparent acrossboth sexes, it is particularly salient for women:conservative white women are about 30 percentmore likely to express a preference for same-racepartners than their liberal counterparts (56 per-cent vs. 43 percent). Likewise, conservative blackwomen are substantially more likely to state asame-race preference than liberal black women(42 percent vs. 30 percent). The percentage ofwhite men with a stated same-race preference is24 percent and 18 percent for conservatives and

2We separately construct “typical” men and women:height, number of profile views, and number of statedpreferences are set to the gender-specific medians; allother attributes are set to the median values over theentire sample.

3Owing to the extremely large sample size, most dif-ferences between percentages are highly statistically sig-nificant at conventional levels. Accordingly, we focus onthe substantive effects and only note when differences arenot significant.

liberals, respectively. We even find this patternfor black men, the group with the lowest propen-sity to state a same-race preference: 9 percentof conservative black men state a same-race pref-erence, compared to 7 percent of liberal blackmen.4

Although the tendency of political conserva-tives to state same-race preferences at higherrates than political liberals is striking, the under-lying cause remains unclear. One possibility isthat conservatives are more selective in general—on a variety of traits—and that their same-racepreferences are simply a manifestation of this ten-dency. Indeed, as we have already noted, womenin our population are heavily outnumbered bymen, and men also tend to be far more activein contacting or approaching women than the re-verse. For both these reasons, it is plausible thatwomen, seeking to exploit their “market power,”or simply to reduce their cognitive load, may electto state more preferences, including a same-racepreference. Possibly, therefore, the observed ef-fect of political ideology can also be explained interms of overall selectivity, not selectivity on racespecifically.

We note that our model includes the numberof nonrace preferences that each individual states,so that if conservatives were, on average, sim-ply more likely to express any preference, theseestimates account for this simple difference inselectivity. Our model also includes a measureof differences in the racial composition of thedating pool in different areas, which mitigatesagainst the possibility that liberals are simplyconcentrated in areas where expressing a racialpreference is less necessary because of greaterracial homogeneity in the dating pool. Never-theless, to further investigate the possibility ofdifferences in overall choosiness by ideology andgender, we measure selectivity by examining thenumber of attributes other than race (e.g., height,income, education, smoking habits, body type)for which users express preferences, again brokendown by sex, race, and political ideology. Analo-gous to our analysis framework, we fit regressionmodels to estimate selectivity as a function of

4Even though we observe similar effects for both whitesand blacks, the estimated effects for blacks are harder togeneralize from our sample to the population at largebecause minorities primarily interested in dating withintheir own race group may seek out racially specific datingsites.

sociological science | www.sociologicalscience.com 33 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

individual attributes. Given that the outcomevariable of interest (i.e., the number of nonracestated preferences) is integer valued, we use Pois-son regression. Again omitting the individualsubscript qi for clarity, the form of the models is

Number of nonrace stated “nice-to-have”and “must-have” preferences =

Poisson(exp

(βrace×sex + βpolitical×sex

+βeducation×sex + βage×sex + βage2×sex+βheight×sex + βheight2×sex + . . .

)).

(2)We also separately fit a model to estimate onlythe number of “must-have” preferences.

Selected model coefficients are listed in TableA2 in the supplement. As before, we also plotmodel estimates for a prototypical individual,varying sex, race, and political ideology. As ex-pected, Figure 3 shows that women—both whiteand black—state “must-have” or “nice-to-have"preferences more than men. Women state such apreference for approximately 60 percent of theseattributes, compared to only about 45 percentfor men. Conservatives, however, state no morepreferences on average than liberals (white men,47 percent vs. 45 percent; white women, 62 per-cent vs. 59 percent). The observed propensityof conservatives to state same-race preferences,therefore, is not attributable to some more gen-eral selectivity but rather is specific to race.

Revealed PreferencesOur results thus far are based entirely on self-reported racial preferences. But are they accu-rate proxies for behavior? Prior research suggeststhat a potential problem with using self-reporteddata is that they may not reflect people’s truepreferences. It could be the case, for example,that liberals have exactly the same preferencesas conservatives for same-race relationships butare not inclined to state, or even acknowledge,those views (Sniderman and Carmines 1997). Weaddress this issue by measuring a user’s revealedpreferences: the relative likelihood that a userviews a candidate’s profile given that the candi-date is the same race versus a different race asthe querier himself or herself. Specifically, forqueriers qi searching for candidates ci whose pro-files are available on the dating site, we estimatethe relative risk RRR of selecting racially congru-ent profiles. Our use of relative risk, motivated

by its application in epidemiology, is defined for-mally as

RRR =

Pr[qi views profile of ci | qi is the same-race as ci]Pr[qi views profile of ci | qi is a different race than ci]

.

A relative risk RRR greater than 1 means thatthe querier is disproportionately inclined to viewsame-race candidates, and hence exhibits a same-race preference, whereas RRR = 1 indicates theabsence of such a preference. (RRR < 1 wouldindicate a preference for partners of a differentrace.)

To estimate RRR, we must address three com-plications. First, we need to specify exactly whichquerier–candidate pairs to consider. One couldnaively consider all possible pairs of users to bepotential matches, but geographic constraintsalone suggest that choice is ill-suited for our anal-ysis. In response, we restrict the candidate setto members of the opposite sex living within 25miles of the querier and who meet the querier’sstated age requirements. These constraints—sex,age, and geography—are ostensibly the most im-portant in initial evaluations of candidates, andall queriers are required to specify these. Thegoal of these constraints was to winnow the fullset of dyads to a manageable level while notputting unneeded restrictions on observing po-tential matches. Accordingly, we did not restrictdyads based on, for example, education, althoughwe do control for such other variables in the anal-yses. We note, however, that on the actual site,users were only presented with profiles of userswho satisfied their age, sex, and geography con-straints as well as their must-have preferences.Therefore, the “broad pool” of candidates we con-sider includes some individuals who are, by de-fault, not presented to the querier by the website,although users could still find these candidatesby using the site’s search functionality. To verifythat our results are not being driven by the site’sdesign, we repeat our analysis for a “narrow pool”of candidates who also meet a querier’s must-havepreferences. As shown in the appendix, these twopools of candidates lead to similar results, andso we focus on the broad pool for our primaryanalysis.

A second issue is that even with appropriatelydefined candidate sets, the number of querier–candidate pairs is large (in the tens of millions).

sociological science | www.sociologicalscience.com 34 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

Male Female

● ● ●● ●

● ● ●● ●

● ● ●● ●

● ● ●● ●

● ● ●●

● ● ●● ●

● ● ●●

● ● ●●

0%

25%

50%

75%

100%

0%

25%

50%

75%

100%

White

Black

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Fra

ctio

n of

non

−ra

ce a

ttrib

utes

Total nice− and must−have prefs

Only must−have prefs

Figure 3: Nonrace selectivity: Estimated fraction of nonrace attributes for which user states preferences,by sex, race, and political ideology. Bars indicate 95 percent confidence intervals. See Table A2 formodel estimates.

We therefore employ a case control design (Kingand Zeng 2001), first constructing a much smallersubset of instances in which a user viewed the pro-file of another member, and then augmenting thisset with randomly selected instances in which aquerier did not elect to view a candidate’s profile.The size of this latter component is chosen to beapproximately three times as large as the former.This selection procedure clearly curbs our abilityto estimate either the numerator or denominatorin RRR. However, as has been observed in thestatistics and epidemiology literature, the oddsratio can still be estimated from such a sample:

ROR =ps/(1− ps)pd/(1− pd)

, (3)

where

ps = Pr[qi views profile of ci|qi is the same race as ci]

pd = Pr[qi views profile of ci|qi is a different race than ci]

Moreover, given the large number of profiles, psand pd are both relatively small, and so

ROR = RRR

(1− pd1− ps

)≈ RRR.

That is, the odds ratio ROR—which we can effi-ciently estimate—approximates RRR.

Finally, as in our analysis of stated prefer-ences, we would like to control for potential con-founding variables and isolate the effects of cer-tain key factors of interest, namely, sex, race, andpolitical affiliation. To do so, observe that thelog odds ratio can be written as

log(ROR) =

logit(Pr[qi views profile of ci |qi is the same race as ci]) −

logit(Pr[qi views profile of ci |qi is a different race than ci])

(4)

where logit(x) = log(x/(1 − x)). We thus esti-mate log(ROR) by first fitting a logistic regression

sociological science | www.sociologicalscience.com 35 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

model that predicts whether any given querier qiviews the profile of a candidate ci, and then ex-amining the difference between model estimateswhen the candidate ci is assumed to be the samerace versus a different race than the querier qi,holding all other traits constant. By then vary-ing the demographic attributes of qi and ci, thisapproach allows us to investigate how revealedpreferences change across subpopulations.

This logistic regression model includes sep-arate terms for the demographic attributes ofboth the querier and the candidate as well asjoint querier–candidate features (i.e., interactionterms between the querier’s features and thoseof the candidate). In particular, analogous tothe stated preferences model, for both querierand candidate, we include age, height, education,income, religion, body type, employment status,drinking habits, smoking habits, existence of chil-dren, desire to have more children, marital status,population density in the querier’s zip code, frac-tion of population in the querier’s zip code whoare of the same ethnicity, whether or not the zipcode is classified as urban, and the number ofnonrace nice-to-have and must-have preferences,all of which are interacted with the sex of thequerier. The joint querier–candidate attributesindicate whether the users have the same politi-cal affiliation, level of education, marital status,smoking habits, drinking habits, religion, income,body type, employment status, existence of chil-dren, and desire to have more children; we alsoinclude the (continuous) distance between thequerier and the candidate. Finally, our modelsinclude three additional interaction terms: weinteract the querier’s sex, race, and political af-filiation with a variable indicating whether thequerier–candidate pair is of the same race.

Figure 4 plots model estimates for ROR asa function of political ideology, broken down bygender and race.5 Consistent with our findingsregarding stated preferences, the lines in Figure 4slope upward for all gender–race groups, indicat-ing that more conservative individuals exhibit agreater behavioral tendency to select same-racepartners. On average, for example, conservative

5Because ROR is the difference of model estimates (asdescribed in Eq. [4]), it depends only on the coefficientsfor the three characteristics (race, gender, and ideology)that are interacted with the variable indicating whetherthe querier is of the same race as the candidate.

white men have an estimated odds ratio of 3.3,compared to 2.6 for liberal white men. Likewise,conservative white women have an odds ratio of3.6, compared to 2.9 for liberal white women.Similar patterns exist for all other gender andrace groupings. Interestingly, however, our earlierfinding that women are more likely than men tostate a same-race preference is not present in thebehavioral data; both black and white men arejust as likely to reveal a same-race preference astheir female counterparts.

Figure 5 casts additional light on these con-trasting findings—that the effect of political ideol-ogy is consistent across stated and revealed pref-erences but that the effect of gender is starklydifferent—by further breaking down revealed pref-erences by the stated preference of each gender–race–ideology category.6 (The vertical axis ofFigure 5 is expanded substantially relative to Fig-ure 4; it now ranges up to 256 rather than just 6,and is scaled exponentially.) We note three mainresults. First, we find that individuals of bothgenders and races who explicitly state not havinga same-race preference do in fact exhibit a sub-stantial tendency to favor same-race partners. Inquantitative terms, ROR for these “no preference”individuals (dashed line) ranges between 2 and 3across demographic groups, meaning that evenafter controlling for a host of other factors, theyare two to three times more likely to select a can-didate of the same race (all of these differencesare statistically distinguishable from 1).

Second, we find that when a same-race prefer-ence is stated, it is highly informative of behavior,particularly for men. In Figure 5, that is, thesolid line is the estimated behaviorally revealedpreference for individuals stating a “must-have”same-race preference, whereas the dotted line isthe same quantity for those expressing “nice-to-have” preferences. Apart from white women, forwhom these two lines cross and point estimatesare statistically indistinguishable from one an-other, the solid line is consistently above the dot-ted line, which is consistently above the dashed

6To generate these estimates, we modify the revealedpreferences model to now include three three-way inter-action terms: we interact the querier’s sex, race, andpolitical affiliation with both the querier’s stated racepreference (i.e., no preference, nice-to-have, or must-have)and a variable indicating whether the querier–candidatepair is of the same race. Fitted values for these interactionterms are listed in Table A3 in the supplement.

sociological science | www.sociologicalscience.com 36 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

Male Female

●●

●●

●●

●●

2

4

6

2

4

6

White

Black

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Rev

eale

d sa

me−

race

pre

fere

nce

Figure 4: Estimated revealed preferences for same-race partners. Bars indicate 95 percent confidenceintervals.

line. Thus stating “must-have” is associated withchoosing same-race candidates at higher ratesrelative to those stating “nice-to-have,” which isassociated with choosing same-race candidatesat higher rates than those stating no same-racepreference. In terms of magnitude, for whitemen, those who say “must-have” are about 20times more likely to select same-race candidatesthan different-race candidates, and those saying“nice-to-have” are about 6 times more likely to se-lect same-race candidates. Both effects are muchlarger than the estimates for the “no preference”category. Similar patterns are apparent for bothblack men and black women. For white women,the effects of “must-have” and “nice-to-have” areindistinguishable, although for most ideologicalcategories, each is distinguishable from the oddsratio among those stating no same-race prefer-ence.7

7The effect for those stating “must-have” may be partlydue to the mechanics of the site design, because for thosestating a must-have preference, the site automaticallydisplayed only same-race candidates, unless the user con-

Finally, Figure 5 provides no evidence of dif-ferences across ideological groups in the meaningof distinct statements of racial preference. If lib-erals were in fact more racially discriminatingfor a given level of stated racial preference, wewould expect the lines for different stated prefer-ences to slope downward. Liberals who declareda “nice-to-have” same-race preference, for exam-ple, would show stronger patterns of same-racebehaviors than conservatives who had expressedthat preference. In fact, across ideological groups,the lines are largely flat for all four race–gendergroups. Thus our findings indicate that liberalsand conservatives—unlike men and women—arenot using these terms in different ways, which in

ducted a custom search. However, the effect for nice-to-have preferences is not due to preferential ranking,because “nice-to-have” preferences have no effect on howcandidates are displayed to the user. We also assessedthe robustness of these results using a different samplingmethod that accounts for which profiles were shown inthe list presented to the users and found similar results(see the appendix).

sociological science | www.sociologicalscience.com 37 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

Male Female

● ● ● ●●

● ● ●

●● ●

●●

● ● ● ●●

● ● ●

●● ●

●●

● ● ● ●●

● ● ●

●● ●

●●

● ● ● ●●

● ● ●

●● ●

●●

1

2

4

8

16

32

64

128

256

1

2

4

8

16

32

64

128

256

White

Black

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Very

liber

al

Liber

al

Midd

le of

the

road

Conse

rvat

ive

Very

cons

erva

tive

Rev

eale

d sa

me−

race

pre

fere

nce

Must−have

Nice−to−have

No preference

Figure 5: Estimated revealed preferences for same-race partners by stated same-race preference. Barsindicate 95 percent confidence intervals.

turn suggests that there is little ideological effecton a willingness to express same-race preferencesrelative to acting on them.

DiscussionReturning to our initial motivation, our resultssuggest that individual preferences are an impor-tant explanation for the relative dearth of same-race relationships. In particular, we find not onlythat a large proportion of our population statesa same-race preference and acts on it but thateven individuals who state that they do not havea preference act as if they do. To the extent thatwe see a discrepancy between stated and revealedpreferences, this difference may be due to uncon-scious bias of which the respondent is unaware.Alternatively, it could be that these individualshave some acknowledged level of same-race pref-erence but believe it is weaker than “nice-to-have,”and so continue to state “no preference.”

We also find that although women are sub-stantially more likely than men to state a same-race preference, for any given stated preferencelevel, men display a stronger propensity to act ina same-race preferential manner and that thesecompeting effects largely cancel out; consequently,overall revealed same-race preferences for menand women are very similar. There also appearsto be little difference between blacks and whitesin these patterns. For political ideology, mean-while, we see a different pattern: conservativesare more likely than liberals to state a same-racepreference, but for any given level of stated pref-erences, both conservatives and liberals show asimilar propensity to act; thus both revealed andstated preferences for same-race partners increasewith political conservatism.

We close by noting that the patterns we haveobserved have implications for any policy inter-ventions designed to influence homogamy levels.In particular, our findings imply that merely alter-ing the structural environment—say, by creating

sociological science | www.sociologicalscience.com 38 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

more opportunities for individuals of differentraces to interact—would not necessarily amelio-rate persistent patterns of racial homogamy inromantic relationships. Especially among politi-cal conservatives, homogamy appears to derivein part from same-race preferences, where thesepreferences are shared broadly across racial andgender groups. Finally, a general behavioral pref-erence for racial homogamy is evident across allgroups, even among those that state a lack of suchpreference, implying that survey data underesti-mate the proportion of the population that willchoose same-race romantic partners. Althoughour study is silent on the malleability of thesepreferences, or how they might change with moreinterracial contact, our results nonetheless implythat opportunity alone will not eliminate the pre-ponderance of same-race romantic relationships.

ReferencesBaldassarri, Delia and Peter Bearman. 2007. “Dy-

namics of Political Polarization.” American So-ciological Review 72:784–811. http://dx.doi.org/10.1177/000312240707200507

Bernard, H. Russell, Peter Killworth, David Kro-nenfeld, and Lee Sailer. 1984. “The Prob-lem of Informant Accuracy: The Validity ofRetrospective Data.” Annual Review of An-thropology 13:495–517. http://dx.doi.org/10.1146/annurev.an.13.100184.002431

Blackwell, Debra L. and Daniel T. Lichter.2004. “Homogamy among Dating, Cohabiting,and Married Couples.” Sociological Quarterly45: 719–37. http://dx.doi.org/10.1111/j.1533-8525.2004.tb02311.x

Campbell, Karen E., Peter V. Marsden, andJeanne S. Hurlbert. 1986. “Social Re-sources and Socioeconomic Status.” SocialNetworks 8:97–117. http://dx.doi.org/10.1016/S0378-8733(86)80017-X

Coleman, James S. 1988. “Social Capital in theCreation of Human Capital.” American Jour-nal of Sociology 94:95–120. http://dx.doi.org/10.1086/228943

Crowne, Doulgas P. and David Marlowe. 1998.The Approval Motive: Studies in EvaluativeDependence. New York: John Wiley.

Federico, Christopher M. and Jim Sidanius. 2002.“Sophistication and the Antecedents of Whites’Racial Policy Attitudes: Racism, Ideology, andAffirmative Action in America.” Public OpinionQuarterly 66:145–76. http://dx.doi.org/10.1086/339848

Feld, Scott L. 1981. “The Focused Organiza-tion of Social Ties.” American Journal of So-ciology 86:1015–35. http://dx.doi.org/10.1086/227352

Feldman, Stanley and Leonie Huddy. 2005.“Racial Resentment and White Oppositionto Race-Conscious Programs: Principles orPrejudice?” American Journal of Politi-cal Science 49:168–83. http://dx.doi.org/10.1111/j.0092-5853.2005.00117.x

Fu, Xuanning and Tim B. Heaton. 2008. “Racialand Educational Homogamy: 1980 to 2000.”Sociological Perspectives 51:735–58. http://dx.doi.org/10.1525/sop.2008.51.4.735

Gilbert, Daniel. 2006. Stumbling on Happiness.New York: Vintage.

Grodsky, Eric and Devah Pager. 2001. “The Struc-ture of Disadvantage: Individual and Occupa-tional Determinants of the Black-White WageGap.” American Sociological Review 66:542–67.http://dx.doi.org/10.2307/3088922

Hitsch2010b Hitsch, Gunter J., Ali Hortacsu,and Dan Ariely. 2010. “What Makes YouClick? Mate Preferences and Matching Out-comes in Online Dating.” Quantitative Mar-keting and Economics 8:393–427. http://dx.doi.org/10.1007/s11129-010-9088-6

King, Gary and Langche Zeng. 2001. “Logistic Re-gression in Rare Events Data.” Political Anal-ysis 9:137–63. http://dx.doi.org/10.1093/oxfordjournals.pan.a004868

Kossinets, Gueorgi and Duncan J. Watts. 2009.“Origins of Homophily in an Evolving So-cial Network.” American Journal of So-ciology 115:405–50. http://dx.doi.org/10.1086/599247

sociological science | www.sociologicalscience.com 39 February 2014 | Volume 1

Anderson et al. Racial Preferences in Online Dating

Krosnick, Jon A. 1999. “Survey Research.”Annual Review of Psychology 50:537–67.http://dx.doi.org/10.1146/annurev.psych.50.1.537

Lazer, David, Alex (Sandy) Pentland, LadaAdamic, Sinan Aral, Albert Laszlo Barabasi,Devon Brewer, Nicholas Christakis, NoshirContractor, James Fowler, Myron Gut-mann, Tony Jebara, Gary King, MichaelMacy, Deb Roy, and Marshall Van Alstyne.2009. “Life in the Network: The Com-ing Age of Computational Social Science.”Science 323:721–23. http://dx.doi.org/10.1126/science.1167742

McPherson, Miller, Lynn Smith-Lovin, andJames M. Cook. 2001. “Birds of a Feather:Homophily in Social Networks.” Annual Re-view of Sociology 27:415–44. http://dx.doi.org/10.1146/annurev.soc.27.1.415

Nail, Paul R., Helen C. Harton, and Brian P.Decker. 2003. “Political Orientation and Mod-ern versus Aversive Racism: Tests of Do-vidio and Gaertner’s (1998) Integrated Model.”Journal of Personality and Social Psychol-ogy 84:754–70. http://dx.doi.org/10.1037/0022-3514.84.4.754

Nisbett, Richard E. and Timothy D. Wilson.1977. “Telling More Than We Can Know: Ver-bal Reports on Mental Processes.” Psycholog-ical Review 84:231–59. http://dx.doi.org/10.1037/0033-295X.84.3.231

Passel, Jeffrey S., Wendy Wang, and Paul Taylor.2010. “Marrying Out: One-in-Seven New USMarriages Is Interracial or Interethnic.” PewResearch Center. http://pewsocialtrends.org/assets/pdf/755-marryingout.pdf.

Portes, Alejandro. 1998. “Social Capital: ItsOrigins and Applications in Modern Sociology.”Annual Review of Sociology 24:1–24. http://dx.doi.org/10.1146/annurev.soc.24.1.1

Schoen, Robert and John Wooldredge. 1989.“Marriage Choices in North Carolina and Vir-ginia, 1969–71 and 1979–81.” Journal of Mar-riage and the Family 51:465–81. http://dx.doi.org/10.2307/352508

Sidanius, Jim, Felicia Pratto, and Lawrence Bobo.1996. “Racism, Conservatism, Affirmative Ac-tion, and Intellectual Sophistication: A Matterof Principled Conservatism or Group Domi-nance?” Journal of Personality and SocialPsychology 70:476–90. http://dx.doi.org/10.1037/0022-3514.70.3.476

Sniderman, Paul M. and Edward G. Carmines.1997. Reaching beyond Race. Cambridge, MA:Harvard University Press.

Whitley, Bernard E., Jr. 1999. “Right-Wing Au-thoritarianism, Social Dominance Orientation,and Prejudice.” Journal of Personality andSocial Psychology 77:126–34. http://dx.doi.org/10.1037/0022-3514.77.1.126

Wimmer, Andreas and Kevin Lewis. 2010. “Be-yond and Below Racial Homophily: ERGModels of a Friendship Network Documentedon Facebook.” American Journal of Soci-ology 116:583–642. http://dx.doi.org/10.1086/653658

Ashton Anderson: Department of ComputerScience, Stanford University. E-mail: [email protected].

Sharad Goel: Microsoft Research. E-mail: [email protected].

Gregory Huber: Department of Political Science,Yale University. E-mail: [email protected].

Neil Malhotra: Graduate School of Business, Stan-ford University. E-mail: [email protected].

Duncan J. Watts :Microsoft Research. E-mail: [email protected].

sociological science | www.sociologicalscience.com 40 February 2014 | Volume 1