Group comparisons
Section 4.2 of Mize and Han (forthcoming): separate models for each group, and which observations to use
Group comparisons span interaction effects and cross-model comparisons. A single interaction term lets one variable’s effect differ across groups; at the other end of the spectrum, separate models for each group let every effect differ. Mize and Han (forthcoming) discuss the trade-off and show both. This page reproduces Table 6 of the chapter (separate models for men and women) and its Section 4.2.c example of the choice of observations when a single model is used.
Separate models for each group (Table 6)
The same happiness model as the chapter’s Section 4.1 (the nested models example) is fit separately for men and for women, on the 2000–2021 GSS respondents who are employed. The groups option tells mecompare that the two models were fit to non-overlapping samples, so each model’s marginal effects are averaged over its own group; groupnames() labels the rows. With no varlist, every predictor is reported; amount(sd) sets a standard-deviation change for age, and pwcompare gives all of the race-ethnicity contrasts.
use "https://tdmize.github.io/data/data/gss_cme", clear(gss_cme.dta | GSS 1972 - 2016 Weighted | 2018-07-10)
drop if year < 2000(38,116 observations deleted)
drop if employed != 1(9,556 observations deleted)
drop if missing(vhappy, college, wages, occprest, age, married, parent, woman, conserv, race)(5,043 observations deleted)
.
quietly logit vhappy i.college i.married i.parent i.conserv i.race c.age##c.age ///
if woman == 0, vce(robust)
estimates store menmod.
quietly logit vhappy i.college i.married i.parent i.conserv i.race c.age##c.age ///
if woman == 1, vce(robust)
estimates store wommodmecompare, models(menmod wommod) groups groupnames(Men Women) amount(sd) pwcomparePredicting: Pr(vhappy)
Marginal effects and cross-model differences (N_Men=4932) (N_Women=4819)
| ME # Estimate Robust SE P>|z|
---------------------------------+---------------------------------------
college |
College Deg - No Col Deg |
Men | 1 0.041 0.014 0.004
Women | 2 0.064 0.014 0.000
Difference | 3 -0.023 0.020 0.245
---------------------------------+---------------------------------------
married |
Married - Not Married |
Men | 4 0.197 0.014 0.000
Women | 5 0.204 0.014 0.000
Difference | 6 -0.008 0.020 0.703
---------------------------------+---------------------------------------
parent |
Parent - No Children |
Men | 7 0.023 0.016 0.153
Women | 8 -0.045 0.017 0.006
Difference | 9 0.068 0.023 0.003
---------------------------------+---------------------------------------
conserv |
Conservativ - Not Conser |
Men | 10 0.046 0.014 0.001
Women | 11 0.046 0.014 0.001
Difference | 12 -0.000 0.020 0.988
---------------------------------+---------------------------------------
race |
black - white |
Men | 13 -0.002 0.021 0.911
Women | 14 -0.020 0.019 0.291
Difference | 15 0.017 0.028 0.541
other - white |
Men | 16 -0.035 0.021 0.096
Women | 17 -0.015 0.023 0.529
Difference | 18 -0.020 0.031 0.523
other - black |
Men | 19 -0.032 0.028 0.244
Women | 20 0.005 0.028 0.858
Difference | 21 -0.037 0.039 0.341
---------------------------------+---------------------------------------
age + SD (centered) |
Men | 22 -0.010 0.008 0.176
Women | 23 -0.005 0.008 0.465
Difference | 24 -0.005 0.011 0.654
NOTE: SD's are based on all 9751 observations pooled across both groups.
Each block is one marginal effect: the effect for men, the effect for women, and the Difference (men minus women, the order of models()). Both married men and married women report being happier than their unmarried counterparts, and college degrees and conservatism have similar positive effects for both genders. The one significant gender difference is the effect of being a parent: it does not affect the probability of being very happy for men, but women who are parents are less likely to be very happy than women who are not.
Which observations: a single model with by(), over(), or separate models
When the groups are in one model, there is a choice of which observations to use when calculating each group’s marginal effect. The chapter’s example compares Germany and Nigeria in the World Values Survey, two countries with very different age distributions, in a linear regression of acceptance of same-sex relationships on age (with a squared term) and country.
use "https://tdmize.github.io/data/data/lvm_wvs6", clear(World Values Survey 6 2010 - 2014 | LVM | 2019-01-28)
drop if age < 18(93 observations deleted)
keep if country == 276 | country == 566(85,669 observations deleted)
.
quietly regress Jgaylesbian c.age##c.age i.country, vce(robust)
estimates store glmodThe default – and what by() does – is to use all observations for every group’s marginal effect: everyone is set to Germany, then everyone to Nigeria, and the effect of a ten-year increase in age is averaged over the whole sample each time. This controls for the two countries’ different age distributions. Because this model has no age-by-country interaction, the counterfactual effects are identical:
mecompare age, models(glmod) amount(10) by(country)Predicting: Linear prediction
Marginal effects (N_glmod=3735)
| ME # Estimate Robust SE P>|z|
---------------------------------+---------------------------------------
age + 10 (centered) |
Germany | 1 -0.174 0.033 0.000
Nigeria | 2 -0.174 0.033 0.000
over() instead uses only each country’s own respondents. The effect of age is nonlinear, so where a country’s respondents sit on the age curve changes the average effect, and the two countries now differ even though they share every coefficient:
mecompare age, models(glmod) amount(10) over(country)Predicting: Linear prediction
Marginal effects (N_glmod=3735)
| ME # Estimate Robust SE P>|z|
---------------------------------+---------------------------------------
age + 10 (centered) |
Germany | 1 -0.311 0.032 0.000
Nigeria | 2 -0.020 0.051 0.701
Separate models with groups, as in Table 6 above, go one step further: each country has its own coefficients and its own observations, so the cross-model difference reflects both.
quietly regress Jgaylesbian c.age##c.age if country == 276, vce(robust)
estimates store germany
quietly regress Jgaylesbian c.age##c.age if country == 566, vce(robust)
estimates store nigeria.
mecompare age, models(germany nigeria) groups groupnames(Germany Nigeria) amount(10)Predicting: Linear prediction
Marginal effects and cross-model differences (N_Germany=1976) (N_Nigeria=1759)
| ME # Estimate Robust SE P>|z|
---------------------------------+---------------------------------------
age + 10 (centered) |
Germany | 1 -0.338 0.040 0.000
Nigeria | 2 -0.104 0.049 0.035
Difference | 3 -0.234 0.064 0.000
Which to report depends on the question. If the goal is to control for all differences across the groups, including the distributions of the focal and control variables, use a single model and all observations (by()). If the groups’ different compositions are part of the question – for example, which group an intervention would matter more for – use group-specific observations (over()) or separate models (groups). The chapter’s Section 4.2.b–c discusses the two causal questions behind this choice; the Options page has more on by() versus over().
Reference
Mize, Trenton D. and Bing Han. Forthcoming. “Marginal effects: flexible methods for interpretation across linear and nonlinear models.” In Understanding Data Modeling and Data Analysis, edited by David Weakliem. Edward Elgar Publishing.