Standardization and decomposition: A conceptual introduction

What are Standardization and Decomposition

Standardization

  • Originally used to adjust mortality comparisons between populations with different age structures (Keiding and Clayton 2014)
  • Lets you “compar[e] rates between populations with different age structures by applying age-specific rates to a single ‘target’ age structure and, thereafter, comparing predicted marginal1 summaries in this target population.” (Keiding and Clayton 2014, p529)
  • More generally - lets you compare rates of any outcome between populations with different compositions of any confounder

Standardization

  • Conceptually the focus is on aggregate/marginal rates, not individual effects
  • Individual effects can aggregate in non-intuitive ways Simpson (1951)
  • Aggregate rates are really common as performance indicators: the reconviction rate, child poverty rate, access to greenspace, proportion of people in unmanageable debt, victimization rate… were all part of Scotland’s National Performance Framework

An example

  • “For example, before comparing the death rates for the residents of two areas, demographers frequently control the factors of differences between the areas in age, sex and race [sic] composition” Kitagawa (1955)
  • Take the observed age/sex/race death rates for your two areas
  • Multiply these by some ‘standard’ age, sex and race [sic] profile
  • You have ‘standardized’ rates!

Picking a standard population

  • Say you wanted to compare mortality rates in Glasgow and Shetland, controlling for differences in age sex and ethnicity
  • Which population would you use as a standard one?
    • The Glasgow population?
    • The Shetland population?
    • The average of the two populations?
    • Something else?

Student discussion

  • Have a think about it, talk it over if you like (in person or the Teams chat)
  • Vote at this menti link

Decomposition

  • So you’ve got some standardized rates - now what?
  • Decomposition asks a related question to standardization: How much of the observed difference in rates is due to differences in the factors you’ve standardized by?
  • From the previous example: how much of the difference in death rates for the residents of two areas is due to each of differences between the areas in age, sex and race [sic] composition?

Why decomposition?

  • Standardized rates are ‘artificial’, summaries of a world we do not live in
  • Why should we care what the mortality rate in Scotland would be if it had, say, the Estonian age distribution?
  • Decomposition makes comparisons of standardized rates more analytically useful by showing the contributions of the factors you standardize by:
  • “A systematic statement of relationships between crude and standardized rates for two or more groups may help to bridge the gap between the observed’ crude (or total) rates and the ‘artificial’ standardized rates” (Kitagawa 1955, pp1169–1170)

Decomposition

  • Decomposition lets you say how much of the difference between crude rates of your outcome between your comparison cases are due to the various factors you choose to control for
  • In Kitagawa’s (1955) example, how much of the difference in mortality rates between areas is due to differences in their age, sex and race [sic] composition, and how much is ‘left over’

A conjecture

  • A lot of social science research either aims to identify the ‘effect’ of some independent variable on an outcome
  • In either an associational or causal way (Holland 1986)
  • But this is only half the picture

A conjecture

  • The What is Your Estimand? approach to social science splits a target ‘estimand’ into “a unit-specific quantity … aggregated over a target population of units” (Lundberg, Johnson, and Stewart 2021, p535)
  • It’s useful to be able to answer questions about macro-level phenomena, and understand the processes of aggregation which connect individual or sub-group effects to aggregate outcomes
  • Standardization and decomposition are one way to do this (there are others!)

A conjecture

  • You can be interested in (average?) individual effects, or the aggregation of individual effects in different populations
  • We can think of these as ‘direct’ and ‘compositional’ effects respectively (Vaupel and Canudas-Romo 2002)
  • But you need to be clear what you are interested in

My motivating example

  • I (Ben) wrote my PhD thesis on the ‘crime drop’ in Scotland and how the falls in convictions overall were reflected in changing demographics of crime and patterns of crime at the individual level (Matthews 2017)
  • There is a tension at the heart of this analysis I struggled to get to grips with - my focus was both at the aggregate level (the crime drop literature is about aggregate crime rates) but also at the individual level (conviction rates for people of different ages)
  • Standardization and decomposition appealed to me because I was trying to separate out how much of the change in the aggregate conviction rate was due to underlying changes in demographics, convictions prevalence and the number of convictions per person convicted

Any questions?

An exercise

Thinking about direct and compositional effects

  • Next we’re going to look at a worked example about age and reconviction
  • There is a direct effect of age on reconviction: younger people typically are reconvicted at higher rates than older people
  • There is also a compositional effect: typically younger people make up more of ‘reconviction cohorts’ because younger people are more likely to be convicted in the first place than older people

Thinking about direct and compositional effects

  • Imagine you are researching whether breastfeeding practices can explain differences in children’s height at age two between Scotland and Argentina
  • What would be a direct effect? And what would be a compositional effect?
  • Take a minute then vote on the options at the menti

Thinking about direct and compositional effects

  • Take two minutes to think about your own research topic
  • What is your key ‘independent’ variable?
  • What might that look like as a direct effect? And what might it look like as a compositional effect?

A caveat

  • Lundberg, Johnson, and Stewart (2021) are actually quite critical of standardization and decomposition
  • Whilst we have a counterfactual it is only a descriptive counterfactual, not a causal one
  • If used purely descriptively “the link to theory is weak: why exactly do we care about that reweighting of the Mexican population, given that Mexico does not have the age distribution of the U.S.?”
  • If used causally - “the causal difference between U.S. mortality and the counterfactual mortality that U.S. individuals would experience under an intervention to move them to Mexico” - the “link to evidence is weak”

So what’s it good for?

  • Lundberg, Johnson, and Stewart (2021) are right about causality - you don’t get causal effects out of standardization and decomposition (indeed Das Gupta (1993) says as much)
  • But I think the descriptive use is still important (we don’t all have to do causal analysis all the time)
  • And I think sensitising researchers to possible role of compositional effects is very valuable too

References

Das Gupta, Prithwis. 1993. Standardization and Decomposition of Rates: A User’s Manual.
Gelman, Andrew. 2006. “Marginal and Marginal.” Statistical Modeling, Causal Inference, and Social Science.
Good, I. J., and Y. Mittal. 1987. “The Amalgamation and Geometry of Two-by-Two Contingency Tables.” The Annals of Statistics 15 (2): 694–711. https://www.jstor.org/stable/2241334.
Holland, Paul W. 1986. “Statistics and Causal Inference.” Journal of the American Statistical Association 81 (396): 945–60. https://doi.org/10.2307/2289064.
Keiding, Niels, and David Clayton. 2014. “Standardization and Control for Confounding in Observational Studies: A Historical Perspective.” Statistical Science 29 (4): 529–58. https://doi.org/10.1214/13-STS453.
Kitagawa, Evelyn M. 1955. “Components of a Difference Between Two Rates.” Journal of the American Statistical Association 50 (272): 1168–94. https://doi.org/10.1080/01621459.1955.10501299.
Lundberg, Ian, Rebecca Johnson, and Brandon M. Stewart. 2021. “What Is Your Estimand? Defining the Target Quantity Connects Statistical Evidence to Theory.” American Sociological Review 86 (3): 532–65. https://doi.org/10.1177/00031224211004187.
Matthews, Ben. 2017. “Criminal Careers and the Crime Drop in Scotland, 1989-2011: An Exploration of Conviction Trends Across Age and Sex.” PhD thesis, University of Edinburgh.
Rose, Geoffrey. 1985. “Sick Individuals and Sick Populations.” International Journal of Epidemiology 14 (1): 32–38. https://doi.org/10.1093/ije/14.1.32.
Simpson, E. H. 1951. “The Interpretation of Interaction in Contingency Tables.” Journal of the Royal Statistical Society: Series B (Methodological) 13 (2): 238–41. https://doi.org/10.1111/j.2517-6161.1951.tb00088.x.
Vaupel, James W., and Vladimir Canudas-Romo. 2002. “Decomposing Demographic Change into Direct Vs. Compositional Components.” Demographic Research 7 (July): 1–14. https://doi.org/10.4054/DemRes.2002.7.1.