Our group’s research programs aim to explore and expand upon how advances in causal inference, causal machine learning and semi-parametric estimation, non-parametric regression and statistical machine learning, and computational statistics catalyze discovery in the health and medical sciences. The methodological research program emphasizes assumption-lean frameworks for inference and applies a translational perspective to formulate causal-analytic, statistical methods tailored to help answer substantive questions that arise in our collaborative applied science research. Broadly, our approach draws upon principles from causal inference to translate scientific questions into precise, interpretable statistical estimands, which one can then learn from data generated by observational studies or randomized controlled trials via analytic methods that
- avoid imposing restrictions not justified by available domain knowledge;
- incorporate flexible, adaptive modeling strategies (e.g., machine learning); and
- apply semi-parametric efficiency theory for best-in-class uncertainty quantification.
Thus, thematically, our research program integrates core aspects of causal inference, to define and target interpretable estimands, with tools from non-parametric estimation and statistical machine learning, to avoid unnecessary and restrictive modeling assumptions, and applies semi-parametric theory for asymptotically efficient estimation. This line of work has yielded novel insights applicable for causal (i.e., de-biased, targeted) machine learning (e.g., targeted minimum loss estimation, sieve estimation); non-parametric causal mediation analysis to study questions of mechanism using novel direct and indirect effect estimands; treatment effect heterogeneity and effect modification analyses to inform stratified medicine; variance-moderated semi-parametric estimation for flexible and stable biomarker discovery; corrections necessary to reliably draw accurate inferences from data collected using two-phase sampling designs; and novel causal effect estimands for continuous exposures.
Our collaborative applied science research program concentrates primarily on problems in chronic and infectious diseases, though we have also contributed to work in environmental health, cancer, substance use epidemiology, and nutrition, among others. We are often interested in and open to working in new substantive domains—wherever rigorous statistical thinking is welcome.
Here are a few highlights from research projects completed over the last few years:
A secondary theme of our research centers on the role of high-performance numerical computing and the development of open-source software tools for statistical science. While distinct, these areas are unified by the overarching aims of pushing the boundaries of statistical methods development and promoting reproducibility and transparency in the practice of applied statistics. Consistent with our commitment to open science, new methods developed by members of the lab are accompanied by open-source software implementations both to ensure replicability of the reported work and to facilitate widespread use of the newly developed techniques.
Browse more about our work on statistical methods innovations, their use in applied sciences, and in developing free and open-source software for statistical science.




