Showing posts with label Nicolas Arning. Show all posts
Showing posts with label Nicolas Arning. Show all posts

Monday, 5 January 2026

New paper: Identifying direct risk factors for COVID-19 hospitalization in UK Biobank using Doublethink

This week sees publication of our new paper Identifying direct risk factors in UK Biobank via simultaneous Bayesian-frequentist model-averaged hypothesis testing using Doublethink in Proceedings of the National Academy of Sciences (PNAS). This work was joint with Nicolas Arning and Helen Fryer.

In this study we applied a novel approach called Doublehthink to implement an exposome-wide association study (ExWAS) of non-genetic risk factors that influence risk of COVID-19 hospitalization in UK Biobank (UKB). Inspired by Bayesian model averaging, our approach enhances power by testing both individual variables and arbitrary groups of variables.

We employed Doublethink to reveal exposome-wide significant signals across nine individual variables and seven groups of variables, notably factors like aging, dementia and prior infection overlooked by 85% of previous studies of the UK Biobank. We found significant direct effects among some commonly reported risk factors like age, sex, and obesity, but not others like cardiovascular disease. The effects of hypertension, depression, and diabetes appeared to be mediated via general comorbidity. 

Graph of variables with direct effects on COVID-19 hospitalization risk in UK Biobank, quantified by log sub 10 posterior odds and adjusted P-value.

Biobank-scale epidemiology has transformed the study of common diseases, particularly through the discovery of genetic risk factors using genome-wide association studies (GWAS) methods. ExWAS applies similiar logic to pursue agnostic, data-driven discovery of non-genetic risk factors.

Doublethink offers an ExWAS method that can test thousands of variables in hundreds of thousands of UK Biobank participants. It controls both Bayesian (false discovery rate; FDR) and classical (familywise error rate; FWER) measures of false positives. Its Bayesian model-averaging approach enables an agnostic approach to variable selection, but it also addresses drawbacks of Bayesian methods like computational burden and reliance on prior assumptions for its false positive control.

Through its capacity to interpret biobank-scale data, our new Doublethink-based ExWAS approach paves the way for future systematic analyses of risk factors for infectious and chronic diseases in UK Biobank and beyond.

Tuesday, 14 October 2025

New paper: Machine learning and statistical inference in microbial population genomics

We have published a new review article in Genome Biology contrasting machine learning and statistics in microbial genomics. This is joint work with Sam Sheppard, Nick Arning and David Eyre.

The availability of large genome datasets has changed the microbiology research landscape. Analyzing such data requires computationally demanding analyses, and new approaches have come from different data analysis philosophies. Machine learning and statistical inference have overlapping knowledge discovery aims and approaches.

In this review, we highlight how machine learning focuses on optimizing prediction, whereas statistical inference focuses on understanding the processes relating variables. We outline the different aims, assumptions, and resulting methodologies, with examples from microbial genomics. These approaches are essentially complementary, and we argue that exploiting both machine learning and statistics - selecting the right tool for the job - has the greatest potential for advancing pathogen research in the big data era.

Monday, 12 May 2025

Doublethink: simultaneous Bayesian-frequentist model-averaged hypothesis testing

Helen Fryer, Nick Arning and I have posted our new preprint to arxiv. This is the first version of the paper that we have submitted for peer review. Doublethink addresses some long-standing questions in assessing evidence for the purposes of hypothesis testing.

Hypothesis testing is central to scientific enquiry, but conclusions can be heavily influenced by model specification, particularly which variables are included. Bayesian model-averaged hypothesis testing offers a solution, but the sensitivity of posterior odds and Bayesian false discovery rate (FDR) guarantees to prior assumptions limit the appeal. In hypothesis testing, we lack unifying results – like Bernstein-von-Mises’ Theorem – that predict convergence of Bayesian and frequentist results, even in large samples.

Our paper introduces new theory and a practical method, Doublethink, motivated by these issues:
  • A key, and perhaps surprising, result is that Bayesian model-averaged hypothesis testing natively controls not only the Bayesian FDR, but also the frequentist strong-sense familywise error rate (FWER). This duality – which is general – seems to be unknown, or forgotten.
  • For practical application, we derive large-sample asymptotic theory to quantify the rate at which the FWER is controlled. Specifically, we use a BIC-like model to characterize the tail probability of the model-averaged posterior odds via a chi-squared distribution.
  • This result enables simultaneous control of Bayesian FDR and frequentist FWER at quantifiable levels and – equivalently – simultaneous reporting of posterior odds and asymptotic p-values.
  • We explore the method’s benefits – like post-hoc variable selection – and limitations – like inflation – through a Mendelian Randomization study and detailed simulations, comparing Doublethink to Lasso, stepwise regression, the Benjamini-Hochberg procedure and e-values.
Besides the practical benefits of model-averaged hypothesis testing with frequentist guarantees, and the implications that entails for objective Bayesian hypothesis testing, these results offer fundamental insights likely to trigger renewed discussion of FDR, FWER and the reconcilability of p-values with evidence.

Doublethink is a novel addition to the emerging class of heavy-tailed combination tests. Since 2019, methods like the Cauchy combination test and harmonic mean p-value have surfaced as powerful tools for combining hypothesis tests despite inter-test dependence. Doublethink improves on these methods by allowing model uncertainty in the null hypothesis and by improving power.

We believe this paper will be of broad interest, addressing questions of importance to statistical methodology, big data analysis and scientific enquiry more generally.

Explanation of variables above

  • The model-averaged p-value, adjusted for multiple testing, is p*.
  • The model-averaged posterior odds, calculated from a Bayesian analysis, is PO.
  • The number of variables in the analysis is ν.
  • The prior odds of including each variable are μ.
  • The sample size n is represented by ξn, which decreases as √n increases.

Wednesday, 2 April 2025

Machine Learning versus Statistical Inference in Microbial Genomics

My talk given today at the 2025 Microbiology Society Conference in Liverpool:

Abstract

The advent of vast genomic datasets has transformed microbiology, presenting opportunities and challenges for data analysis. The distinct philosophies of machine learning (ML) and statistical inference gives them complementarity strengths and weaknesses in tackling big data problems in pathogen research. While statistical inference prioritizes understanding underlying relationships, ML focuses on optimizing predictive performance. In this talk I will contrast the approaches and offer a view on their relative utility for three problems: source attribution, bacterial genome-wide association studies, and predicting antimicrobial resistance phenotypes from whole genome sequences.

Monday, 29 July 2024

Doublethink methods paper

Today we release the first full draft of the Doublethink methods paper. This is an evolution of what was originally conceived as the supplement to the Doublethink COVID-19 paper. The wider significance of the results persuaded us to separate the two, which now focus on:

  • Doublethink methods paper: Broad connections between Bayesian and classical hypothesis testing that we hope bring the best of both world by enabling scientists to simultaneously control the Bayesian false discovery rate and the classical familywise error rate, in big data settings.
  • Doublethink COVID-19 paper: Identifying direct risk factors for COVID-19 hospitalization among 2000 candidate variables in 200,000 UK Biobank participants. Compares results to the literature and considers the limitations imposed by mediation and complex 'exposome-wide' association studies.
After soliciting colleagues for comments and another round of editing, we will move toward submission in later this year.

Wednesday, 3 January 2024

Introducing Doublethink: joint Bayesian-frequentist model-averaged hypothesis testing

This week Nick Arning, Helen Fryer and I released two related preprints describing a new method called Doublethink, and its application to identifying risk factors for COVID-19 hospitalization in UK Biobank:

Doublethink: Bayesian-frequentist model-averaged hypothesis testing

Doublethink enables joint Bayesian and frequentist hypothesis testing when there is model uncertainty by interconverting Bayesian posterior odds and classical (frequentist) p-values. It has broad implications because (i) it reveals connections between the Bayesian approach to model averaging and the classical approach to multiple testing, and (ii) it brings the benefits of Bayesian model averaging to classical statistics.
Doublethink addresses two fundamental problems in hypothesis testing:
  1. In classical tests, the statistical evidence that one variable directly affects an outcome generally depends on which other variables are assumed to directly affect it.
  2. In Bayesian tests, the statistical evidence that one variable directly affects an outcome depends on the prior assumptions.
These issues are addressed by computing p-values from Bayesian model-averaged posterior odds, which (1) account for model uncertainty and (2) are theoretically invariant to prior assumptions, assuming large sample sizes.
Doublethink simultaneously controls the frequentist family-wise error rate (FWER) and the Bayesian false discovery rate (FDR). It builds on Johnson's Bayesian tests based on likelihood ratio statistics, and Karamata's theory of regular variation.

Identifying direct risk factors in UK Biobank with Doublethink

We applied Doublethink to identify direct risk factors for COVID-19 hospitalization in UK Biobank. This is a well-studied problem but we took an 'exposome-wide' approach in which we evaluated whether 1,900 variables measured in the UK Biobank each affected the outcome. This is still an under-utilized approach in epidemiology, which usually focuses on candidate risk factors.
Exposome-wide approaches have potential benefits over candidate risk factor approaches, including:
  • The ability to discover unexpected results.
  • Stringent control for multiple testing.
  • Avoidance of bias in choosing candidate risk factors or deciding to publish.
However, we only studied the direct effects of variables on the outcome. This means we cannot make statements about the total (direct and indirect) effects of a variable, e.g. smoking, on the outcome, which are needed in applications like assessing potential interventions.
We identified individual variables and groups of variables that were 'exposome-wide significant' at 9% FDR and 0.05% FWER, after accounting for the direct effects of all other variables.

Comparing our results to over 100 published studies of COVID-19 in UK Biobank, we
  • Recapitulated several commonly reported direct risk factors, e.g. age, sex, and obesity.
  • Excluded others, e.g. diabetes, cardiovascular disease, and hypertension, which might be mediated through other variables that measure general comorbidity.
  • Identified some infrequently reported direct risk factors, both individually, e.g. lung infection, and as groups, e.g. constipation/urinary tract infection, which might reflect underlying kidney disease.
The ability to test groups of variables, which increases sensitivity, was one of the benefits of Doublethink's model-averaging approach. It is particularly helpful in large biobanks that measure thousands of variables, because correlation between variables is pervasive, and can dilute the significance of individual variables that measure similar phenomena, like the numerous types of deprivation index. It serves as a flexible alternative to pre-analysis variable filtering algorithms, while controlling the risk of false positives by pre-defining significance thresholds for all possible tests.
To read more, please check out the preprints here and here.

Tuesday, 7 December 2021

New paper: Machine learning to predict the source of campylobacteriosis using whole genome data

This study, published in October in PLOS Genetics, brings together machine learning, large bacterial isolate collections and whole genome sequencing to address the general problem of how to trace the source of human infections.

Specifically, we investigated campylobacteriosis, a common infection of animal origin causing ~1.5 million cases of gastroenteritis and 10,000 hospitalizations every year in the United States alone. We show that our combined machine learning/genomics analyses:

  • Improve the accuracy with which infections can be traced back to farm reservoirs.
  • Identify evolutionary shifts in bacterial affinity for livestock host species.
  • Detect changes in human infection capability within related strains.

These results will improve understanding not only of Campylobacter, but more generally as these technologies can readily be applied to other important bacterial pathogen species.

This paper builds on previous work published by the group, including our well cited Tracing the source of campylobacteriosis (Wilson et al 2008, PLOS Genetics 4:e1000203). The use of these methods for tracing infection has influenced public health policy and contributed to reducing disease burden.

This work demonstrates the potential for modern genomics and artificial intelligence approaches to address common and serious problems that affect our everyday lives. The awareness of the importance of infection to society has rarely been higher than in 2021, and while the current pandemic imposes an acute global problem, other infections continue to present long-term threats to health and productivity.

This work was led by Nicolas Arning, in collaboration with David Clifton and Sam Sheppard.