Welcome to Shirley's website

Shirley Xiaoqi Liu

Research

Empirical advances in machine learning have outpaced our theoretical understanding, leaving open central questions about when, why, and how modern machine learning systems work. My research aims to develop statistical and information-theoretic foundations for machine learning that help address these questions. A recurring theme of my research is learning under data heterogeneity since real-world data varies across time, latent groups, and problem instances, violating the stationarity and homogeneity assumptions on which much of classical theory rests.

Drawing on statistical learning theory, high-dimensional statistics, information theory, probability theory, and optimisation, my recent work is organised around three inter-connected threads:

Training dynamics of canonical statistical & machine learning models

A key question in learning theory is how the model class and training procedure jointly determine the learned solution and its generalisation properties. Current theory, both asymptotic and finite-sample, extends primarily to canonical models such as Gaussian mixtures and multi-index models (i.e., two-layer neural networks), since precise analyses are demanding even in these cases. Nevertheless, such analyses have repeatedly sharpened our understanding of phenomena including early stopping, benign overfitting, and feature learning, and informed practical choices of model and training procedure. Two of my recent and ongoing projects contribute to this area.

Sequential decision-making under data heterogeneity

Sequential decision-making arises wherever actions must be taken one at a time, under partial feedback, with each action shaping what is observed next: clinical trials allocating patients as evidence accrues, recommender systems learning preferences from the choices they elicit, or inference pipelines escalating a query through a cascade of increasingly costly models, deciding at each stage whether to exit or defer. What makes such problems hard is that the information needed to act optimally is unavailable at the time of action, so practical systems rely heavily on heuristics. My work develops procedures with provable optimality guarantees that can inform or supersede heuristics.

Uncertainty quantification under weak distributional assumptions

Quantifying uncertainty is relatively straightforward when the model is well-specified and much harder otherwise. My ongoing research asks what can still be certified when the usual distributional assumptions are withdrawn. The main tools I use are e-values and e-processes, conformal prediction, and prediction-powered inference, which together deliver inference that remains valid at arbitrary, data-dependent sample sizes, under minimal data distributional assumptions.


During my PhD, I focused on message passing algorithms for high-dimensional statistical estimation and inference: a family of computationally efficient, iterative algorithms that provably achieve statistical optimality across a range of problems. I applied these to changepoint localisation, sketching of sparse low-rank matrices, and reliable communication for large user networks.

Interests:

  • Training dynamics: early stopping, benign overfitting, feature learning.
  • First-order optimisation methods: gradient descent, approximate message passing
  • Sequential decision-making: bandits, online learning, adaptive inference
  • Uncertainty quantification: e-values, conformal prediction, prediction-powered inference
  • Information theory and communication systems

Selected publications:

  1. X. Liu, D. Baudry, J. Zimmert, P. Rebeschini, A. Akhavan, “Non-stationary Bandit Convex Optimization: A Comprehensive Study”, Proceedings of the 39th Annual Conference on Neural Information Processing Systems, 2025. (talk, poster)
  2. In alphabetical order: J. Allison, P. Anderson, E. Aranas, Y. Assaf, M. Caballero, J. Chattaway, A Chatzieleftheriou, J. Clegg, B. Cooper, T. Deegan, A. Donnelly, R. Drevinskas, C. Gkantsidis, A. G. Diaz, I. Haller, F. Hong, T. Ilieva, R. Joyce, V. Kapitany, W. Kunkel, D. Lara, T. Lawson, S. Legtchenko, F. Liu, X. Liu, B. Magalhaes, S. Nowozin, H. Overweg, A. Rowstron, M. Sakakura, N. Schreiner, A. Smith, O. Snowdon, I. Stefanovici, D. Sweeney, G. Verkes, P. Wainman, C. Whittaker, P. W. Berenguer, H. Williams, T. Winkler, S. Winzeck, R. Black, B. Canakci, D. Cletheroe, Z. Feng, “Laser writing in glass for dense, fast and efficient archival data storage”, Nature 650, 606–612, 2026.
  3. X. Liu, P. Pascual Cobo and R. Venkataramanan, “Many-user multiple access with random user activity: achievability bounds and efficient schemes”, IEEE Transactions on Information Theory, vol. 72, no. 1, pp. 383-414, Jan. 2026, doi: 10.1109/TIT.2025.3622969. (conference version, talk, poster)
  4. X. Liu, K. Hsieh and R. Venkataramanan, “Coded many-user multiple access via Approximate Messsage Passing”, to appear in Information Theory, Probability and Statistical Learning: A Festschrift in Honor of Andrew Barron, 2025. (conference version, talk)
  5. X. Liu, “Message Passing Algorithms for Statistical Estimation and Communication”, PhD thesis, Apollo-University of Cambridge Repository, 2024.
  6. G. Arpino, X. Liu, J. Gontarek and R. Venkataramanan, “Inferring Change Points in High-Dimensional Regression via Approximate Message Passing”, Journal of Machine Learning Research, 26(225), pp.1-49, 2025.
  7. G. Arpino, X. Liu and R. Venkataramanan, “Inferring Change Points in High-Dimensional Linear Regression via Approximate Message Passing”, Proceedings of the 41st International Conference on Machine Learning, PMLR 235:1841-1864, 2024. (talk, poster, code)
  8. X. Liu and R. Venkataramanan, “Sketching Sparse Low-Rank Matrices With Near-Optimal Sample- and Time-Complexity Using Message Passing”, IEEE Transactions on Information Theory, vol. 69, no. 9, pp. 6071-6097, Sept. 2023, doi: 10.1109/TIT.2023.3273181. (conference version)

Selected talks:

  1. “Bayes-Assisted Confidence Sequences for CDFs”, Safe Anytime-Valid Inference (SAVI) Conference, University of Twente, July 2026.
  2. “Mind the U Curve: Why We Stop Gradient Descent Early”, Meet Our Group Seminar Series, Department of Statistics, University of Oxford, May 2026.
  3. “Learning Under Non-Stationarity: Statistical Inference & Bandit Convex Optimization”, Warwick Algorithms and Computationally Intensive Inference Seminar, Department of Statistics, University of Warwick, March 2026.
  4. “A Friendly Talk on Bandit Convex Optimization in Changing Environments”, Oxford Young Statisticians Seminar, Department of Statistics, University of Oxford, March 2026.
  5. “Tight Confidence Sequences From Universal Portfolio”, Hypothesis Testing with E-values Reading Group, Department of Statistics, University of Oxford, March 2026.
  6. “Tutorial on Sequential Anytime-Valid Inference Using E-processes”, Statistics in AI CDT Module, Department of Statistics, University of Oxford, November 2025.
  7. “Sequential Anytime-Valid Inference Using E-processes”, Hypothesis Testing with E-values Reading Group, Department of Statistics, University of Oxford, October 2025.)
  8. “Loss landscapes and optimizaton in over-parameterized non-linear systems and neural networks”, Learning Theory and Statistical Optimization Reading Group, Department of Statistics, University of Oxford, November 2024.
  9. “Communication over many-user channels via Approximate Message Passing”, Information Theory Seminar, Department of Mathematics, University of Cambridge, May 2024.

Service:

I review papers for several conferences and journals including the Conference on Neural Information Processing Systems (NeurIPS) (top reviewer 2024), International Conference on Machine Learning (ICML), Transactions on Machine Learning Research (TMLR), and International Symposium on Information Theory (ISIT).