James Ming Liang Ang

Research

How learning occurs in neural networks, and how understanding that process can help us design better learning algorithms.

How does learning occur in neural networks?

Neural networks learn through sequences of local parameter updates, yet those updates can produce representations and behaviours that generalize far beyond the training data. I study the dynamics of this process: how optimization algorithms, training data geometry, and model architecture interact during training; which properties emerge at different stages; and which mechanisms persist across models and scales. The aim is to move beyond describing a trained model towards explaining the process that produced it.

How can we design better learning algorithms?

Understanding learning should help us design algorithms, not merely diagnose them after the fact. I am interested in deriving training rules from principles—particularly Bayesian inference, optimization, and information geometry—so that we can explain why an update works and extend it systematically.

Case study: variance reduction as posterior correction. Stochastic variance-reduced gradient methods such as SVRG speed up training by correcting each mini-batch gradient with a full-batch gradient computed at an older iterate. With Nico Daheim, Thomas Möllenhoff, and Emtiyaz Khan, I showed that SVRG can be derived as a special case of Bayesian posterior correction [1]. In variational form, the update reads

qin ← qin1−η q̂outη exp( −ηN [ ℓ̂i|in − ℓ̂i|out ] ), (1)

where the site functions ℓ̂ are linear surrogates of each example’s loss at the current and stored posteriors. Choosing an isotropic Gaussian for q and removing its sampling noise recovers SVRG exactly. Richer exponential families then yield new algorithms: a Newton-like variant that corrects Hessians as well as gradients, and an Adam-like variant built on the IVON optimizer, which we test on GPT-2 pretraining and fine-tuning.

Selected Publications

  1. SVRG and Beyond via Posterior Correction N. Daheim, T. Möllenhoff, M. L. Ang, M. E. Khan International Conference on Machine Learning (ICML), 2026. Oral. arXiv:2512.01930
  2. Explanation in an Emerging Science of Large Language Models J. M. L. Ang ICML 2026 Workshop, Philosophy Meets Machine Learning: What Counts as Trustworthy?