Research directions
How can learning systems scale reliably?
My work treats scaling as a mathematical question about representation, geometry, simulation,
and resource allocation rather than an empirical slogan.
01
Inference Time Scaling
Monte Carlo, control, and sequential methods that convert additional inference computation into reliable accuracy gains.
Read more →
Selected works (2)
Generative AI · 2026
URGE
Unbiased derivative free inference time scaling for diffusion models through sequential Monte Carlo on path measures.
Paper ↗
LLM Reasoning Theory · 2026
On the Power of Approximate Reward Models for Inference Time Scaling
A theory of when approximate reward models reduce the complexity of long-horizon LLM reasoning from exponential to polynomial through SMC inference-time scaling.
Paper ↗
02
Scientific Machine Learning
Structure preserving learning for PDEs, operator learning, uncertainty quantification, and simulation calibrated correction.
Read more →
Selected works (5)
AI for Science · ICLR 2026
Simulation Calibrated Scientific ML
Inference time defect correction improves high dimensional PDE solvers without retraining the learned model.
Paper ↗
Equation Discovery · ICML 2018
PDE Net: Learning PDEs from Data
Learns differential operators and nonlinear dynamics through constrained convolution filters to identify PDEs from observed data.
Paper ↗
Equation Discovery · JCP 2019
PDE Net 2.0
Combines learnable differential operators with symbolic neural networks to recover explicit PDE models and predict their dynamics.
Paper ↗
PDE Learning Theory · ICLR 2022
Machine Learning for Elliptic PDEs
Establishes sharp generalization bounds and minimax rates for PINNs and a modified Deep Ritz method in a prototype elliptic PDE setting.
Paper ↗
Operator Learning · ICLR 2023 Spotlight
Minimax Optimal Kernel Operator Learning via Multilevel Training
Develops minimax optimal rates and a multilevel algorithm for learning linear operators between infinite dimensional function spaces.
Paper ↗
03
Optimization and Reliability
Width and depth stable optimization geometry, predictable hyperparameter transfer, and robust learning algorithms.
Read more →
Selected works (3)
Optimization · 2026
Scaling Neural Optimizers
Matrix operator norm geometry explains width scaling, row and column normalization, and hyperparameter transfer.
Paper ↗
Statistics · 2026
Fragility of Interpolators
Heavy tailed risk and high dimensional large deviations reveal failure modes hidden by benign average case behavior.
Paper ↗
Feature Geometry · ICLR 2022
An Unconstrained Layer Peeled Perspective on Neural Collapse
Studies neural collapse through the implicit bias of gradient flow and the optimization geometry of an unconstrained model of features and classifiers.
Paper ↗
04
Agentic Mathematical Reasoning
Representations and search procedures that help AI systems discover, verify, and communicate mathematical structure.
Read more →
Selected works (1)
Probability · 2026
Signed BAR Conjecture
Uniqueness in the Harrison–Reiman class and a completely S class obstruction for a longstanding problem in reflected Brownian motion.
Paper ↗
05
Differential Equations for ML
Numerical differential equations, optimal control, and mean field limits provide principles for neural network architecture, efficient training, and optimization theory.
Read more →
Selected works (3)
Network Architectures · ICML 2018
Beyond Finite Layer Neural Networks
Interprets neural architectures as numerical discretizations of differential equations and uses linear multistep methods to design more efficient residual networks.
Paper ↗
Optimal Control · NeurIPS 2019
You Only Propagate Once (YOPO)
Formulates adversarial training as a differential game and uses Pontryagin’s maximum principle to reduce repeated propagation through the full network.
Paper ↗
Optimization Theory · ICML 2020
A Mean Field Analysis of Deep ResNet and Beyond
Develops a continuum model of deep residual networks and establishes optimization guarantees through mean field analysis and overparameterization from depth.
Paper ↗