Mingcheng Lu, Fudan University Math BS'27
Teaching, Mentoring, and Publications.
Teaching
- Mathematical Data Science.
- [Spring 2025, 2024] TA, “Advanced Topics on Deep Learning” (CS 496), Northwestern University CS.
- [Spring 2026, 2025, 2024] TA, “Mathematical Foundations of Machine Learning” (CS 416 & STAT 435), Northwestern University CS & Stats & Data Science.
- [Winter 2025, 2022] TA, “Data Science Pipeline” (CS 326), Northwestern University CS.
- [Fall 2020] Instructor, “Vibrations, Waves, & Electromagnetism.” Notes & recordings upon request, University of Maryland College Park Physics.
- [2018–2021] TA for multiple courses (PHYS121, PHYS122, PHYS131, PHYS132, PHYS260, PHYS261, PHYS603...), University of Maryland College Park Physics.
- [Spring 2018] TA, “Gravity, S-Matrix & String Theory,” NTU Physics.
- [Fall 2017] TA, “Quantum Field Theory I,” NTU Physics.
Collaboration and Mentoring
In reverse chronological order.
Xiwen Zhang, Fudan University Math BS'27
Po-Chiao Lin, NTU Physics BS'25
- Universal Approximation, Compositional Generalization, and Algorithm Emulation All In-Context [ICML'26]
Venkat Sripad Ganti, Northwestern Math + Statistics + Economics BS'27
- Chain-of-Thought Gradient Descent [ICML'26]
Jennifer Yuntong Zhang, University of Toronto BAS'26 Engineering Science → Yale CS MS (Fall'26)
- In-Context Algorithm Emulation in Fixed-Weight Transformers [ICLR'26]
Hude Liu, Fudan University Math BS'25
- In-Context Algorithm Emulation in Fixed-Weight Transformers [ICLR'26]
- Attention Mechanism, Max-Affine Partition, and Universal Approximation [NeurIPS'25a]
- Universal Approximation with Softmax Attention [ICML'26]
Maojiang Su, University of Science and Technology of China (School of the Gifted Young) BS'25 → CS PhD at Northwestern (Fall'25)
- A Theoretical Analysis of Discrete Flow Matching Generative Models [arXiv]
- High-Order Flow Matching: Unified Framework and Sharp Statistical Rates [NeurIPS'25b]
- In-Context Deep Learning via Transformer Models [ICML'25b]
- Computational Limits of Low-Rank Adaptation (LoRA) for Transformer-Based Models [ICLR'25a]
Zoe Mehta, High School Outreach Student @ Vernon Hills High School → MIT (Class of 2029)
- Fast and Low-Cost Genomic Foundation Models via Outlier Removal [ICML'25a]
Sophia Pi, Northwestern CS + Economics + Mathematical Methods in Social Science BS'26 → UPenn CS PhD (Fall'26)
- Learning Manifold Data with Flow Matching [ICML'26]
- On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs) [NeurIPS'24b]
- On Flow Matching KL Divergence [arXiv]
Chenghao Qiu, Tianjin University CS BS'25 → CS PhD study at TAMU (Fall'25)
- Fast and Low-Cost Genomic Foundation Models via Outlier Removal [ICML'25a]
Thomas Yuan-Lung Lin, High School Outreach Student @ WLSH'23 → NTU (transferred) → University of Washington Physics (Class of 2027)
- Latent Variable Estimation in Bayesian Black-Litterman Models [ICML'25e]
- On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis [ICML'24a]
Stephen Cheng, Northwestern EE BS'25 + CS MS'25 → CS PhD at UMD (Fall'25)
- Financial Data Prediction Models, Statistical Theory of Diffusion Models and Transformer
Teng-Yun Hsiao, NTU Physics BS'26
- In-Context Learning as Conditioned Associative Memory Retrieval [ICML'25c]
- Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models [ICML'24d]
Wei-Po Wang, NTU Physics BS'24 → Physics PhD at Johns Hopkins University (Fall'25)
- Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency [ICLR'25b]
- Outlier-Efficient Hopfield Layers for Large Transformer-Based Models [ICML'24b]
Morris Huang, NTU Physics MS'24 → CS PhD at North Carolina, Chapel Hill (Fall'25)
- On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality [ICLR'25c]
- BiSHop: Bi-Directional Cellular Learning for Tabular Data with Generalized Sparse Hopfield Model [ICML'24]
Hong-Yu Chen, NTU Physics MS'24 → CS PhD study at Northwestern (Fall'24)
- Universal Approximation of Softmax Attention [ICML'26]
- Outlier-Efficient Hopfield Layers for Large Transformer-Based Models [ICML'24b]
Yi-Chen Lee, NTU Physics BS'26 → CS PhD study at Johns Hopkins University (Fall'26)
- High-Order Flow Matching: Unified Framework and Sharp Statistical Rates [NeurIPS'25b]
- On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality [ICLR'25c]
Bo-Yu Chen, High School Outreach Student @ HSNU'23 → NTU Physics + CS (Class of 2027) with NTU Fu Bell Scholarship
- Nonparametric Modern Hopfield Models [ICML'25c]
- STanHop: Sparse Tandem Hopfield Model for Memory-Enhanced Time Series Prediction [ICLR'24]
- On Sparse Modern Hopfield Model [NeurIPS'23]
Chenwei Xu, MSCS'24 at Northwestern → Stats & DS PhD study at Northwestern (Fall'24)
- BiSHop: Bi-Directional Cellular Learning for Tabular Data with Generalized Sparse Hopfield Model [ICML'24c]
- On Sparse Modern Hopfield Model [NeurIPS'23]
- Beyond PID Controllers: PPO with Neuralized PID Policy for Proton Beam Intensity Control in Mu2e [ML4Phys Workshop @ NeurIPS'23] [READS Collaboration, Fermilab]
- Feature Programming for Multivariate Time Series Prediction [ICML'23]
Dennis Wu, MSCS'24 at Northwestern → CS PhD study at Northwestern (Fall'24)
- Provably Optimal Memory Capacity for Modern Hopfield Models [NeurIPS'24]
- Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models [ICML'24d]
- STanHop: Sparse Tandem Hopfield Model for Memory-Enhanced Time Series Prediction [ICLR'24]
- On Sparse Modern Hopfield Model [NeurIPS'23]
Zhenyu Pan, MSECE'24 at University of Rochester → CS PhD study at Northwestern (Fall'24)
Zhenji Wang, UMD'21 Math → MS'23 at Columbia University → ML PhD study at University of Tsukuba
- Differential Geometry and Heat Kernel Expansion, 2021 Spring
Resources
A nice PhD checklist from Aaditya Ramdas
- Checklist for research ethics
- Checklist for a well-rounded, balanced PhD experience
- Checklist for effectively writing papers
What are expected for graduate research
- A Note to Prospective Students by Bert Huang
- A Note to Prospective Students by Sanmi Koyejo
An open letter to graduate students and other procrastinators: it’s time to write
10 easy ways to fail a Ph.D.
Publications
Please see Google Scholar for the latest publications. (* denotes equal contribution)
2026
-
ICML'26
On Structure State-Space Duality
International Conference on Machine Learning, 2026.
-
ICML'26
In-Context Universal Approximation, Compositional Generalization, and Algorithm Emulation
International Conference on Machine Learning, 2026.
-
ICML'26
Chain-of-Thought Gradient Descent
International Conference on Machine Learning, 2026.
-
ICML'26
Universal Approximation with Softmax Attention
International Conference on Machine Learning, 2026.
-
ICML'26
Genome-Factory: An Integrated Python Library for Tuning, Deploying and Interpreting Genomic Models
International Conference on Machine Learning, 2026.
-
ICML'26
Learning Manifold Data with Flow Matching
International Conference on Machine Learning, 2026.
-
ICLR'26
In-Context Algorithm Emulation in Fixed-Weight Transformers
International Conference on Learning Representations, 2026.
-
KDD'26
POLO: Preference-Guided Multi-Turn Reinforcement Learning for Sample-Efficient Lead Optimization
ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2026.
2025
-
NeurIPS'25
Attention Mechanism, Max-Affine Partition, and Universal Approximation
Advances in Neural Information Processing Systems, 2025.
-
NeurIPS'25
High-Order Flow Matching: Unified Framework and Sharp Statistical Rates
Advances in Neural Information Processing Systems, 2025.
-
NeurIPS'25
Pareto-Optimal Energy Alignment for Designing Nature-Like Antibodies
Advances in Neural Information Processing Systems, 2025.
-
USENIX'25
Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Ethical Boundaries
USENIX Security Symposium, 2025.
-
ICML'25
Fast and Low-Cost Genomic Foundation Models via Outlier Removal
International Conference on Machine Learning, 2025.
-
ICML'25
Latent Variable Estimation in Bayesian Black-Litterman Models
International Conference on Machine Learning, 2025.
-
ICML'25
In-Context Deep Learning via Transformer Models
International Conference on Machine Learning, 2025.
-
ICML'25
In-Context Learning as Conditioned Associative Memory Retrieval
International Conference on Machine Learning, 2025.
-
ICML'25
Nonparametric Modern Hopfield Models
International Conference on Machine Learning, 2025.
-
ICLR'25
On Statistical Rates of Conditional Diffusion Transformers
International Conference on Learning Representations, 2025.
-
ICLR'25
Computational Limits of Low-Rank Adaptation for Transformer-Based Models
International Conference on Learning Representations, 2025.
-
ICLR'25
Fundamental Limits of Prompt Tuning: Universality, Capacity and Efficiency
International Conference on Learning Representations, 2025.
2024
-
NeurIPS'24
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers
Advances in Neural Information Processing Systems, 2024.
-
NeurIPS'24
Provably Optimal Memory Capacity for Modern Hopfield Models
Advances in Neural Information Processing Systems, 2024.
-
ICML'24
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
International Conference on Machine Learning, 2024.
-
ICML'24
BiSHop: Bi-Directional Cellular Learning for Tabular Data with Generalized Sparse Modern Hopfield Model
International Conference on Machine Learning, 2024.
-
ICML'24
Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models
International Conference on Machine Learning, 2024.
-
ICML'24
On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
International Conference on Machine Learning, 2024.
-
ICLR'24
STanHop: Sparse Tandem Hopfield Model for Memory-Enhanced Time Series Prediction
International Conference on Learning Representations, 2024.
2023
-
NeurIPS'23
On Sparse Modern Hopfield Model
Advances in Neural Information Processing Systems, 2023.
-
ICML'23
Feature Programming for Multivariate Time Series Prediction
International Conference on Machine Learning, 2023.
-
AAAI'23
Ising-traffic: Using Ising Machine Learning to Predict Traffic Congestion under Uncertainty
AAAI Conference on Artificial Intelligence, 2023.
-
IBIC'23
FPGA Architectures for Distributed ML Systems for Real-time Beam Loss De-blending
International Beam Instrumentation Conference, Report number: FERMILAB-CONF-23-281-AD, 2023.
-
ICALEPCS'23
Disentangling Beam Losses in the Fermilab Main Injector Enclosure Using Real-Time Edge AI
International Conference on Accelerator and Large Experimental Physics Control Systems, 2023.