Ming Shi Safe learning and online optimization for networked autonomous systems.
Reinforcement learning, online optimization, and bandits under hard constraints, limited information, and costly adaptation.
I develop algorithms and fundamental performance guarantees for systems that must learn and adapt without compromising safety, reliability, or resource efficiency.
Assistant Professor of Electrical and Computer Engineering at UB
Affiliated with the Institute for Artificial Intelligence and Data Science.
Ming Shi received his Ph.D. degree in Electrical and Computer Engineering from Purdue University in 2022, advised by Prof. Xiaojun Lin.
From 2022 to 2024, he was a Post-Doctoral Scholar in Electrical and Computer Engineering at The Ohio State University and was affiliated with the NSF AI-EDGE Institute, advised by Prof. Ness B. Shroff and Prof. Yingbin Liang.
Safe learning, resource-aware adaptation, and decision-making with limited information.
These directions address how networked and autonomous systems can learn under uncertainty while satisfying safety, information, and resource constraints.
Learning under instantaneous hard constraints
Algorithms and limits for reinforcement learning when unsafe exploration is unacceptable and guarantees must hold throughout learning—not only on average or asymptotically.
Adaptation with switching and reconfiguration costs
Learning and optimization for systems where changing actions, models, schedules, or configurations consumes time, energy, bandwidth, or operational capacity.
Learning from partial state and imperfect feedback
Decision-making with partial state information, imperfect preferences, conversational queries, and heterogeneous feedback across distributed and multi-objective systems.
Results in safe learning, partial observability, and resource-aware decision-making.
These papers illustrate the program’s progression from fundamental guarantees to algorithms for networked and autonomous systems.
Bi-Level Online Provisioning and Scheduling with Switching Costs and Cross-Level Constraints
24th International Symposium on Modeling and Optimization in Mobile, Ad hoc, and Wireless Networks, Columbus, Ohio, USA, June 2026. [WiOPT Best Paper Award Runner-Up]
Reinforcement Learning with Partial Online State Information in POMDPs: Regret Bounds and Limits
IEEE Transactions on Information Theory, May 2026, DOI: 10.1109/TIT.2026.3694700. [IEEE TIT]
Provably Efficient Reinforcement Learning for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces
42nd International Conference on Machine Learning, Vancouver, Canada, July 2025. [ICML] (Acceptance rate: 26.9%.)
Power-of-2-Arms for Adversarial Bandit Learning with Switching Costs
IEEE/ACM Transactions on Networking, vol. 33, no. 3, pp. 1112-1127, June 2025, DOI: 10.1109/TON.2024.3522073. [IEEE/ACM ToN]
Theoretical guarantees for dynamic networked and autonomous systems.
The work connects learning-theoretic analysis with operational constraints from wireless, edge-AI, data-center, distributed, and human-in-the-loop systems.
Learning-theoretic guarantees
Regret, competitive analysis, sample complexity, impossibility results, and performance bounds characterize what is achievable and at what information or resource cost.
Nonstationary and adversarial environments
Models account explicitly for drift, adversarial inputs, partial observability, regime changes, and cross-level constraints.
Network resource allocation and coordination
Resource allocation, scheduling, communication, sensing, and distributed coordination connect algorithmic choices to system-level reliability and efficiency.
Preference-aware and human-in-the-loop learning
Preference feedback, conversational queries, safety constraints, and multi-objective learning support systems that better reflect human goals and operational requirements.
Awards, invited talks, and conference presentations.
Use the filters to view honors, invited talks, research presentations, and professional leadership activities.
Conference presentation: “Probe-then-Commit Multi-Objective Bandits: Theoretical Benefits of Limited Multi-Arm Feedback” — IEEE/IFIP WiOpt, Columbus, OH, June 2026.
PresentationInvited talk: “Reinforcement Learning under Switching Costs, Adversarial Environments, and Multi-Source Imperfect Preference Feedback” — Rochester Institute of Technology (RIT), Rochester, NY, April 2026.
Invited talkInvited talk: “Partially Observable Reinforcement Learning with Partial Online State Information” — University of California San Diego (UCSD), San Diego, CA, February 2026.
Invited talkSession chair: Federated and Distributed Learning — ACM MobiHoc, Houston, TX, October 2025.
LeadershipConference presentation: “Online Learning for Optimizing AoI-Energy Tradeoff under Unknown Channel Statistics” — ACM MobiHoc, Houston, TX, October 2025.
PresentationLightning talk: “Power-of-M in Reinforcement Learning for Autonomous Edge-AI Systems” — ACM SIGMETRICS, Stony Brook, NY, June 2025.
PresentationPanel talk: “Thriving in Today's Job Market: Finding, Excelling, and Evolving in Your Career” — Artificial Intelligence Modeling, Analysis, and Control of Complex Systems Workshop (AIMACCS), Columbus, OH, May 8–9, 2025.
LeadershipConference presentation: “Designing Near-Optimal Partially Observable Reinforcement Learning” — IEEE Military Communications Conference, Washington, DC, October 2024.
PresentationInvited talk: “RL for Networking: Safety, Adversarial Inputs, Partial Observability, and Human Feedback” — Arizona State University, Tempe, AZ, May 2024.
Invited talkInvited talk: “Autonomous Networked Systems: Adversarial Online RL Under Limited Defender Resources” — New York University, New York, NY, November 2023.
Invited talkInvited talk: “AI-Powered Autonomous Systems with Switching Costs: Power-of-2-Arms” — Massachusetts Institute of Technology, Cambridge, MA, May 2023.
Invited talkInvited talk: “Near-Optimal Adversarial Reinforcement Learning with Switching Costs” — California Institute of Technology, Pasadena, CA, July 2023.
Invited talkInvited talk: “RL under Instantaneous Hard Safety Constraints and POMDPs: From Wireless Communications to Smart Health” — Northeastern University, Boston, MA, July 2023.
Invited talkConference presentation: “A Near-Optimal Algorithm for Safe Reinforcement Learning Under Instantaneous Hard Constraints” — International Conference on Machine Learning, Honolulu, HI, July 2023.
PresentationConference presentation: “Near-Optimal Adversarial Reinforcement Learning with Switching Costs” — ICLR Spotlight presentation, May 2023.
HonorSpotlight paper (notable top 25%), International Conference on Learning Representations, January 2023.
HonorConference presentation: “Leveraging Synergies Between AI and Networking to Build Next Generation Edge Networks” — IEEE Conference on Collaboration and Internet Computing, December 2022.
PresentationConference presentation: “Power-of-2-Arms for Bandit Learning with Switching Costs” — ACM MobiHoc, Seoul, South Korea, October 2022.
PresentationInvited talk: “Power-of-2-Arms for Bandit Learning with Switching Costs” — California Institute of Technology, Pasadena, CA, May 2022.
Invited talkConference presentation: “Combining Regularization with Look-Ahead for Competitive Online Convex Optimization” — IEEE INFOCOM, May 2021.
PresentationBilsland Dissertation Fellowship, Purdue University, April 2021.
HonorIEEE INFOCOM Student Conference Award, U.S. National Science Foundation, March 2021.
HonorConference presentation: “On the Value of Look-Ahead in Competitive Online Convex Optimization” — ACM SIGMETRICS / IFIP Performance, Phoenix, AZ, June 2019.
PresentationACM SIGMETRICS Student Travel Grant, U.S. National Science Foundation, May 2019.
HonorConference presentation: “Competitive Online Convex Optimization with Switching Costs and Ramp Constraints” — IEEE INFOCOM, Honolulu, HI, April 2018.
PresentationIEEE INFOCOM Student Travel Grant, U.S. National Science Foundation, March 2018.
HonorProspective students and research collaborators.
Research inquiries are most useful when they identify a specific problem, publication, or research direction connected to the program.