Research overview

Safe learning and resource-adaptive optimization for networked systems.

I study algorithms and fundamental limits for sequential decisions under uncertainty, hard constraints, limited information, and costly adaptation.

Reinforcement learningOnline optimizationBandit learningNetwork optimizationDistributed systems
Diagram connecting uncertainty, safe decisions, resource adaptation, learning feedback, and networked autonomous systems
Observepartial, delayed, noisy, or strategic information
Learnmodels, policies, preferences, and uncertainty sets
Decideonline under safety, resource, and coupling constraints
Adaptto drift, adversaries, switching costs, and regimes
Guaranteereliability, efficiency, regret, and system performance
01Safe reinforcement learning

Learning under instantaneous hard constraints

Safety guarantees must hold throughout learning—not only after convergence.

Many autonomous systems cannot afford unsafe exploration. The goal is to learn effectively while respecting instantaneous constraints in uncertain, partially observed, non-convex, or adversarial environments.

Research questions

  • How can a learner explore without violating hard constraints?
  • Which safety guarantees are achievable under limited information?
  • How do adversarial transitions, modeling errors, and non-convex feature spaces alter the fundamental limits?

Selected publications

02Online optimization and bandits

Adaptation with switching, reconfiguration, and coupling costs

Changing actions, schedules, models, or configurations consumes resources.

Networked systems continuously reconfigure schedules, models, routes, sensing modes, and computing resources. This thrust studies how to balance immediate performance against the operational cost and cross-layer consequences of change.

Research questions

  • When is additional prediction, feedback, or limited multi-arm information worth its cost?
  • How should provisioning and scheduling be coordinated across time scales?
  • What competitive or regret guarantees remain possible with ramp and coupling constraints?

Selected publications

03Partial observability and preference learning

Learning from partial state and imperfect feedback

The learner may not observe the full state or know the objective exactly.

This thrust develops learning methods for partially observable systems and settings where preferences, objectives, or feedback sources are imperfect, heterogeneous, personalized, or costly to query.

Research questions

  • How much online state information is necessary in a POMDP?
  • How should imperfect preference sources be combined without losing statistical efficiency?
  • When do proactive conversational queries improve personalized multi-objective learning?

Selected publications

Application domains

Wireless, edge-AI, distributed, and autonomous systems.

These domains motivate the safety, information, communication, and resource constraints studied in the theoretical models.

Networks and computing

Wireless, edge-AI, and data-center systems

Resource allocation, age of information, sensing, communication, scheduling, and adaptive computing under uncertain demand and channel conditions.

Distributed autonomy

Multi-agent and networked decision systems

Coordination with limited communication, partial observability, dynamic constraints, strategic corruption, and changing operational regimes.

Emerging AI systems

AI-enabled cyber-physical and computing platforms

Extensions to secure systems, recommendation and preference learning, large language models, quantum networking, and quantum machine learning.

Publication record

Browse all publications and technical reports.

Search by topic, year, publication type, or status; open available papers; and copy formatted citations.