Loading Events

« All Events

Virtual Event Virtual Event
  • This event has passed.

Ph.D. (Engg): Development of Learning-based Strategies for Reconnaissance with Multi-Robot Systems

September 22 @ 10:30 AM - 12:30 PM

Virtual Event Virtual Event

Reconnaissance is the contest for information about a territory: one side tries to observe an asset or a guarded area, the other tries to deny it. Low-cost drones have made this contest cheap and constant, and protecting critical infrastructure against it is now a crucial problem, because an intruder no longer needs to strike a target to threaten it. This thesis develops learning-based strategies for multi-robot reconnaissance from both of its perspectives. In the defense perspective, a team of defenders must deny reconnaissance of a protected territory by intercepting intruders before they cross its perimeter. In the adversarial perspective, an agent inside a guarded region must escape to a safe area with the information it has gathered before it is neutralized. Both are hard for the same real-world reasons: opponents arrive from any direction at any time, so the environment is non-stationary; each robot senses only a limited range, so the state is partially observed; communication may be unavailable; and the opponent’s numbers and strategy are unknown. Classical game-theoretic and assignment-based methods assume away one or more of these conditions, which motivates strategies that learn from local observations.
For the defense perspective, the thesis first formulates perimeter defense as a decentralized assignment learning problem and develops the Context-aware Deep Assignment Network (CDAN). Each defender encodes its limited field of view as a spatio-temporal context map of past observations, the present, and a predicted future with position uncertainty; pseudo-values around each intruder counter the sparsity of this map and allow a 3D convolutional network, trained by imitating a centralized solution, to converge and to be reused by the whole team. CDAN captures about 6% more intruders than the best decentralized baseline (73.4% against 67.5%) and generalizes over team size, perimeter length, and intruder speed and maneuvers. Because assignment quality inherits the reliability of the communication channel, the second contribution, CARE (Communication-free planning using Adaptive Regions of Engagement), resolves the assignment from each defender’s own observations: a defender senses over its full detection radius but commits only within an engagement region set by the distance to its nearest observed teammate, and drifts its rest position toward the arrival directions it has itself observed. Without a single message, CARE matches or exceeds communicating baselines over about 200 scenarios and 38,600 seed-paired episodes, and is demonstrated on a Crazyflie quadrotor team.
For the adversarial perspective, the third contribution gives the first reinforcement learning formulation of the confinement escape problem, with a constant-size LiDAR-based state that is independent of the number of pursuers and the shape of the region, and proposes Scaffolding Reflection based Reinforcement Learning (SR2L), in which a simple motion-planner scaffold guides the learner only when its suggestion is clearly better, so that it can accelerate learning but never damage it. SR2L converges in about half the episodes of standalone actor-critic methods and escapes faster against three pursuit strategies, with the lowest variance in every case.
Because no single learned policy is best under all conditions, the fourth contribution fuses a pool of pre-trained policies online in a non-stationary environment. Standard multiplicative-weights fusion collapses onto one policy and performs worse than doing nothing. Regularized Reward-aware Online LEarning (ROLE) adapts the fusion weights from the reward of the executed action alone, without ground truth, and bounds every weight with an anti-collapse cap. ROLE raises perfect escape runs from 41% for the best single policy to 68% on healthy pools and degrades gracefully when rogue policies contaminate the pool.
Together, these contributions treat reconnaissance in multi-robot systems as one problem seen from two perspectives, and show that learning from local observations, with minimal or no information exchange, can both deny reconnaissance of a protected territory and accomplish it from inside a guarded one.

Speaker : Vignesh Gurumurthy

Research Supervisor: Prof. Suresh Sundaram

 

Details

Date:
September 22
Time:
10:30 AM - 12:30 PM
Event Category:
Watch

Other

Speaker
 Vignesh Gurumurthy
Scroll to Top