RL4SN

Frugal Reinforcement Learning for Stochastic Networks

Objectives of the chair

Markov decision processes (MDPs) and their equivalents in reinforcement learning have been highly effective in solving large-scale problems over the past two decades. However, their success often depends on exceptional computational resources, limiting their applicability in contexts with restricted data volume or computing power.

In contrast with this general trend, RL4SN aims to develop learning algorithms suitable for situations where data is scarce or computational power is limited, with a focus on stochastic networks. These are problems that are relevant from a practical perspective -for instance data centers provide the main infrastructure to support Internet applications and scientific computations– but for which the RL techniques developed so far are not directly applicable.

For instance, Stochastic networks may have sparse and rare non-zero rewards, not all policies are stable (in the sense of keeping the number of jobs bounded), optimal policies exhibiting clear structures in disjoint regions, and so-called index policies are known to perform very well. RL4SN is motivated by the potential for significant improvement that these properties offer to learning algorithms.

The main objective of RL4SN is thus to leverage the specific structures of the underlying MDPs of stochastic network problems to develop tailored learning algorithms.

To achieve this ambitious goal, RL4SN is organized into three main tasks, each targeting a distinct objective: (i) improving the exploration of learning algorithms in stochastic networks, (ii) creating more efficient and frugal learning algorithms, and (iii) developing ensemble algorithms to effectively learn indexing policies in stochastic networks.

Research objectives

  • Devise algorithms that can improve exploration in stochastic networks, while reducing the data needed to learn the optimal policy
  • Exploiter le processus de décision de Markov sous-jacent pour développer des algorithmes d’apprentissage plus efficaces.
  • Advance the applicability of index policies to continuous action
  • Apply function approximation and reinforcement learning techniques in calculating index policies.

Main scientific tasks

  • Improving exploration through particle systems
  • Optimal policy as a classification problem
  • Function approximation and index policies

Collaborations

The project has obtained the support of the industrial partner Nokia Bell Labs represented by the research scientist Lorenzo Maggi. The participation of NOKIA Bell Lab stems from the current interest of L. Maggi in learning algorithms and their use for resolving very concrete problems arising in the area of communication networks.

The interactions of the PhD students with the industrial partner NOKIA
will give them a feeling for the research environment and problems tackled in the industry.