
ThèseInformatiqueInria
Inria – NEO (Sophia Antipolis)
France
dimanche 25 octobre 2026
Duration: 36 months
Type de contrat : CDD Contexte et atouts du poste This PhD thesis is part of the Inria–Hivenet Challenge Cupseli: Collaborative Unified Platform for a Scalable and Efficient Learning Infrastructure. Hivenet, which will soon become Antimatter, aims to develop scalable, efficient, and secure solutions for running AI training and inference workloads on distributed, heterogeneous, and volatile computing resources. The PhD candidate will be hired by Hivenet and hosted mostly by the NEO project-team at the Inria Centre at Université Côte d’Azur, in Sophia Antipolis. The thesis will be carried out in close collaboration with the Inria ARGO project-team and with Hivenet engineers. The work will focus on the design of new distributed training algorithms tailored to environments where participants share computing and storage resources only for limited periods of time. The research activity will be supervised by: Giovanni Neglia, Inria NEO, https://www-sop.inria.fr/members/Giovanni.Neglia/ Laurent Massoulié, Inria ARGO, https://www.di.ens.fr/laurent.massoulie/ Igor Carrara, Hivenet / Antimatter, https://www.linkedin.com/in/carraraig/ The candidate will benefit from a highly interdisciplinary environment combining expertise in machine learning, distributed optimization, federated learning, stochastic systems, and large-scale distributed infrastructures. Mission confiée Context The Hivenet distributed platform allows participants to share computing and storage resources for limited periods of time. These resources may become available only at specific moments, or may be withdrawn unexpectedly. This creates a highly dynamic training environment, where the system must continuously decide which resources to use, when to use them, and how to exploit them efficiently. This challenge is closely related to cross-device federated learning, where client participation often depends on external constraints such as battery level, WiFi connectivity, device usage, or time-of-day patterns. In operational federated learning systems, volatility is often handled by recruiting a large number of clients at each training round and proceeding once enough responses have been received. While effective at large scale, this strategy can be inefficient and may introduce statistical biases in the learning process. The Hivenet setting differs from standard cross-device federated learning in several important ways. First, the number of available participants may be smaller, making it costly or impossible to rely on massive redundancy. Second, while federated learning usually assumes that data cannot be moved across clients, Hivenet may allow portions of data to be transferred to selected participants. Third, the availability of Hivenet resources may sometimes be known in advance or predicted with some accuracy, creating new opportunities for resource-aware training algorithms. Beyond training time and model accuracy, such systems also raise questions of energy efficiency and environmental impact. Recent work on carbon-aware federated learning has shown that client and time-slot scheduling can exploit temporal and geographical variations in carbon intensity to reduce the carbon footprint of training [arputharaj25]. This perspective is particularly relevant in a distributed infrastructure such as Hivenet, where resource availability, energy characteristics, and communication conditions may vary significantly over time. These specific features call for new distributed learning methods that jointly account for statistical efficiency, system constraints, resource availability, communication cost, and, when relevant, energy or carbon-related criteria. Research objectives The objective of this PhD is to develop new distributed training algorithms for volatile and resource-constrained environments such as Hivenet, taking into account not only training time and accuracy, but also communication cost and, when relevant, energy or carbon-related objectives. A first objective will be to model the availability of participating resources. The candidate will study how to characterize resource availability patterns using real-world traces, measurements, or insights from Hivenet engineers. This modeling phase will aim to capture both predictable availability, such as daily or weekly patterns, and unpredictable events, such as sudden resource withdrawals. A second objective will be to design training algorithms that exploit these availability models. The algorithms will decide which resources should participate in training, when they should be activated, and how they should be used. In particular, the thesis will study questions such as how much data should be transferred to a participant, how many local model updates should be performed before communication, and how to balance computation, communication, and statistical progress. A third objective will be to provide theoretical guarantees for the proposed methods. The candidate will analyze the convergence and tr
Source : Inria · Récupérée le 30 septembre 2026