
StageInformatiqueCEA
CEA Saclay
France
Stage Steering LLM to inhibit biases- Saclay-H/F Multimodal Large Language Models (LLMs) are increasingly capable of processing and integrating information from multiple modalities, including text, images, audio, and speech. However, these models can inherit and amplify biases across modalities, potentially affecting the fairness, reliability, and robustness of their outputs. This internship will investigate steering-based approaches to identify and inhibit such biases at the model level, with the goal of developing more controllable and robust multimodal LLMs. As an intern at the CEA, you will have the opportunity to work in a world-renowned research environment. Our teams consist of passionate and dedicated experts, providing an environment conducive to learning and collaboration. You will have access to state-of-the-art equipment and top-tier research resources to carry out your assignments. The work performed may potentially lead to a scientific publication. ContextThrough the thesis of Clément Cornet, the team has already developped several approaches of steering and other works in mechanistic interpretability [1,2]. A large part of the work is integrated into a light python library that can serve as basis for the work. What do we expect from you ?The intern will work on the following tasks :Conduct a literature review on methods for bias in inhibition in multimodal LLMs with steering Conduct experiments with available steering approaches to inhibate biases, including a rigorous quantitative evaluation on well chosen models and modalitiesDevelop novel approaches to inhibate biases with steering, in particular to determine its strength automaticallyDevelop a demonstrator to showcase the work carried out Depending on the profile and motivation of the intern, the work may lead to a scientific publication and may be pursued with a PhD focused on a similar topic. The person will work in collaboration with Clement Cornet, Hervé Le Borgne, Romaric Besançon and possibly other researchers of the lab, depending on the direction of the work. [1] Cornet et al (2025) Explaining How Visual, Textual and Multimodal Encoders Share Concepts, CoRR:2507.18512 [2] Cornet et al (2026) The Deleuzian Representation Hypothesis, ICLR #Cea List Profil :Students in their final year of studies (M2 or last year of engineering school)Strong foundations in machine learning and deep learningInterest in mechanistic interpretability and bias of AI modelsPython proficiency in pytorch
Source : CEA · Récupérée le 2 octobre 2026