Candidatures closes. La date limite est passée : cette offre n'est plus proposée. Voir les offres ouvertes →

ThèseInformatiqueInria
Inria – WIMMICS (Sophia Antipolis)
France
dimanche 11 octobre 2026
Gross Salary per month: 2300 €
Type de contrat : CDD Contexte et atouts du poste INRIA is the French national research institute dedicated to computer science and applied mathematics and is a founding member of the World-Wide Web Consortium (W3C). The Inria centre at Université Côte d'Azur includes 42 research teams and 9 support services. The centre's staff (about 500 people) is made up of scientists of different nationalities, engineers, technicians, and administrative staff. The teams are mainly located on the university campuses of Sophia Antipolis and Nice as well as Montpellier, in close collaboration with research and higher education laboratories and establishments (Université Côte d'Azur, CNRS, INRAE, INSERM ...), but also with the regional economic players. The Wimmics team works on the topic of AI on the Web, in particular knowledge graphs and (linked) data representation and processing in the Semantic Web. Wimmics contributes to knowledge formalization and semantic-based methods to extract, control, query, validate, infer, explain and interact with knowledge in epistemic communities on the Web. Wimmics has been involved in a large number of European research projects and national projects. This PhD position takes place within the context of the national ANR project KGTWIN, with partners in Nantes and Nancy. Mission confiée Large-scale collaborative platforms such as Wikipedia have demonstrated that distributed communities can maintain shared knowledge at scale. Yet the emergence of Large Language Models (LLMs) redefines the conditions of collaboration: trained on human knowledge, these models can now produce and revise it, raising new questions about coherence, authorship, and trust. The project KGTWIN addresses these questions through a new paradigm of neuro-symbolic collaborative editing, in which free text and structured Knowledge Graphs (KGs) co-evolve as complementary representations of the same knowledge. The shared KG acts as the semantic backbone of human–AI collaboration. It maintains consistency across documents, exposes contradictions between texts that refer to the same concepts, and links contributors working on related entities. In this context, the KG provides a verifiable, traceable foundation for shared knowledge, transforming human–AI interaction from sequential exchange into continuous co-evolution. KGTWIN models this co-evolution as an iterative cycle combining controlled extraction, semantic synchronization, and grounded generation. Controlled extraction uses LLMs to instantiate ontologies and thesauri from collaboratively written texts while preserving provenance. Semantic synchronization detects inconsistencies and propagates updates between textual and graph representations. Grounded generation produces text fragments from coherent subgraphs. Together, these processes maintain alignment between natural-language narratives and structured knowledge, establishing a framework where reasoning, learning, and collaboration converge. In the KGTWIN project, one of the work packages will design a language-model-based pipeline able to extract structured knowledge from unstructured text while keeping users aware of the extraction process and its uncertainties. This work package aims to develop principled methods for extracting RDF knowledge from free text under the guidance of a predefined ontology. The goal is to ensure semantic fidelity, ontological compliance, and traceability of the extracted triples. It builds on the emergence of generative models guided by explicit semantic constraints. Language Models (LMs) can produce rich relational information from unstructured text, but they remain prone to hallucinations, schema violations, and reasoning errors. Conversely, traditional rule-based and supervised extraction systems offer strong guarantees of correctness but are costly to maintain and brittle when applied to new domains. This work package investigates how to combine symbolic knowledge with the language models, to obtain the best of both worlds: scalable extraction, but still traceable. The goal is to systematically evaluate extraction accuracy and recall according to the target ontologies and thesauri, the chosen model, and the pipeline used. The work package will produce metrics and visualization interfaces to expose what is extracted, what is missing, extraction confidence and highlight potential semantic conflicts with the current KG. Principales activités The PhD subject is about “extracting from and verbalizing RDF to text to support collaborative co-editing of wikis and their corresponding linked data” with the core questions: • Methods to establish SHACL shapes describing targeted graph patterns, including domain and range restrictions, property cardinalities, and integrity constraints. The method should help define the formal targets that will guide knowledge extraction, starting from ontologies and targeted texts. • Methods exploiting targeted graph patterns in providing training and validati
Source : Inria · Récupérée le 30 septembre 2026