
Post-dochistoireInria
Inria – ALMANACH (Paris)
France
lundi 30 novembre 2026
Type de contrat : CDD Contexte et atouts du poste This position is part of the ANR project ROMAM², led by Thibault Clérice at Inria Paris within the ALMAnaCH project-team (Automatic Language Modelling and Analysis & Computational Humanities, led by Benoît Sagot). ROMAM² treats pre-editorial normalisation (PEN) of graphemic automatic text recognition (ATR) output as a dedicated NLP task. PEN is a traceable process in which every editorial inference (abbreviation expansion, post-correction of recognition errors, spelling normalisation) stays anchored to the manuscript it comes from. The project works on two medieval languages: Latin, which is heavily abbreviated, and Old French, whose spelling varies in linguistically meaningful ways. A central claim of the project is that editorial normalisation is not neutral. Printed critical editions silently expand abbreviations and regularise spelling. In doing so, they erase variation that is evidence for the history of the language. Dees' quantitative geography of Old French and its successors rest largely on such editions, and Morin has shown how this can distort dialectal conclusions. However, no study has yet measured this distortion on a controlled parallel corpus. This postdoctoral position is designed to produce that study and the gold data it requires. The postdoctoral researcher will be supervised by Thibault Clérice. They will work closely with: • the project's PhD candidate in NLP, co-supervised by Thibault Clérice, Benoît Sagot and Rachel Bawden. The PhD candidate will use the gold data and evaluation framework produced by the postdoc to train and evaluate normalisation models. • David Smith (Northeastern University), a specialist in aligning noisy historical data. • the ANR JCJC project Phil•IA, coordinated by Ariane Pinche at CIHAM (UMR 5648, ENS de Lyon). A regular collaboration is expected on Old French graphemic transcription, digital editing and TEI encoding. The position is based at Inria Paris, within ALMAnaCH. The team brings together researchers in NLP, language modelling and computational humanities, and offers a rare environment for philologists and linguists who want to work directly with NLP researcher. The postdoc will also benefit from existing community resources developed by the team: the CATMuS dataset (the largest ATR dataset for medieval manuscripts), the CoMMA corpus (3.3 billion tokens of Latin and Old French from over 32,000 manuscripts) and the upcoming work on Biblissima-Textes. The position is for 18 months, starting March 2027. Bibliography • Clérice, T., Bawden, R., Glaise, A., Pinche, A., & Smith, D. (2026). Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin. In Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC 2026. https://arxiv.org/abs/2602.13905 • Clérice, T., Pinche, A., Vlachou-Efstathiou, M., Chagué, A., Camps, J.-B., et al. (2024). CATMuS Medieval: A multilingual large-scale cross-century dataset in Latin script for handwritten text recognition and beyond. In Proceedings of ICDAR 2024 (LNCS 14806, pp. 174–194). Springer. https://doi.org/10.1007/978-3-031-70543-4_11 • Clérice, T., Gabay, S., Vlachou-Efstathiou, M., Pinche, A., & Sagot, B. (2026). CoMMA, a Large-scale Corpus of Multilingual Medieval Archives. In Proceedings of the Fifteenth Language Resources and Evaluation Conference. ELRA. https://inria.hal.science/hal-05299220 • Dees, A. (1985). Dialectes et scriptae à l'époque de l'ancien français. Revue de Linguistique Romane, 49(193–194), 87–117. • Morin, Y. C. (2006). Histoire du corpus d'Amsterdam : le traitement des données dialectales. In Le Nouveau Corpus d'Amsterdam. Actes de l'atelier de Lauterbad. • Scheer, T., & Brun-Trigaud, G. (2022). L'atlas Dees électronique. Concordial, Grenoble. https://hal.science/hal-03912660 • Kuparinen, O., & Scherrer, Y. (2024). Corpus-based dialectometry with topic models. Journal of Linguistic Geography, 12(1), 1–12. • Pinche, A. (2022). Guide de transcription pour les manuscrits du Xe au XVe siècle. https://hal.archives-ouvertes.fr/hal-03697382 • Duval, F. (2012). Transcrire le français médiéval : de l'« Instruction » de Paul Meyer à la description linguistique contemporaine. Bibliothèque de l'École des chartes, 170(2), 321–342. Mission confiée The postdoctoral researcher, trained in philology or historical linguistics with skills in digital humanities, will build the philological foundations and evaluation resources of ROMAM² and carry out a controlled study of how editorial practices affect the dialectometry of Old French. The work involves: • Building a multi-layer gold corpus from the Nouveau Corpus d'Amsterdam (NCA, formerly Dees' corpus). Using Tobias Scheer's Atlas Dees Électronique mapping between NCA editions and their source manuscripts, the postdoc will sample each text (~500 words per document, ~100,000 words in total) and transcribe them in eScri
Source : Inria · Récupérée le 3 octobre 2026