
ThèseInformatiqueInria
Inria – COMPACT (Rennes)
France
vendredi 20 novembre 2026
2 300€ per month
Type de contrat : CDD Contexte et atouts du poste Context The volume of data generated worldwide is projected to approach 180 zettabytes (ZB) per year by 2025 [1]. However, current storage technologies face significant limitations in scaling sustainably to such volumes. One promising solution to address these challenges is DNA-based data storage, which offers several advantages, including extremely high data density, long-term retention, and low energy consumption [2]. From a density perspective, DNA can theoretically store up to 10 terabytes per mm³, which would allow all data generated throughout human history to be stored within a cube of approximately 30 cm per side [3]. In terms of retention, DNA can remain readable for centuries under suitable conditions, whereas conventional storage media typically degrade within decades [3]. Furthermore, DNA storage is energy-efficient, as it can be preserved at ambient temperature provided it is protected from light and humidity. Mission confiée Goal The goal of the project is to develop an algorithm to allow robust retrieval of data in the context of DNA-based data storage. Principales activités Challenges and envisaged approach Despite its potential, making DNA a practical and efficient storage medium requires overcoming several key challenges: (i) Data transformation: converting digital data into a quaternary alphabet (A, C, G, T). (ii) DNA synthesis: writing data through the physical synthesis of DNA strands. (iii) DNA sequencing: reading the stored data by sequencing DNA. (iv) Data retrieval: reconstructing the original digital data from the sequenced symbols. This PhD project focuses on the first and fourth challenges, by developing joint compression and error-correction algorithms that are robust to sequencing errors arising during step (iii). Efficient DNA storage critically depends on fast sequencing technologies, which often come at the cost of increased error rates. For example, nanopore sequencing, developed by Oxford Nanopore Technologies (ONT), enables real-time analysis but introduces relatively high error rates [4,5]. Unlike traditional sequencing technologies, nanopore sequencing produces not only substitution errors but also insertion and deletion errors. Deletions are particularly challenging, as they differ from erasure errors where the position of missing data is known e.g., packet losses in digital communications). In the case of deletions, neither the existence nor the location of the missing symbols is known, significantly complicating error correction. Two main approaches have been proposed in the literature to address these errors. The first consists of avoiding error-prone patterns by designing constrained DNA sequences, such as limiting homopolymers or enforcing balanced GC content [6]. The second approach embraces sequencing errors and focuses on correcting them using coding techniques [7]. These strategies are typically considered contradictory, as one seeks to prevent errors while the other assumes their presence [8]. In this project, we propose to explore an alternative route that combines both strategies: avoiding the majority of sequencing errors while correcting the remaining ones. This will be achieved by jointly structuring the compressed DNA stream and designing error-correction mechanisms tailored to nanopore sequencing. In particular, we will exploit the properties of the quaternary alphabet and the interaction between compression and error-correction algorithms. The work will build upon concepts from non-binary channel coding [10,11] and DNA-specific transcoding techniques [6], while also integrating efficient post-sequencing data retrieval methods such as those proposed in [9]. Bibliography [1] David Reinsel-John Gantz-John Rydning, John Reinsel, and John Gantz. “The digitization of the world from edge to core.Framingham: International Data Corporation”, 16:1–28, 2018. [2] Luis Ceze, Jeff Nivala, and Karin Strauss. “Molecular digital data storage using DNA”. NatureReviews Genetics, 20(8):456–466, 2019. [3] Victor Zhirnov, Reza M Zadegan, Gurtej S Sandhu, George M Church, and William L Hughes. “Nucleic acid memory”. Nature materials, 15(4):366–370, 2016. [4] Delahaye, Clara, and Jacques Nicolas. “Nanopore MinION Long Read Sequencer: An Overview of Its Error Landscape,” November 23, 2020. https://hal.inria.fr/hal-03123133. [5] ———. “Sequencing DNA with Nanopores: Troubles and Biases.” PLoS ONE, October 1, 202 [6] S. Al Sayyed and A. Roumy and T. Maugey. ``Efficient constraining of transcoding for DNA-based data storage'', IEEE International Conference on Image Processing (ICIP), 2025. [7] R. Khabbaz, M. Antonini, S. Kas Hanna, Marker Guess & Check Plus (MGC+): An Efficient Short Blocklength Code for Random Edit Errors, International Symposium on Topics in Coding (ISTC), 2025. [8] F. Weindel, A. L. Gimpel, R. N. Grass and R. Heckel, ”Embracing errors is more effective than avoiding them through constrained cod
Source : Inria · Récupérée le 30 septembre 2026