Luca Denti, Paola Bonizzoni, Brona Brejova, Rayan Chikhi, Thomas Krannich, Tomas Vinar, Fereydoun Hormozdiari. Pangenome graph augmentation from unassembled long reads. Technical Report 2025.02.07.637057, bioRxiv, 2025.

Download preprint: not available

Download from publisher: https://doi.org/10.1101/2025.02.07.637057

Related web page: not available

Bibliography entry: BibTeX

Abstract:

Pangenomes are becoming increasingly popular data structures for genomics 
analyses due to their ability to compactly represent the genetic diversity 
within populations. Constructing a pangenome graph, however, is still a 
time-consuming and expensive process. A promising approach for pangenome 
construction consists of progressively augmenting a pangenome graph with 
additional high-quality assemblies. Currently, there is no method for 
augmenting a pangenome graph with unassembled reads from newly sequenced 
samples without first aligning the reads to a reference genome and 
performing variant calling and genotyping on the new individuals.

In this work, we present the first assembly-free and mapping-free approach 
for augmenting an existing pangenome graph using unassembled long reads from 
an individual not already present in the pangenome. Our approach consists of 
finding sample specific sequences in reads using efficient indexes, 
clustering reads corresponding to the same novel variant(s), and then 
building a consensus sequence to be added to the pangenome graph for each 
variant separately.

Using simulated reads based on Human Pangenome Reference Consortium (HPRC) 
assemblies, we demonstrate the effectiveness of the proposed approach for 
progressively augmenting the pangenome with long reads, without the need for 
de novo assembly or predicting genetic variants of the new sample. The 
software is freely available at https://github.com/ldenti/palss.