Skip to content
deep-learning-tool-uses-long-read-sequencing-to-detect-cancer-mutations

Deep-Learning Tool Uses Long-Read Sequencing to Detect Cancer Mutations

By Bio-IT World News Staff 

September 2, 2026 | Analyzing genetic changes in cancer cells can provide insight into how tumors develop and inform precision oncology. However, identifying cancer-related mutations can be challenging, particularly when they occur at low frequencies or in complex regions of the genome. Short-read sequencing can struggle with repetitive and difficult-to-map regions, but long-read sequencing can help resolve these regions. 

Researchers at The University of Hong Kong have developed ClairS, a tool designed to detect small cancer-related genetic mutations from long-read sequencing data. The team trained the model using synthetic data to overcome the limited availability of high-quality somatic mutation data. Led by Professor Ruibang Luo, assistant director (Learning Experience & Student Enrichment) and associate head of the department of AI and data science at the School of Computing and Data Science, the study was published this summer in Nature Methods (DOI: 10.1038/s41592-026-03152-4). 

The team had previously developed deep-learning germline variant callers, but somatic mutation detection presented a different obstacle because high-quality training data are scarce. ClairS addresses this challenge through a multistep workflow that combines phasing with two neural networks: a pileup-based model that examines groups of sequencing reads at a specific genomic position, and a full-alignment model that analyzes individual reads in greater detail. The tool also uses ancestral haplotype information to identify support for more distant variants on the same chromosome. 

A key feature of ClairS is its synthetic data training strategy. Rather than relying on real tumor samples, the researchers generated synthetic tumor and normal samples using real sequencing data. The researchers used variants present in one sample but that were absent from the other as simulated somatic mutations across different levels of tumor DNA, sequencing coverage, and normal-cell contamination. 

They found that carefully designed synthetic data, reliable training labels, and phasing enabled ClairS to make accurate predictions of real somatic variants across different sequencing coverages and tumor purities, even without using real tumor samples during training. Haplotype information from long reads also improved detection of low-frequency variants. 

As long-read sequencing continues to evolve, tools such as ClairS could help researchers uncover cancer mutations that have remained difficult to detect with conventional approaches.  

To learn more about the technology, its limitations, and its potential applications in cancer research, read the full story by Irene Yeh at Diagnostics World News. 

colind88

Back To Top