智能化、基于匹配的基因组压缩算法（AMGC） (AMGC: Adaptive match-based genomic compression algorithm) - 专知论文

会员服务 ·

0

压缩算法 · 算法 · 压缩率 · 序列 · 错误率 ·

2023 年 4 月 3 日

AMGC: Adaptive match-based genomic compression algorithm

翻译：智能化、基于匹配的基因组压缩算法（AMGC）

Jia Wang,Yi Niu,Tianyi Xu,Mingming Ma,Dahua Gao,Guangming Shi

Motivation: Despite significant advances in Third-Generation Sequencing (TGS) technologies, Next-Generation Sequencing (NGS) technologies remain dominant in the current sequencing market. This is due to the lower error rates and richer analytical software of NGS than that of TGS. NGS technologies generate vast amounts of genomic data including short reads, quality values and read identifiers. As a result, efficient compression of such data has become a pressing need, leading to extensive research efforts focused on designing FASTQ compressors. Previous researches show that lossless compression of quality values seems to reach its limits. But there remain lots of room for the compression of the reads part. Results: By investigating the characters of the sequencing process, we present a new algorithm for compressing reads in FASTQ files, which can be integrated into various genomic compression tools. We first reviewed the pipeline of reference-based algorithms and identified three key components that heavily impact storage: the matching positions of reads on the reference sequence(refpos), the mismatched positions of bases on reads(mispos) and the matching failed reads(unmapseq). To reduce their sizes, we conducted a detailed analysis of the distribution of matching positions and sequencing errors and then developed the three modules of AMGC. According to the experiment results, AMGC outperformed the current state-of-the-art methods, achieving an 81.23% gain in compression ratio on average compared with the second-best-performing compressor.

翻译：动机：尽管第三代测序（TGS）技术取得了显著的进展，但比起TGS来说，下一代测序（NGS）技术在当前测序市场上仍然占主导地位，这是由于NGS的比TGS更低的错误率和更丰富的分析软件。 NGS技术产生了大量的基因组数据，包括短读取、质量值和读取标识符。因此，有效压缩这种数据已成为迫切需要，导致广泛的研究工作致力于设计FASTQ压缩器。以往的研究表明，质量值的无损压缩似乎已达到了极限。但对于reads部分的压缩仍具有较大的潜力。结果：通过调查测序过程的特征，我们提出了一种新的用于压缩FASTQ文件中reads的算法，该算法可以集成到各种基因组压缩工具中。我们首先审查了基于参考序列的算法的流程，并确定了三个关键组成部分，这些组成部分对存储产生了重大影响：reads在参考序列上的匹配位置（refpos）、reads上的错配位置（mispos）和匹配失败的reads（unmapseq）。为了减少它们的大小，我们对匹配位置和测序错误的分布进行了详细的分析，然后开发了AMGC的三个模块。根据实验结果，AMGC优于现有最先进的方法，平均压缩率相比于次佳性能压缩器提高了81.23％。

0

相关内容

压缩算法

Science Advances | 基于片段的、针对结构化表位的抗体计算设计

Science Advances | 基于片段的、针对结构化表位的抗体计算设计

专知会员服务

5+阅读 · 2022年12月5日

Meta最新WWW2022《联邦计算导论》教程，附77页ppt

Meta最新WWW2022《联邦计算导论》教程，附77页ppt

专知会员服务

60+阅读 · 2022年5月5日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【MIT】自监督几何感知，22页ppt，Self-supervised Geometric Perception

【MIT】自监督几何感知，22页ppt，Self-supervised Geometric Perception

专知会员服务

23+阅读 · 2021年6月3日

【NLP模型压缩方法综述】《A Survey of Methods for Model Compression in NLP》by Madison May

【NLP模型压缩方法综述】《A Survey of Methods for Model Compression in NLP》by Madison May

专知会员服务

43+阅读 · 2020年4月22日

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

专知会员服务

76+阅读 · 2020年4月10日

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

专知会员服务

15+阅读 · 2020年3月7日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【泡泡一分钟】用于评估视觉惯性里程计的TUM VI数据集

【泡泡一分钟】用于评估视觉惯性里程计的TUM VI数据集

泡泡机器人SLAM

11+阅读 · 2019年1月4日

【泡泡前沿追踪】跟踪SLAM前沿动态系列之IROS2018

【泡泡前沿追踪】跟踪SLAM前沿动态系列之IROS2018

泡泡机器人SLAM

29+阅读 · 2018年10月28日

AI实战圣经《Machine Learning Yearning》第1-52章中英文版pdf分享

AI实战圣经《Machine Learning Yearning》第1-52章中英文版pdf分享

深度学习与NLP

15+阅读 · 2018年9月8日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

LibRec 精选：推荐的可解释性[综述]

LibRec 精选：推荐的可解释性[综述]

LibRec智能推荐

10+阅读 · 2018年5月4日

【论文推荐】最新七篇图像检索相关论文—草图、Tie-Aware、场景图解析、叠加跨注意力机制、深度哈希、人群估计

【论文推荐】最新七篇图像检索相关论文—草图、Tie-Aware、场景图解析、叠加跨注意力机制、深度哈希、人群估计

专知

10+阅读 · 2018年4月22日

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

专知

15+阅读 · 2018年2月3日

拷贝数变异在中国遗传性耳聋人群中的分布及筛查策略研究

国家自然科学基金

0+阅读 · 2015年12月31日

面向众核处理器的HEVC并行编码关键技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

新生犊牛小肠上皮受体介导IgG转运通路的研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于核酸等温信号放大技术检测病原微生物的高灵敏SERS传感新方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

中国汉族人群尼古丁依赖的易感基因位点关联分析及易感基因功能研究

国家自然科学基金

0+阅读 · 2012年12月31日

胰安肽（Aglycin）治疗2型糖尿病的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

全基因组甲基化CpG岛扩增技术的建立及在食管癌早期诊断中的应用

国家自然科学基金

0+阅读 · 2011年12月31日

雄激素受体AR介导基因转录调控的辅调节因子筛选及生物学功能分析

国家自然科学基金

0+阅读 · 2008年12月31日

车用Ad Hoc网络的隐私与安全技术研究

国家自然科学基金

1+阅读 · 2008年12月31日

Model-Based Performance Analysis of the HyTeG Finite Element Framework

Arxiv

0+阅读 · 2023年5月24日

FedZero: Leveraging Renewable Excess Energy in Federated Learning

Arxiv

0+阅读 · 2023年5月24日

Regret Matching+: (In)Stability and Fast Convergence in Games

Arxiv

0+阅读 · 2023年5月24日

Prototype Adaption and Projection for Few- and Zero-shot 3D Point Cloud Semantic Segmentation

Arxiv

0+阅读 · 2023年5月23日

Notes on Causation, Comparison, and Regression

Arxiv

0+阅读 · 2023年5月23日

Adversarial Color Projection: A Projector-based Physical Attack to DNNs

Arxiv

0+阅读 · 2023年5月23日

Goldfish: No More Attacks on Proof-of-Stake Ethereum

Arxiv

0+阅读 · 2023年5月23日

Vision-Language Pre-training: Basics, Recent Advances, and Future Trends

Arxiv

28+阅读 · 2022年10月17日

A Survey of Model Compression and Acceleration for Deep Neural Networks

Arxiv

66+阅读 · 2019年9月8日

Matching Networks for One Shot Learning

Arxiv

10+阅读 · 2017年12月29日

VIP会员

文章信息

相关主题

相关VIP内容

Science Advances | 基于片段的、针对结构化表位的抗体计算设计

Science Advances | 基于片段的、针对结构化表位的抗体计算设计

专知会员服务

5+阅读 · 2022年12月5日

Meta最新WWW2022《联邦计算导论》教程，附77页ppt

Meta最新WWW2022《联邦计算导论》教程，附77页ppt

专知会员服务

60+阅读 · 2022年5月5日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【MIT】自监督几何感知，22页ppt，Self-supervised Geometric Perception

【MIT】自监督几何感知，22页ppt，Self-supervised Geometric Perception

专知会员服务

23+阅读 · 2021年6月3日

【NLP模型压缩方法综述】《A Survey of Methods for Model Compression in NLP》by Madison May

【NLP模型压缩方法综述】《A Survey of Methods for Model Compression in NLP》by Madison May

专知会员服务

43+阅读 · 2020年4月22日

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

专知会员服务

76+阅读 · 2020年4月10日

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

专知会员服务

15+阅读 · 2020年3月7日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《多域空战指挥体系：驾驭复杂性的艺术》

构建军事人工智能信任体系始于破除黑盒机制

《生态建模密码破译：建模与编程实践》美陆军最新报告

《战争形态演变：合成兵种防御主导模式探析》48页slides

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【泡泡一分钟】用于评估视觉惯性里程计的TUM VI数据集

【泡泡一分钟】用于评估视觉惯性里程计的TUM VI数据集

泡泡机器人SLAM

11+阅读 · 2019年1月4日

【泡泡前沿追踪】跟踪SLAM前沿动态系列之IROS2018

【泡泡前沿追踪】跟踪SLAM前沿动态系列之IROS2018

泡泡机器人SLAM

29+阅读 · 2018年10月28日

AI实战圣经《Machine Learning Yearning》第1-52章中英文版pdf分享

AI实战圣经《Machine Learning Yearning》第1-52章中英文版pdf分享

深度学习与NLP

15+阅读 · 2018年9月8日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

LibRec 精选：推荐的可解释性[综述]

LibRec 精选：推荐的可解释性[综述]

LibRec智能推荐

10+阅读 · 2018年5月4日

【论文推荐】最新七篇图像检索相关论文—草图、Tie-Aware、场景图解析、叠加跨注意力机制、深度哈希、人群估计

【论文推荐】最新七篇图像检索相关论文—草图、Tie-Aware、场景图解析、叠加跨注意力机制、深度哈希、人群估计

专知

10+阅读 · 2018年4月22日

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

专知

15+阅读 · 2018年2月3日

相关论文

Model-Based Performance Analysis of the HyTeG Finite Element Framework

Arxiv

0+阅读 · 2023年5月24日

FedZero: Leveraging Renewable Excess Energy in Federated Learning

Arxiv

0+阅读 · 2023年5月24日

Regret Matching+: (In)Stability and Fast Convergence in Games

Arxiv

0+阅读 · 2023年5月24日

Prototype Adaption and Projection for Few- and Zero-shot 3D Point Cloud Semantic Segmentation

Arxiv

0+阅读 · 2023年5月23日

Notes on Causation, Comparison, and Regression

Arxiv

0+阅读 · 2023年5月23日

Adversarial Color Projection: A Projector-based Physical Attack to DNNs

Arxiv

0+阅读 · 2023年5月23日

Goldfish: No More Attacks on Proof-of-Stake Ethereum

Arxiv

0+阅读 · 2023年5月23日

Vision-Language Pre-training: Basics, Recent Advances, and Future Trends

Arxiv

28+阅读 · 2022年10月17日

A Survey of Model Compression and Acceleration for Deep Neural Networks

Arxiv

66+阅读 · 2019年9月8日

Matching Networks for One Shot Learning

Arxiv

10+阅读 · 2017年12月29日

相关基金

拷贝数变异在中国遗传性耳聋人群中的分布及筛查策略研究

国家自然科学基金

0+阅读 · 2015年12月31日

面向众核处理器的HEVC并行编码关键技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

新生犊牛小肠上皮受体介导IgG转运通路的研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于核酸等温信号放大技术检测病原微生物的高灵敏SERS传感新方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

中国汉族人群尼古丁依赖的易感基因位点关联分析及易感基因功能研究

国家自然科学基金

0+阅读 · 2012年12月31日

胰安肽（Aglycin）治疗2型糖尿病的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

函数域中的Vinogradov中值定理

国家自然科学基金

0+阅读 · 2012年12月31日

全基因组甲基化CpG岛扩增技术的建立及在食管癌早期诊断中的应用

国家自然科学基金

0+阅读 · 2011年12月31日

雄激素受体AR介导基因转录调控的辅调节因子筛选及生物学功能分析

国家自然科学基金

0+阅读 · 2008年12月31日

车用Ad Hoc网络的隐私与安全技术研究

国家自然科学基金

1+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员