A crowdsourced set of curated structural variants for the human genome

Lesley M. Chapman; Noah Spies; Patrick Pai; Andrew Carroll; Marc L. Salit; Justin M. Zook

Official websites use .gov
A .gov website belongs to an official government organization in the United States.

Secure .gov websites use HTTPS
A lock ( ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.

PUBLICATIONS

A crowdsourced set of curated structural variants for the human genome

Published

June 19, 2020

Author(s)

Lesley M. Chapman, Noah Spies, Patrick Pai, Andrew Carroll, Marc L. Salit, Justin M. Zook

Abstract

A high quality benchmark for small variants encompassing 88 to 90% of the reference genome has been developed for seven Genome in a Bottle (GIAB) reference samples. However a reliable benchmark for large indels and structural variants (SVs) is more challenging. In this study, we manually curated 1235 SVs, which can ultimately be used to evaluate SV callers or train machine learning models. We developed a crowdsourcing appSVCuratorto help GIAB curators manually review large indels and SVs within the human genome, and report their genotype and size accuracy. SVCurator displays images from short, long, and linked read sequencing data from the GIAB Ashkenazi Jewish Trio son [NIST RM 8391/HG002]. We asked curators to assign labels describing SV type (deletion or insertion), size accuracy, and genotype for 1235 putative insertions and deletions sampled from different size bins between 20 and 892,149 bp. Expert curators were 93% concordant with each other, and 37 of the 61 curators had at least 78% concordance with a set of expert curators. The curators were least concordant for complex SVs and SVs that had inaccurate breakpoints or size predictions. After filtering events with low concordance among curators, we produced high confidence labels for 935 events. The SVCurator crowdsourced labels were 94.5% concordant with the heuristic-based draft benchmark SV callset from GIAB. We found that curators can successfully evaluate putative SVs when given evidence from multiple sequencing technologies.

Citation

PLOS Computational Biology

Pub Type

Journals

Download Paper

https://doi.org/10.1371/journal.pcbi.1007933

Local Download

Keywords

genomic measurements, machine learning

Citation

Chapman, L. , Spies, N. , Pai, P. , Carroll, A. , Salit, M. and Zook, J. (2020), A crowdsourced set of curated structural variants for the human genome, PLOS Computational Biology, [online], https://doi.org/10.1371/journal.pcbi.1007933 (Accessed July 28, 2025)

Issues

If you have any questions about this publication or are having problems accessing it, please contact [email protected].

Created June 18, 2020, Updated July 15, 2020

Was this page helpful?

A crowdsourced set of curated structural variants for the human genome

Author(s)

Abstract

Download Paper

Keywords

Citation

Additional citation formats

Issues