WO2015073711A1

Compositions and methods for identification of a duplicate sequencing read

Abstract

The present invention provides methods, compositions and kits for detecting duplicate sequencing reads. In some embodiments, the duplicate sequencing reads are removed. The present invention is based, in part, on compositions and methods for discerning duplicate sequencing reads from a population of sequencing reads. The detection and/or removal of duplicate sequencing reads presented herein is a novel approach to increasing the efficacy of evaluating data generated from high throughput sequence reactions, including complex multiplex sequence reactions.

WO2015073711A1, drawing sheet 1
Sheet 1 of 6

Term

No projected expiry on record.

  1. Priority
  2. Filed
  3. Published
  4. Today

1 claim: 1 independent, 0 dependent

  1. 1
    CLAIMS A method for detecting a duplicate sequencing read from a population of sample sequencing reads comprising:a) ligating an adaptor to a 5' end of each nucleic acid fragment of a plurality of nucleic acid fragments from one or more samples, wherein the adaptor comprises: (i) an indexing primer binding site;(ii) an indexing site;(iii) an identifier site;and (iv) a target sequence primer binding site;b) amplifying the adapter-nucleic acid fragment ligated products;c) generating a population of sequencing reads from amplified adapter-nucleic acid fragment ligated products;and d) detecting the population of sequencing reads comprising a sequencing read with a duplicate identifier site and target sequence. The method of claim 1 , wherein the method further comprises removing from the population of sequence reads the sequencing read with a duplicate identifier site and target sequence. The method of claim 1, wherein the identifier site is sequenced with the indexing site. The method of claim 1, wherein the identifier site is sequenced separately from the indexing site. The method of claim 1, wherein the identifier site is sequenced with the target sequence. The method of claim 1, wherein the identifier site is sequenced separately from the target sequence. The method of claim 1 , wherein the adaptor comprises from 5 ' to 3 ' : (i) the indexing primer binding site;(ii) the indexing site;(iii) the identifier site;and (iv) the target sequence primer binding site. 8. The method of claim 1, wherein the adaptor comprises from 5' to 3': (i) the indexing primer binding site;(ii) the indexing site;(iii) the target sequence primer binding site;and (iv) the identifier site. 9. The method of claim 1 , wherein the plurality of nucleic acid fragments is generated from more than one sample. 10. The method of claim 9, wherein the nucleic acid fragments from each sample has the same indexing site. 11. The method of claim 10, wherein the sequencing reads are separated based on the indexing site. 12. The method of claim 11 , wherein the separation of sequencing reads is performed prior to step d). 13. The method of claim 1 , wherein the nucleic acid fragments are DNA fragments, RNA fragments, or DNA RNA fragments. 14. The method of claim 13, wherein the nucleic acid fragments are genomic DNA fragments or cDNA fragments. 15. The method of claim 1, wherein the indexing site is between 2 and 8 nucleotides in length. 16. The method of claim 1 , wherein the indexing site is about 6 nucleotides in length. 17. The method of claim 1, wherein the identifier site is between 1 and 8 nucleotides in length. 18. The method of claim 1, wherein the identifier site is about 8 nucleotides in length. 19. The method of claim 1 , wherein the indexing primer binding site is a universal indexing primer binding site. 20. The method of claim 1 , wherein the target sequence primer binding site is a universal target sequence primer binding site. 21. A kit comprising a plurality of adaptors, wherein each adaptor comprises: (i) an indexing primer binding site;(ϋ) an indexing site;(iii) an identifier site;and (iv) a target sequencing primer binding site.