EP3087204A1

Methods and systems for detecting genetic variants

Abstract

This record has no abstract on file.

EP3087204A1, drawing sheet 1
Sheet 1 of 12

Term

Projected expiry 24 December 2034.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

92 claims: 16 independent, 76 dependent

  1. 1
    Claims of equivalent WO 2015100427 A1 CLAIMS WHAT IS CLAIMED IS:1. A method for determining a quantitative measure indicative of a number of individual double-stranded deoxyribonucleic acid (DNA) molecules in a sample, comprising: (a) determining a quantitative measure of individual DNA molecules for which both strands are detected;(b) determining a quantitative measure of individual DNA molecules for which only one of the DNA strands is detected;(c) inferring from (a) and (b) above a quantitative measure of individual DNA molecules for which neither strand was detected;and (d) using (a)-(c) to determine the quantitative measure indicative of a number of individual double-stranded DNA molecules in the sample.
  2. 2
    The method of Claim 1 , further comprising detecting copy number variation in said sample by determining a normalized quantitative measure determined in step (d) at each of one or more genetic loci and determining copy number variation based on the normalized measure.
  3. 3
    The method of Claim 1, wherein said sample comprises double-stranded polynucleotide molecules sourced substantially from cell-free nucleic acids.
  4. 4
    The method of Claim 1, wherein determining said quantitative measure of individual DNA molecules comprises tagging said DNA molecules with a set of duplex tags, wherein each duplex tag differently tags complementary strands of a double-stranded DNA molecule in said sample to provide tagged strands.
  5. 5
    The method of Claim 4, further comprising sequencing at least some of said tagged strands to produce a set of sequence reads.
  6. 6
    The method of Claim 5, further comprising sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in said set, and (ii) each unpaired read represents a first tagged strand having no second differently tag complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads.
  7. 7
    The method of Claim 6, further comprising determining quantitative measures of (i) said paired reads and (ii) said unpaired reads that map to each of one or more genetic loci to determine a quantitative measure of total double-stranded DNA molecules in said sample that map to each of said one or more genetic loci based on said quantitative measure of paired reads and unpaired reads mapping to each locus.
  8. 8
    A method for reducing distortion in a sequencing assay, comprising:(a) tagging control parent polynucleotides with a first tag set to produce tagged control parent polynucleotides;(b) tagging test parent polynucleotides with a second tag set to produce tagged test parent polynucleotides;(c) mixing tagged control parent polynucleotides with tagged test parent polynucleotides to form a pool;(d) determining quantities of tagged control parent polynucleotides and tagged test parent polynucleotides;and (e) using the quantities of tagged control parent polynucleotides to reduce distortion in the quantities of tagged test parent polynucleotides.
  9. 9
    The method of Claim 8, wherein said first tag set comprises a plurality of tags, wherein each tag in said first tag set comprises a same control tag and an identifying tag, and wherein said first tag set comprises a plurality of different identifying tags.
  10. 10
    The method of Claim 9, wherein said second tag set comprises a plurality of tags, wherein each tag in said second tag set comprises a same test tag and an identifying tag, wherein said test tag is distinguishable from said control tag, and wherein said second tag set comprises a plurality of different identifying tags.
  11. 11
    The method of Claim 9, wherein (d) comprises amplifying tagged parent polynucleotides in said pool to form a pool of amplified, tagged polynucleotides, and sequencing amplified, tagged polynucleotides in said amplified pool to produce a plurality of sequence reads.
  12. 12
    The method of Claim 11, further comprising grouping sequence reads into families, each family comprising sequence reads generated from a same parent polynucleotide, which grouping is optionally based on information from an identifying tag and from start/end sequences of said parent polynucleotides, and, optionally, determining a consensus sequence for each of a plurality of parent polynucleotides from said plurality of sequence reads in a group.
  13. 13
    The method of Claim 8, wherein (d) comprises determining copy number variation in said test parent polynucleotides at greater than or equal to one locus based on relative quantity of test parent polynucleotides and control parent polynucleotides mapping to said locus.
  14. 14
    A set of library adaptors comprising a plurality of polynucleotide molecules with molecular barcodes, wherein said plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length, wherein said molecular barcodes are at least 4 nucleotide bases in length, and wherein:(a) said molecular barcodes are different from one another and have an edit distance of at least 1 between one another;(b) said molecular barcodes are located at least one nucleotide base away from a terminal end of their respective polynucleotide molecules;(c) optionally, at least one terminal base is identical in all of said polynucleotide molecules;and (d) none of said polynucleotide molecules contains a complete sequencer motif.
  15. 15
    The set of library adaptors of Claim 14, wherein said polynucleotide molecules are identical but for said molecular barcodes.
  16. 16
    The set of library adaptors of Claim 14, wherein each of said plurality of polynucleotide molecules has a double stranded portion and at least one single-stranded portion.
  17. 17
    The set of library adaptors of Claim 16, wherein said double-stranded portion has a molecular barcode among said molecular barcodes.
  18. 18
    The set of library adaptors of Claim 17, wherein said given molecular barcode is a randomer.
  19. 19
    The set of library adaptors of Claim 16, wherein each of said plurality of polynucleotide molecules further comprises a strand-identification barcode on said at least one single-stranded portion.
  20. 20
    The set of library adaptors of Claim 19, wherein said strand-identification barcode includes at least 4 nucleotide bases.
  21. 21
    The set of library adaptors of Claim 16, wherein said single-stranded portion has a partial sequencer motif.
  22. 22
    The set of library adaptors of Claim 14, wherein said polynucleotide molecules have a sequence of terminal nucleotides that are the same.
  23. 23
    The set of library adaptors of Claim 14, wherein each of said plurality of polynucleotide molecules is Y-shaped, bubble shaped or hairpin shaped.
  24. 24
    The set of library adaptors of Claim 14, wherein none of said polynucleotide molecules contains a sample identification motif.
  25. 25
    The set of library adaptors of Claim 14, wherein said molecular barcodes are at least 10 nucleotide bases in length.
  26. 26
    The set of library adaptors of Claim 14, wherein each of said plurality of polynucleotide molecules is from 10 nucleotide bases to 60 nucleotide bases in length.
  27. 27
    The set of library adaptors of Claim 14, where said at least one terminal base is identical in all of said polynucleotide molecules.
  28. 28
    The set of library adaptors of Claim 14, wherein said molecular barcodes are located at least 10 nucleotide base away from a terminal end of their respective polynucleotide molecules.
  29. 29
    The set of library adaptors of Claim 14, consisting essentially of said plurality of polynucleotide molecules.
  30. 30
    A method, comprising:(a) tagging a collection of polynucleotides with a plurality of polynucleotide molecules from a library of adaptors as in Claim 14 to create a collection of tagged polynucleotides;and (b) amplifying said collection of tagged polynucleotides in the presence of sequencing adaptors, wherein said sequencing adaptors have primers with nucleotide sequences that are selectively hybridizable to complementary sequences in said plurality of polynucleotide molecules.
  31. 31
    A method for detecting or quantifying rare deoxyribonucleic acid (DNA) in a heterogeneous population of original DNA fragments, wherein said rare DNA has a concentration that is less than 1%, the method comprising:(a) tagging said original DNA fragments in a single reaction such that greater than 30% of said original DNA fragments are tagged at both ends with library adaptors that comprise molecular barcodes, thereby providing tagged DNA fragments;(b) performing high-fidelity amplification on said tagged DNA fragments;(c) optionally, selectively enriching a subset of said tagged DNA fragments;(d) sequencing one or both strands of said tagged, amplified and optionally selectively enriched DNA fragments to obtain sequence reads comprising nucleotide sequences of said molecular barcodes and at least a portion of said original DNA fragments;(e) from said sequence reads, determining consensus reads that are representative of single-strands of said original DNA fragments;and (f) quantifying said consensus reads to detect or quantify said rare DNA at a specificity that is greater than 99.9%.
  32. 32
    The method of Claim 31 , wherein step (e) comprises comparing sequence reads having the same or similar molecular barcodes and the same or similar end of fragment sequences.
  33. 33
    The method of Claim 32, wherein said comparing further comprises performing a phylogentic analysis on said sequence reads having the same or similar molecular barcodes.
  34. 34
    The method of Claim 32, wherein said molecular barcodes include a barcode having an edit distance of up to 3.
  35. 35
    The method of Claim 31 , wherein said end of fragment sequence includes fragment sequences having an edit distance of up to 3.
  36. 36
    The method of Claim 31 , further comprising sorting sequence reads into paired reads and unpaired reads, and quantifying a number of paired reads and unpaired reads that map to each of one or more genetic loci.
  37. 37
    The method of Claim 31 , wherein said tagging occurs by having an excess amount of library adaptors as compared to original DNA fragments.
  38. 38
    The method of Claim 31 , further comprising binning said sequence reads according to said molecular barcodes and sequence information from at least one end of each of said original DNA fragments to create bins of single stranded reads.
  39. 39
    The method of Claim 38, further comprising, in each bin, determining a sequence of a given original DNA fragment among said original DNA fragments by analyzing sequence reads.
  40. 40
    The method of Claim 39, further comprising detecting or quantifying said rare DNA by comparing a number of times each base occurs at each position of a genome represented by said tagged, amplified, and optionally enriched DNA fragments.
  41. 41
    The method of Claim 31 , further comprising selectively enriching a subset of said tagged DNA fragments.
  42. 42
    The method of Claim 41, further comprising, after enriching, amplifying the enriched tagged DNA fragments in the presence of sequencing adaptors comprising primers.
  43. 43
    The method of Claim 31 , wherein said DNA fragments are tagged with polynucleotide molecules from a library of adaptors as in Claim 1.
  44. 44
    A method for processing and/or analyzing a nucleic acid sample of a subject, comprising:(a) exposing polynucleotide fragments from said nucleic acid sample to a set of library adaptors to generate tagged polynucleotide fragments;and (b) subjecting said tagged polynucleotide fragments to nucleic acid amplification reactions under conditions that yield amplified polynucleotide fragments as amplification products of said tagged polynucleotide fragments, wherein said set of library adaptors comprises a plurality of polynucleotide molecules with molecular barcodes, wherein said plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length, wherein said molecular barcodes are at least 4 nucleotide bases in length, and wherein (1) said molecular barcodes are different from one another and have an edit distance of at least 1 between one another;(2) said molecular barcodes are located at least one nucleotide base away from a terminal end of their respective polynucleotide molecules;(3) optionally, at least one terminal base is identical in all of said polynucleotide molecules;and (4) none of said polynucleotide molecules contains a complete sequencer motif.
  45. 45
    The method of Claim 44, further comprising determining nucleotide sequences of said amplified tagged polynucleotide fragments.
  46. 46
    The method of Claim 45, wherein said nucleotide sequences of said amplified tagged polynucleotide fragments are determined without polymerase chain reaction (PCR).
  47. 47
    The method of Claim 45, further comprising analyzing said nucleotide sequences with a programmed computer processor to identify one or more genetic variants in said nucleotide sample of the subject.
  48. 48
    The method of Claim 44, wherein said nucleic acid sample is a cell-free nucleic acid sample.
  49. 49
    The method of Claim 44, wherein exposing said polynucleotide fragments of said nucleic acid sample to said plurality of polynucleotide molecules yields said tagged polynucleotide fragments with a conversion efficiency of at least 10%.
  50. 50
    The method of Claim 44, wherein said subjecting comprises amplifying said tagged polynucleotide fragments from sequences corresponding to genes selected from the group consisting of ALK, APC, BRAF, CDK 2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CT NB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPNl l, RET,SMAD4, SMARCBl, SMO, SRC, STKl l, VHL, TERT, CCNDl, CDK4, CDKN2B, RAFl, BRCAl, CCND2, CDK6, NFl, TP53, ARIDIA, BRCA2, CCNE1, ESR1, RIT1, GAT A3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1.
  51. 51
    A method, comprising :(a) generating a plurality of sequence reads from a plurality of polynucleotide molecules, wherein said plurality of polynucleotide molecules cover genomic loci of a target genome, wherein said genomic loci correspond to a plurality of genes selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNFIA, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPNl l, RET,SMAD4, SMARCBl, SMO, SRC, STKl l, VHL, TERT, CCNDl, CDK4, CDKN2B, RAFl , BRCAl, CCND2, CDK6, NFl, TP53, ARIDIA, BRCA2, CCNE1, ESR1, RIT1, GAT A3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRKl;(b) grouping with a computer processor said plurality of sequence reads into families, wherein each family comprises sequence reads from one of said template polynucleotides;(c) for each of said families, merging sequence reads to generate a consensus sequence;(d) calling said consensus sequence at a given genomic locus among said genomic loci;(e) detecting at said given genomic locus any of: i. genetic variants among the calls;ii. frequency of a genetic alteration among the calls;iii. total number of calls;and iv. total number of alterations among the calls.
  52. 52
    The method of Claim 51 , wherein each family comprises sequence reads from only one of said template polynucleotides.
  53. 53
    The method of Claim 51 , further comprising performing (d)-(e) at an additional genomic locus among said genomic loci.
  54. 54
    The method of Claim 53, further comprising determining a variation in copy number at one of said given genomic locus and additional genomic locus based on counts at said given genomic locus and additional genomic locus.
  55. 55
    The method of Claim 51 , wherein said grouping comprises classifying said plurality of sequence reads into families by identifying (i) distinct molecular barcodes coupled to said plurality of polynucleotide molecules and (ii) similarities between said plurality of sequence reads, wherein each family includes a plurality of nucleic acid sequences that are associated with a distinct combination of molecular barcodes and similar or identical sequence reads.
  56. 56
    The method of Claim 51 , wherein said consensus sequence is generated by evaluating a quantitative measure or a statistical significance level for each of said sequence reads.
  57. 57
    The system of Claim 51 , wherein said plurality of genes includes at least 10 of said plurality of genes selected from said group.
  58. 58
    A method, comprising :(a) providing template polynucleotide molecules and a set of library adaptors in a single reaction vessel, wherein said library adaptors are polynucleotide molecules that have different molecular barcodes, and wherein none of said library adaptors contains a complete sequencer motif;(b) in said single reaction vessel, coupling said library adaptors to said template polynucleotide molecules at an efficiency of at least 10%, thereby tagging each template polynucleotide with a tagging combination that is among a plurality of different tagging combinations, to produce tagged polynucleotide molecules;(c) subjecting said tagged polynucleotide molecules to an amplification reaction under conditions that yield amplified polynucleotide molecules as amplification products of said tagged polynucleotide molecules;and (d) sequencing said amplified polynucleotide molecules.
  59. 59
    The method of Claim 58, wherein said library adaptors are identical but for said molecular barcodes.
  60. 60
    The method of Claim 58, wherein each of said library adaptors has a double stranded portion and at least one single-stranded portion, and wherein said single-stranded portion has a partial sequencer motif.
  61. 61
    The method of Claim 58, wherein said library adaptors couple to both ends of said template polynucleotide molecules.
  62. 62
    The method of Claim 58, wherein said efficiency is at least 30%.
  63. 63
    The method of Claim 58, further comprising identifying genetic variants upon sequencing said amplified polynucleotide molecules.
  64. 64
    The method of Claim 58, wherein said sequencing comprises (i) subjecting said amplified polynucleotide molecules to an additional amplification reaction under conditions that yield additional amplified polynucleotide molecules as amplification products of said amplified polynucleotide molecules, and (ii) sequencing said additional amplified polynucleotide molecules.
  65. 65
    The method of Claim 64, wherein said additional amplification is performed in the presence of sequencing adaptors.
  66. 66
    The method of Claim 58, wherein (b) and (c) are performed without aliquoting said tagged polynucleotide molecules.
  67. 67
    A system for analyzing a target nucleic acid molecule of a subject, comprising:a communication interface that receives nucleic acid sequence reads for a plurality of polynucleotide molecules that cover genomic loci of a target genome;computer memory that stores said nucleic acid sequence reads for said plurality of polynucleotide molecules received by said communication interface;and a computer processor operatively coupled to said communication interface and said memory and programmed to (i) group said plurality of sequence reads into families, wherein each family comprises sequence reads from one of said template polynucleotides, (ii) for each of said families, merge sequence reads to generate a consensus sequence, (iii) call said consensus sequence at a given genomic locus among said genomic loci, and (iv) detect at said given genomic locus any of genetic variants among the calls, frequency of a genetic alteration among the calls, total number of calls;and total number of alterations among the calls, wherein said genomic loci correspond to a plurality of genes selected from the group consisting of ALK, APC, BRAF, CDK 2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CT NB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET,SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID 1 A, BRCA2, CCNE1, ESR1, RIT1, GAT A3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1.
  68. 68
    A set of oligonucleotide molecules that selectively hybridize to at least 5 genes selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RBI, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET,SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID 1 A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1.
  69. 69
    The set of Claim 68, wherein said oligonucleotide molecules are from 10-200 bases in length.
  70. 70
    The kit of Claim 68, wherein said oligonucleotide molecules selectively hybridize to exon regions of said at least 5 genes.
  71. 71
    The kit of Claim 70, wherein said oligonucleotide molecules selectively hybridize to at least 30 exons in said at least 5 genes.
  72. 72
    The kit of Claim 71, wherein multiple oligonucleotide molecules selectively hybridize to each of said at least 30 exons.
  73. 73
    The kit of Claim 72, wherein said oligonucleotide molecules that hybridize to each exon have sequences that overlaps with at least 1 other oligonucleotide molecule.
  74. 74
    A kit, comprising:a first container containing a plurality of library adaptors each having a different molecular barcode;and a second container containing a plurality of sequencing adaptors, each sequencing adaptor comprising at least a portion of a sequencer motif and optionally a sample barcode.
  75. 75
    The kit of Claim 74, wherein said sequencing adaptor comprises said sample barcode.
  76. 76
    A method for detecting sequence variants in a cell free DNA sample, comprising:detecting rare DNA at a concentration less than 1% with a specificity that is greater than 99.9%.
  77. 77
    A method, comprising:(a) providing a sample comprising a set of double-stranded polynucleotide molecules, each double-stranded polynucleotide molecule including first and second complementary strands;(b) tagging said double-stranded polynucleotide molecules with a set of duplex tags, wherein each duplex tag differently tags said first and second complementary strands of a double-stranded polynucleotide molecule in said set;(c) sequencing at least some of said tagged strands to produce a set of sequence reads;(d) reducing and/or tracking redundancy in said set of sequence reads;(e) sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in said set, and (ii) each unpaired read represents a first tagged strand having no second differently tag complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads;(f) determining quantitative measures of (i) said paired reads and (ii) said unpaired reads that map to each of one or more genetic loci;and (g) estimating with a programmed computer processor a quantitative measure of total double-stranded polynucleotide molecules in said set that map to each of said one or more genetic loci based on said quantitative measure of paired reads and unpaired reads mapping to each locus.
  78. 78
    The method of Claim 77, further comprising (h) detecting copy number variation in said sample by determining a normalized total quantitative measure determined in step (g) at each of said one or more genetic loci and determining copy number variation based on the normalized measure.
  79. 79
    The method of Claim 77, wherein said sample comprises double-stranded polynucleotide molecules sourced substantially from cell-free nucleic acids.
  80. 80
    The method of Claim 77, wherein said duplex tags are not sequencing adaptors.
  81. 81
    The method of Claim 77, wherein reducing redundancy in said set of sequence reads comprises collapsing sequence reads produced from amplified products of an original polynucleotide molecule in said sample back to said original polynucleotide molecule.
  82. 82
    The method of Claim 81, further comprising determining a consensus sequence for said original polynucleotide molecule.
  83. 83
    The method of Claim 82, further comprising identifying polynucleotide molecules at one or more genetic loci comprising a sequence variant.
  84. 84
    The method of Claim 82, further comprising determining a quantitative measure of paired reads that map to a locus, wherein both strands of said pair comprise a sequence variant.
  85. 85
    The method of Claim 84, further comprising determining a quantitative measure of paired molecules in which only one member of said pair bears a sequence variant and/or determining a quantitative measure of unpaired molecules bearing a sequence variant.
  86. 86
    A method, comprising:(a) from a sequencer, receiving into memory a set of sequence reads of polynucleotides tagged with duplex tags;(b) reducing and/or tracking redundancy in said set of sequence reads;(c) sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in said set, and (ii) each unpaired read represents a first tagged strand having no second differently tag complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads;(d) determining quantitative measures of (i) said paired reads and (ii) said unpaired reads that map to each of one or more genetic loci;and (e) estimating a quantitative measure of total double-stranded polynucleotide molecules in said set that map to each of said one or more genetic loci based on said quantitative measure of paired reads and unpaired reads mapping to each locus.
  87. 87
    A method, comprising:(a) providing a sample comprising a set of double-stranded polynucleotide molecules, each double-stranded polynucleotide molecule including first and second complementary strands;(b) tagging said double-stranded polynucleotide molecules with a set of duplex tags, wherein each duplex tag differently tags said first and second complementary strands of a double-stranded polynucleotide molecule in said set;(c) sequencing at least some of said tagged strands to produce a set of sequence reads;(d) reducing and/or tracking redundancy in said set of sequence reads;(e) sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in said set, and (ii) each unpaired read represents a first tagged strand having no second differently tag complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads;and (f) determining quantitative measures of at least two of (i) said paired reads, (ii) said unpaired reads that map to each of one or more genetic loci, (iii) read depth of said paired reads and (iv) read depth of unpaired reads. A method, comprising: (a) tagging control parent polynucleotides with a first tag set to produce tagged control parent polynucleotides, wherein said first tag set comprises a plurality of tags, wherein each tag in said first tag set comprises a same control tag and an identifying tag, and wherein said tag set comprises a plurality of different identifying tags;(b) tagging test parent polynucleotides with a second tag set to produce tagged test parent polynucleotides, wherein said second tag set comprises a plurality of tags, wherein each tag in said second tag set comprises a same test tag that is distinguishable from said control tag and an identifying tag, and wherein said second tag set comprises a plurality of different identifying tags;(c) mixing tagged control parent polynucleotides with tagged test parent polynucleotides to form a pool;(d) amplifying tagged parent polynucleotides in said pool to form a pool of amplified, tagged polynucleotides;(e) sequencing amplified, tagged polynucleotides in said amplified pool to produce a plurality of sequence reads;(f) grouping sequence reads into families, each family comprising sequence reads generated from a same parent polynucleotide, which grouping is optionally based on information from an identifying tag and from start/end sequences of said parent polynucleotides, and, optionally, determining a consensus sequence for each of a plurality of parent polynucleotides from said plurality of sequence reads in a group;(g) classifying each family or consensus sequence as a control parent polynucleotide or as a test parent polynucleotide based on having a test tag or a control tag;(h) determining a quantitative measure of control parent polynucleotides and control test polynucleotides mapping to each of at least two genetic loci;and (i) determining copy number variation in said test parent polynucleotides at at least one locus based on relative quantity of test parent polynucleotides and control parent polynucleotides mapping to said at least one locus.
  88. 88
    89. A method comprising:(a) generating a plurality of sequence reads from a plurality of template polynucleotides, each polynucleotide mapped to a genomic locus;(b) grouping the sequence reads into families, each family comprising sequence reads generated from one of the template polynucleotides;(c) calling a nucleotide base or sequence at the genomic locus for each of the families;(d) detecting at the genomic locus any of: i. genomic alterations among the calls;ii. frequency of a genetic alteration among the calls;iii. total number of calls;iv. total number of alterations among the calls.
  89. 89
    90. The method of Claim 89, wherein calling comprises any of phylogenetic analysis, voting, weighing, assigning a probability to each read at the locus in a family, and calling the nucleotide base with the highest probability.
  90. 90
    91. The method of Claim 89, performed at two loci, comprising determining CNV at one of the loci based on counts at each of the loci.
  91. 91
    92. A method comprising:(a) ligating adaptors to double-stranded deoxyribonucleic acid (DNA) polynucleotides, wherein ligating is performed in a single reaction vessel, and wherein the adaptors comprise molecular barcodes, to produce a tagged library comprising an insert from the double-stranded DNA polynucleotides, and having between 4 and 1 million different tags;(b) generating a plurality of sequence reads for each of said double-stranded DNA polynucleotides in the tagged library;(c) grouping sequence reads into families, each family comprising sequence reads generated from a single DNA polynucleotide among said double-stranded DNA polynucleotides, based on information in a tag and information at an end of the insert;and (d) calling nucleotide bases at each position in the double-stranded DNA molecule based on nucleotide bases at the position in members of a family.
  92. 92
    93. The method of Claim 93, wherein (d) comprises calling a plurality of sequential bases from at least a subset of said sequence reads to identify single nucleotide variations (SNV) in the double-stranded DNA molecule.
Independent claims92