EP4036247A1

Systems and methods to detect rare mutations and copy number variation

Abstract

The present disclosure provides a system and method for the detection of rare mutations and copy number variations in cell free polynucleotides. Generally, the systems and methods comprise sample preparation, or the extraction and isolation of cell free polynucleotide sequences from a bodily fluid; subsequent sequencing of cell free polynucleotides by techniques known in the art; and application of bioinformatics tools to detect rare mutations and copy number variations as compared to a reference. The systems and methods also may contain a database or collection of different rare mutations or copy number variation profiles of different diseases, to be used as additional references in aiding detection of rare mutations, copy number variation profiling or general genetic profiling of a disease.

EP4036247A1, drawing sheet 1
Sheet 1 of 16

Term

6.9 yearsto projected expiry

Projected expiry 4 September 2033, counted from filing; an application has no term until it is granted.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

15 claims: 8 independent, 7 dependent

  1. 1
    A method comprising:(a) providing at least one set of tagged parent polynucleotides by converting initial starting genetic material into the tagged parent polynucleotides, wherein the initial starting genetic material is cell-free nucleic acid, wherein converting comprises any of blunt-end ligation, sticky end ligation and single strand ligation, and wherein each tagged parent polynucleotide in the set is uniquely tagged;and for each set of tagged parent polynucleotides: (b) amplifying the tagged parent polynucleotides in the set to produce a corresponding set of amplified progeny polynucleotides;(c) sequencing a subset of the set of amplified progeny polynucleotides, to produce a set of sequencing reads;and (d) collapsing the set of sequencing reads to generate a set of consensus sequences, each consensus sequence corresponding to a unique polynucleotide among the set of tagged parent polynucleotides;the method comprising enriching the set of amplified progeny polynucleotides for polynucleotides mapping to one or more selected mappable positions in a reference sequence by: (i) selective amplification of sequences from initial starting genetic material converted to tagged parent polynucleotides;(ii) selective amplification of tagged parent polynucleotides;(iii) selective sequence capture of amplified progeny polynucleotides;or (iv) selective sequence capture of initial starting genetic material.
  2. 5
    The method of any preceding claim, wherein the initial starting genetic material comprises no more than 100 ng of polynucleotides.
  3. 6
    The method of any preceding claim, comprising converting the initial starting genetic material into tagged parent polynucleotides with a conversion efficiency of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 80% or at least 90%.
  4. 9
    The method of any preceding claim, wherein the subset of the set of amplified progeny polynucleotides sequenced is of sufficient size so that any nucleotide sequence represented in the set of tagged parent polynucleotides at a percentage that is the same as the percentage per-base sequencing error rate of the sequencing platform used, has at least a 50%, at least a 60%, at least a 70%, at least a 80%, at least a 90% at least a 95%, at least a 98%, at least a 99%, at least a 99.9% or at least a 99.99% chance of being represented among the set of consensus sequences.
  5. 10
    The method of any preceding claim, comprising attaching one or more barcodes to the cell-free nucleic acids prior to any amplification or enrichment step, for example wherein the barcodes comprise oligonucleotides of at least 5, 10, 15, 20 25, 30, 35, 40, 45, or 50mer base pairs in length and/or wherein the barcode comprises random sequence.
  6. 13
    The method of any preceding claim, wherein the polynucleotides are extracted from a sample selected from the group consisting of blood, plasma, serum, urine, saliva, mucosal excretions, sputum, stool, and tears.
  7. 14
    The method of any preceding claim, comprising filtering out reads with an accuracy or quality score of less than a threshold.
  8. 15
    The method of any preceding claim, wherein the number of unique identifiers is:i) at least 3 and at most 100;ii) at least 5 and at most 100;iii) at least 10 and at most 100;iv) at least 15 and at most 100;v) at least 25 and at most 100;vi) at least 3 and at most 1000;vii) at least 5 and at most 1000;viii) at least 10 and at most 1000;ix) at least 15 and at most 1000;x) at least 25 and at most 1000;xi) at least 3 and at most 10,000;xii) at least 5 and at most 10,000;xiii) at least 10 and at most 10,000;xiv) at least 15 and at most 10,000;or xv) at least 25 and at most 10,000.