Nova Patents
CA2869574C

Sequence assembly

Abstract

The invention relates to assembly of sequence reads. The invention provides a method for identifying a mutation in a nucleic acid involving sequencing nucleic acid to generate a plurality of sequence reads. Reads are assembled to form a contig, which is aligned to a reference. Individual reads are aligned to the contig. Mutations are identified based on the alignments to the reference and to the contig.

CA2869574C, drawing sheet 1
Sheet 1 of 3

Term

6.5 yearsleft in the term

Expires 22 March 2033.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

58 claims: 4 independent, 54 dependent

  1. 1
    CLAIMS:1. A method for identifying a mutation in a nucleic acid, the method comprising: sequencing nucleic acid with a sequencer to generate a plurality of sequence reads, the sequencer operably coupled to a computer comprising memory and a processor configured to: create a contig based on the plurality of sequence reads;align the contig to a reference sequence to provide a contig-to-reference alignment;align each read back to the contig to provide read-to-contig alignments;map the read-to-contig alignments to the reference sequence using the contig-toreference alignment;combine differences between the plurality of sequence reads and the contig and differences between the contig and the reference sequence for each read;identify a mutation for each read based on the differences between the reference sequence and the plurality of sequence reads;and output mutation identification results as a report.
  2. 13
    A method for assembling and aligning a plurality of sequence reads having mutations of different types, the method comprising:obtaining a sample comprising a template nucleic acid;sequencing the sample to generate the plurality of sequence reads, the sequencing comprising: fragmenting the template nucleic acid, attaching the fragments to a surface of channels in a flow cell, and amplifying the attached fragments to create clusters, each cluster comprising a plurality of copies of the template nucleic acid in one of the channels in the flow cell;Date Reçue/Date Received 2021-09-29 81783049 inputting a reference genome and the plurality of sequence reads into a computer system comprising a non-transitory memory and a processor coupled to the non-transitory memory, wherein the non-transitory memory has instructions stored thereon that, when executed by the processor, cause the processor to perform the steps of: assembling a contig from at least some of the plurality of sequence reads;identifying a plurality of contig-to-reference descriptions of the mutations by aligning the contig to a sequence of the reference genome, the mutations including a substitution and an indel;identifying a plurality of read-to-contig descriptions by aligning each of the at least some of the plurality of sequence reads to the contig;and generating a read-to-reference description by aligning at least one of the plurality of contig-to-reference descriptions with a corresponding at least one of the plurality of read-tocontig descriptions, wherein the read-to-reference description maps positional information of the mutations found in at least one of the at least some of the plurality of sequence reads relative to the sequence of the reference genome.
  3. 31
    A method for accurately identifying differences between a reference human genome and sequence reads obtained from a biological sample, the biological sample obtained from a human subject, the method comprising:obtaining nucleic acid from the biological sample obtained from the human subject sequencing, by next generation sequencing, the nucleic acid to generate the sequence reads;and genotyping at least some of the sequence reads using a multi-stage alignment, the genotyping performed by one or more computer software programs executing on at least one computer processor coupled to a computer-readable memory, the genotyping comprising: assembling a contig using the at least some of the sequence reads, the contig including information about positions of the at least some of the sequence reads relative to each other or to a reference;aligning, using a first substitution probability and a first gap penalty, the contig to the reference human genome to obtain a reference alignment and storing the reference alignment in the computer-readable memory, the reference alignment indicative of first differences between the contig and the reference human genome;aligning, using a second substitution probability and a second gap penalty, the at least some of the sequence reads to the contig to obtain sequence read alignments and storing the sequence read alignments in the computer-readable memory, the sequence read alignments indicative of second differences between the at least some of the sequence reads and the contig;and genotyping the at least some of the sequence reads by identifying multiple mutations in the at least some of the sequence reads based on the first differences and the second differences, the genotyping comprising mapping the at least some of the sequence reads to the reference genome by combining the reference alignment and the sequence read alignments to determine an identity of each of the multiple mutations and its location in the human Date Reçue/Date Received 2021-09-29 81783049 reference genome, the multiple mutations including mutations of different types including a first substitution and a first indel.
  4. 50
    A method for accurately identifying differences between a human reference genome and sequence reads obtained from a biological sample, each of the sequence reads being associated with a respective barcode in a plurality of barcodes, the biological sample obtained from a human subject, the method comprising:obtaining nucleic acid from the biological sample obtained from the human subject;sequencing, by next generation sequencing, the nucleic acid to generate the sequence reads;and genotyping at least some of the sequence reads using a multi-stage alignment, the genotyping performed by using one or more computer software programs executing on at least one computer processor coupled to a computer-readable memory, the genotyping comprising: assembling the at least some of the sequence reads into multiple contigs by: grouping the at least some of the sequence reads into subsets based on associated barcodes;generating a respective contig for each of the subsets, the respective contig including information about the position of at least some of the sequence reads in the subset relative to each other;aligning, using a first substitution probability and a first gap penalty, the contigs to the reference human genome to obtain reference alignments and storing the reference alignments in the computer-readable memory, the reference alignments indicative of first differences between the multiple contigs and the reference human genome;aligning, using a second substitution probability and a second gap penalty, sequence reads in each of the subsets to the respective contig to obtain sequence read alignments and Date Reçue/Date Received 2021-09-29 81783049 storing the sequence read alignments in the computer-readable memory, the sequence read alignments indicative of second differences between the sequence reads in each of the subsets and the respective contigs in the multiple contigs;and genotyping the at least some of the sequence reads by identifying multiple mutations in the at least some of the sequence reads based on the fust differences and the second differences, the genotyping comprising mapping the at least some of the sequence reads to the reference genome by combining the reference alignments and the sequence read alignments to determine an identity of each of the multiple mutations and its location in the reference human genome, the multiple mutations including mutations of different types including a first substitution and a first indel.