IL281741A

Methylation markers and targeted methylation probe panel

Abstract

This record has no abstract on file.

IL281741A, drawing sheet 1
Sheet 1 of 29

Term

No projected expiry on record.

  1. Priority
  2. Filed
  3. Published
  4. Today

255 claims: 151 independent, 104 dependent

  1. 1
    A bait set for hybridization capture, the bait set comprising a plurality of different oligonucleotide-containing probes, wherein each of the oligonucleotide-containing probes comprises a sequence of at least 30 bases in length that is complementary to either:(1 ) a sequence of a genomic region;or (2 ) a sequence that varies from the sequence of (1) only by one or more transitions, wherein each respective transition of the one or more transitions occurs at a cytosine in the genomic region, and wherein each probe of the different oligonucleotide-containing probes is complementary to a sequence corresponding to a CpG site that is differentially methylated in cancer samples relative to non-cancer samples.
  2. 4
    The bait set of any one of claims 1-3, wherein the CpG site is considered to be differentially methylated in cancer samples relative to non-cancer samples based on criteria comprising Ncancer and Nnon-cancer, wherein:Ncancer is a number of cancer samples that include a cfDNA fragment covering the CpG site that (1) has at least X CpG sites, wherein at least Y% of the CpG sites are methylated or unmethylated, wherein X is at least 4 and Y is at least 70, and (2) has a p-value rarity in non-cancerous samples of below a threshold value;and Nnon-cancer is a number of cancer samples that include a cfDNA fragment covering the CpG site that (1) has at least M CpG sites, wherein at least N% of the sites are methylated or unmethylated, wherein M is at least 4 and N is at least 70, and (2) has a p-value rarity in noncancerous samples of below a threshold value.
  3. 7
    8. The bait set of any one of claims 1-7, wherein, for each of the different oligonucleotide-containing probes, the sequence of at least 30 bases in length is complementary to either (1) a sequence within a genomic region selected from the genomic regions set forth in any one of Lists 1-8;or (2) a sequence that varies from the sequence of (1) only by one or more transitions, wherein each respective transition of the one or more transitions occurs at a cytosine in the genomic region.
  4. 8
    9. The bait set of any one of claims 1-8, wherein the plurality of different oligonucleotide-containing probes are each conjugated to an affinity moiety.
  5. 9
    10. The bait set of claim 9, wherein the affinity moiety is biotin.
  6. 10
    11. The bait set of any one of claims 1-10, wherein, for at least one of the different oligonucleotide-containing probes, the sequence of at least 30 bases is complementary to the sequence that varies from the sequence of (1) only by one or more transitions, wherein each respective transition of the one or more transitions occurs at a cytosine in the genomic region.
  7. 11
    12. The bait set of claim 11, wherein for at least 500, 1,000, 2,000, 2,500, 5,000, 6,000, 10,000, 15,000, 20,000, 25,000, or 50,000 of each of the different oligonucleotidecontaining probes, the sequence of at least 30 bases is complementary to the sequence that varies from the sequence of (1) only by one or more transitions, wherein each respective transition of the one or more transitions occurs at a cytosine in the genomic region.
  8. 12
    13. The bait set of any one of claims 1-12, wherein at least 80%, 90%, or 95% of the oligonucleotide-containing probes in the bait set do not include an at least 30, at least 40, or at least 45 base sequence that has 20 or more off-target regions of the genome.
  9. 13
    14. The bait set of any one of claims 1-13, wherein the oligonucleotide-containing probes in the bait set do not include an at least 30, at least 40, or at least 45 base sequence that has 20 or more off-targets regions of the genome.
  10. 14
    15. The bait set of any one of claims 1-14, wherein the sequence of at least 30 bases is at least 40 bases, at least 45 bases, at least 50 bases, at least 60 bases, at least 75, or at least 100 bases in length.
  11. 15
    16. The bait set of any one of claims 1-14, wherein each of the oligonucleotidecontaining probes has a nucleic acid sequence of at least 45, 40, 75, 100, or 120 bases in length.
  12. 16
    17. The bait set of any one of claims 1-16, wherein each of the oligonucleotidecontaining probes have a nucleic acid sequence of no more than 300, 250, 200, or 150 bases in length.
  13. 17
    18. The bait set of any one of claims 1-17, wherein each of the plurality of different oligonucleotide-containing probes is between 60 and 200 bases in length, between 100 and 150 bases in length, between 110 and 130 bases in length, and/or 120 bases in length.
  14. 18
    19. The bait set of any one of claims 1-18, wherein the different oligonucleotidecontaining probes comprise at least 500, at least 1000, at least 2,000, at least 2,500, at least 5,000, at least 6,000, at least 7,500, and least 10,000, at least 15,000, at least 20,000, or at least 25,000 different pairs of probes, wherein each pair of probes comprises a first probe and second probe, wherein the second probe differs from the first probe and overlaps with the first probe by an overlapping sequence that is at least 30, at least 40, at least 50, or at least 60 nucleotides in length.
  15. 19
    20. The bait set of any one of claims 1-19, wherein the bait set includes oligonucleotide-containing probes that are configured to target at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the genomic regions identified in any one of Lists 1-8.
  16. 20
    21. The bait set of claim 20, wherein the bait set include oligonucleotidecontaining probes that are configured to target at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the genomic regions identified in List 1.
  17. 22
    23. The bait set of any one of claims 1-22, wherein an entirety of oligonucleotide probes in the bait set are configured to hybridize to fragments obtained from cfDNA molecules corresponding to at least 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in a list selected from any one of Lists 1-8.
  18. 23
    24. The bait set of any one of claims 1-23, wherein an entirety of oligonucleotidecontaining probes in the bait set are configured to hybridize to fragments obtained from cfDNA molecules corresponding to at least 500, 1,000, 5000, 10,000, 15,000, 20,000, at least 25,000, or at least 30,000 genomic regions in any one of Lists 1-8.
  19. 24
    25. The bait set of any one of claims 1-24, wherein an entirety of oligonucleotidecontaining probes in the bait set are configured to hybridize to fragments obtained from cfDNA molecules corresponding to at least 50, 60, 70, 80, 90, 100, 120, 150, or 200 genomic regions in any one of Lists 1-8.
  20. 25
    26. The bait set of any one of claims 1-25, wherein the plurality of oligonucleotide-containing probes comprise at least 500, 1,000, 5,000, or 10,000 different subsets of probes, wherein each subset of probes comprises a plurality of probes that collectively extend across a genomic region selected from the genomic regions of any one of Lists 1-8 in a 2x tiled fashion.
  21. 26
    27. The bait set of claim 26, wherein the plurality of probes that collectively extend across the genomic region in a 2x tiled fashion comprises at least one pair of probes that overlap by a sequence of at least 30 bases, at least 40 bases, at least 50 bases, or at least 60 bases in length.
  22. 27
    28. The bait set of any one of claims 1-27, wherein the plurality of probes collectively extend across portions of the genome that collectively are a combined size of between 0.2 and 15 MB, between 0.5 MB and 15 MB, between 1 MB and 15 MB, between 3 MB and 12 MB, between 3 MB and 7, MB, between 5 MB and 9 MB, or between 7 MB and 12 MB.
  23. 28
    29. The bait set of any one of claims 1-28, wherein at least a subset of the different oligonucleotide-containing probes are designed to hybridize to cfDNA fragments derived from one or more genomics region from either List 4 or List 6.
  24. 29
    30. The bait set of claim 29, wherein the subset of the different oligonucleotidecontaining probes are designed to target at least 2, at least 10, at least 50, at least 100, at least 1000, or at least 5000, at least 8000, at least 10,000 or at least 20,000 of the genomic regions from either List 4 or List 6.
  25. 31
    32. The bait set of any one of claims 1-31, wherein each of the different oligonucleotide-containing probes comprises less than 20, 15, 10, 8, or 6 CpG detection sites.
  26. 32
    33. The bait set of any of claims 1-32, wherein at least 80%, 85%, 90%, 92%, 95%, or 98% of the plurality of oligonucleotide-containing probes have exclusively either CpG or CpA on all CpG detection sites.
  27. 33
    34. The bait set of any one of claims 1-33, wherein the oligonucleotide-containing probes of the bait set correspond with a number of genomic regions selected from the genomic regions of any one of Lists 1-8, wherein at least 30% of the genomic regions that correspond with the probes in the bait set are in exons or introns.
  28. 34
    35. The bait set of any one of claims 1-34, wherein the oligonucleotide-containing probes of the bait set correspond with a number of genomic regions, wherein at least 15% or at least 20% of the genomic regions that correspond with probes in the bait set are in exons.
  29. 35
    36. The bait set of any one of claims 1-35, wherein the oligonucleotide-containing probes of the bait set correspond with a number of genomic regions, wherein less than 10% of the genomic regions that correspond with probes in the bait set are intergenic regions.
  30. 36
    37. The bait set of any one of claims 1-36, wherein, for each of the different oligonucleotide-containing probes, the at least 30 nucleotide sequence is complementary to a sequence that varies from the sequence of the genomic region by one or more transitions at all CpG sites within the sequence.
  31. 37
    38. The bait set of any one of claims 1-37, wherein, for oligonucleotidecontaining probes that vary with respect to the sequence within the genomic region by one or more transitions, a transition occurs at each CpG site within the genomic region.
  32. 38
    39. The bait set of any one of claims 1-37, wherein the different oligonucleotidecontaining probes are complementary to cfDNA fragments that have been converted to replace cytosine with uracil, wherein the cfDNA fragments are found at least 2-fold, 10-fold, 20-fold, 50-fold, 100-fold, or 1000-fold more frequently in cfDNA from cancer subjects than from cfDNA from non-cancer subjects.
  33. 39
    40. A mixture comprising:converted cfDNA;and the bait set of any one of claims 1-39.
  34. 40
    41. The mixture of claim 40, wherein the converted cfDNA comprises bisulfiteconverted cfDNA.
  35. 42
    43. A method for enriching a converted cfDNA sample, the method comprising:contacting the converted cell-free DNA sample with the bait set of any one of claims 1-39;and enriching the sample for a first set of genomic regions by hybridization capture.
  36. 43
    44. A method for providing sequence information informative of a presence or absence of a cancer, a stage of cancer, or a type of cancer, the method comprising:processing cell-free DNA from a biological sample with a deaminating agent to generate a cell-free DNA sample comprising deaminated nucleotides;and enriching the cellfree DNA sample for informative cell-free DNA molecules, wherein enriching the cell-free DNA sample for informative cell-free DNA molecules comprises contacting the cell-free DNA with a plurality of probes that are configured to hybridize to cell-free DNA molecules that correspond to regions identified in any one of Lists 1-8;and sequencing the enriched cell-free DNA molecules, thereby obtaining a set of sequence reads informative of a presence or absence of a cancer, a stage of cancer, or a type of cancer.
  37. 44
    45. The method of claim 44, wherein the plurality of probes comprise a plurality of primers, and enriching the cell-free DNA comprises amplifying, via PCR, the cell-free DNA fragments using the primers.
  38. 48
    49. The method of any one of claims 44-48, further comprising determining a cancer classification by evaluating the set of sequence reads, wherein the cancer classification is (a) a presence or absence of cancer;(b) a stage of cancer;or (c) a presence or absence of a type of cancer.
  39. 49
    50. The method of claim 49, wherein the cancer classification is a presence or absence of cancer.
  40. 50
    51. The method of any one of claims 49-50, wherein the step of determining a cancer classification comprises:(a) generating a test feature vector based on the set of sequence reads;and (b) applying the test feature vector to a classifier.
  41. 51
    52. The method of claim 51, wherein the classifier comprises a model that is trained by a training process with a cancer set of fragments from one or more training subjects with cancer and a non-cancer set of fragments from one or more training subjects without cancer, wherein both the cancer set of fragments and the non-cancer set of fragments comprise a plurality of training fragments.
  42. 55
    56. The method of claim 55, wherein the stage of cancer is selected from Stage I, Stage II, Stage III, and Stage IV.
  43. 57
    58. The method of claim 57, wherein the step of determining a cancer classification comprises:(a) generating a test feature vector based on the set of sequence reads;and (b) applying the test feature vector to a classifier.
  44. 58
    59. The method of claim 58, wherein the classifier comprises a model that is trained by a training process with a cancer set of fragments from one or more training subjects with cancer and a non-cancer set of fragments from one or more training subjects without cancer, wherein both the cancer set of fragments and the non-cancer set of fragments comprise a plurality of training fragments.
  45. 59
    60. The method of any one of claims 57-59, wherein the type of cancer is selected from the group consisting of head and neck cancer, liver/bile duct cancer, upper GI cancer, pancreatic/gallbladder cancer;colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoid neoplasms, melanoma, sarcoma, breast cancer, and uterine cancer.
  46. 60
    61. The method of any one of claims 58-60, wherein the type of cancer is head and neck cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 85%, or at least 87%.
  47. 61
    62. The method of any one of claims 58-60, wherein the type of cancer is liver/bile duct cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 60%, at least 65%, at least 70%, or at least 73%.
  48. 62
    63. The method of any one of claims 58-60, wherein the type of cancer is an upper GI tract cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 75%, at least 80%, or at least 85%.
  49. 63
    64. The method of any one of claims 58-60, wherein the type of cancer is a pancreatic or gallbladder cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 85%, or at least 90%.
  50. 64
    65. The method of any one of claims 58-60, wherein the type of cancer is colorectal cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, at least 95%, or at least 98%.
  51. 65
    66. The method of any one of claims 58-60, wherein the type of cancer is ovarian cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 60%, at least 70%, at least 80%, at least 85%, or at least 87%.
  52. 66
    67. The method of any one of claims 58-60, wherein the type of cancer is lung cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, at least 95%, or at least 97%.
  53. 67
    68. The method of any one of claims 58-60, wherein the type of cancer is multiple myeloma, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 85%, or at least 90% or at least 93%.
  54. 68
    69. The method of any one of claims 58-60, wherein the type of cancer is a lymphoid neoplasm, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, or at least 95% or at least 98%.
  55. 69
    70. The method of any one of claims 58-60, wherein the type of cancer is a melanoma, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, or at least 95% or at least 98%.
  56. 70
    71. The method of any one of claims 58-60, wherein the type of cancer is a sarcoma, and the classifier, at 99.4% specificity, has a sensitivity of at least 35%, at least 40%, at least 45%, or at least 50%.
  57. 71
    72. The method of any one of claims 58-60, wherein the type of cancer is breast cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, or at least 95% or at least 98%.
  58. 72
    73. The method of any one of claims 58-60, wherein the type of cancer is uterine cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, or at least 95% or at least 97%.
  59. 73
    74. The method of any one of claims 49-73, wherein the step of determining a cancer classification comprises:(a) generating a test feature vector based on the set of sequence reads;and (b) applying the test feature vector to a model obtained by a training process with a cancer set of fragments from one or more training subjects with a cancer and a non-cancer set of fragments from one or more training subjects without cancer, wherein both the cancer set of fragments and the non-cancer set of fragments comprise a plurality of training fragments.
  60. 74
    75. The method of claim 74, wherein the training process comprises:(a) obtaining sequence information of training fragments from a plurality of training subjects;(b) for each training fragment, determining whether that training fragment is hypomethylated or hypermethylated, wherein each of the hypomethylated and hypermethylated training fragments comprises at least a threshold number of CpG sites with at least a threshold percentage of the CpG sites being unmethylated or methylated, respectively, (c) for each training subject, generating a training feature vector based on the hypomethylated training fragments and hypermethylated training fragments, and (d) training the model with the training feature vectors from the one or more training subjects without cancer and the training feature vectors from the one or more training subjects with cancer.
  61. 76
    77. The method of any one of claims 74-76, wherein the model comprises one of a kernel logistic regression classifier, a random forest classifier, a mixture model, a convolutional neural network, and an autoencoder model.
  62. 77
    78. The method of any one of claims 74-77, further comprising the steps of:(a) obtaining a cancer probability for the test sample based on the model;and (b) comparing the cancer probability to a threshold probability to determine whether the test sample is from a subject with cancer or without cancer.
  63. 78
    79. The method of claim 78, further comprising administering an anti-cancer agent to the subject.
  64. 79
    80. A method of treating a cancer patient, the method comprising:administering an anti-cancer agent to a subject who has been identified as a cancer subject by the method of claim 78.
  65. 80
    81. The method of claim 80, wherein the anti-cancer agent is a chemotherapeutic agent selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, anti-tumor antibiotics, cytoskeletal disruptors (taxans), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, and platinum-based agents.
  66. 81
    82. A method for assessing whether a subject has a cancer, the method comprising:obtaining cfDNA from the subject;isolating a portion of the cfDNA from the subject by hybridization capture;obtaining sequence reads derived from the captured cfDNA to determine methylation states cfDNA fragments;applying a classifier to the sequence reads;and determining whether the subject has cancer based on application of the classifier;wherein the classifier has an area under the receiver operator characteristic curve of greater than 0.70, greater than 0.75, greater than 0.77, greater than 0.80, greater than 0.81, greater than 0.82, or greater than 0.83.
  67. 82
    83. The method of claim 82, further comprising converting unmethylated cytosines in the cfDNA to uracil prior to isolating the portion of the cfDNA from the subject by hybridization capture.
  68. 84
    85. The method of any one of claims 82-84, wherein the classifier is a binary classifier.
  69. 85
    86. The method of any one of claims 82-85, wherein isolating a portion of the cfDNA from the subject by hybridization capture comprises contacting the cell-free DNA with a bait set comprising a plurality of different oligonucleotide-containing probes.
  70. 86
    87. The method of any one of claim 86, wherein the bait set is the bait set of any one of claims 1-39.
  71. 87
    88. A method for identifying genomic regions that exhibit differential methylation in cancer samples relative to non-cancer samples, the method comprising:(a) obtaining sequence reads of converted cfDNA from both cancer subjects and noncancer subjects;(b) identifying, based on the sequence reads, cfDNA fragments that: (i) have a p-value rarity in non-cancerous samples of below a threshold value;and (ii) have at least X CpG sites, wherein at least Y% of the CpG sites are methylated, wherein X is at least 4, 5, 6, 7, 8, 9, or 10 and Y is at least 70;and (c) for each of a plurality of CpG sites in a reference genome, counting both (1) a number of cancer subjects (Ncancer) and (2) a number of non-cancer subjects (Nnoncancer) that have a fragment identified in step (b);(d) for each of the plurality of CpG sites in the reference genome, determining whether the CpG site is differentially methylated in cancer samples based on criteria comprising Ncancer and Nnon-cancer;(e) identifying a genomic region as differentially methylated in cancer based, at least in part, on inclusion of a differentially methylated CpG site within the genomic region.
  72. 88
    89. A method for identifying genomic regions that exhibit differential methylation in cancer samples relative to non-cancer samples, the method comprising:(a) obtaining sequence reads of converted cfDNA from both cancer subjects and noncancer subjects;(b) identifying, based on the sequence reads, cfDNA fragments that: (i) have at least X CpG sites, wherein at least Y% of the CpG sites are unmethylated, wherein X is 4, 5, 6, 7, 8, 9, or 10 and Y is at least 70;and (ii) have a p-value rarity in non-cancerous samples of below a threshold value;(c) for each of a plurality of CpG sites in a reference genome, counting both (1) a number of cancer subjects (Ncancer) and (2) a number of non-cancer subjects (Nnoncancer) that have a fragment identified in step (b);(d) for each of the plurality of CpG sites in the reference genome, determining whether the CpG site is differentially methylated in cancer samples based on criteria comprising Ncancer and Nnon-cancer;(e) identifying a genomic region as differentially methylated in cancer based, at least in part, on inclusion of a differentially methylated CpG site within the genomic region.
  73. 90
    91. The method of claim 90, wherein the CpG site is considered to be differentially methylated when (Ncancer + 1)/(Ncancer + Nnon-cancer + 2) is greater than a threshold value.
  74. 91
    92. The method of any one of claims 89-91, wherein each of the identified genomic regions has at least X CpG sites, wherein X is 4, 5, or 6.
  75. 92
    93. The method of any one of claims 89-92, wherein at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the identified regions are from any one of Lists 1-8.
  76. 93
    94. A method for developing a bait set for hybridization capture of cfDNA from genomic regions that are differentially methylated between cancer and non-cancer, the method comprising:identifying at least 1000, at least 5,000, at least 10,000, at least 25,000, or at least 30,000 differentially methylated genomic regions of the genome by comparison of one or more parameters derived from cfDNA fragments from cancer subject to one or more parameters derived from cfDNA fragments from non-cancer subjects;and designing, in silico, a plurality of oligonucleotide-containing probes that include a sequence of at least 30 bases in length that is complementary to either (1) a sequence of a genomic region or (2) a sequence that differs from the sequence of the genomic region only by one or more transitions, wherein each respective transition occurs at a cytosine in the genomic region.
  77. 94
    95. The method of claim 94, the method comprising removing, in silico, probes that have at least X off-target regions, wherein X is at least one.
  78. 95
    96. The method of claim 95, wherein X is at least 5, at least 10, or at least 20.
  79. 96
    97. The method of any one of claims 94-96, wherein the differentially methylated regions are identified via the method of any one of claims 88-93.
  80. 98
    99. A method for selecting probes for hybridization capture of cfDNA, the method comprising:identifying a first set of genomic regions that are preferentially hypermethylated in cfDNA from cancer subjects relative to non-cancer subjects;identifying a second set of genomic regions that are preferentially hypomethylated in cfDNA from cancer subjects relative to non-cancer subjects;and selecting probes for hybridization capture of cfDNA corresponding to the first set of genomic regions and the second set of genomic regions, wherein the probes comprise a first set of probes for hybridization capture of cfDNA corresponding to the first set of genomic regions and a second set of probes for hybridization capture of cfDNA corresponding to the second set of genomic regions;wherein the probes comprise at least 500, at least 1,000, at least 2,500, at least 5,000, at least 10,000, at least 20,000 subsets of probes, wherein each subset of probes comprises a plurality of probes that extend across a genomic region in a 2x tiled fashion.
  81. 99
    100. The method of claim 99, wherein the second set of probes for hybridization capture comprises selecting probes that differ from a sequence in the genomic region only by one or more transitions, wherein each transition occurs at a nucleotide corresponding to a cytosine in the genomic region.
  82. 100
    101. The method of any one of claims 99-100, wherein selecting probes for hybridization capture comprises filtering out probes that have more than a threshold number of off-target regions.
  83. 101
    102. The method of any one of claims 99-101, wherein each subset of probes comprises at least three probes.
  84. 102
    103. The method of any one of claims 99-102, wherein each probe is between 75 and 200, between 100 and 150, between 110 and 130, or 120 nucleotides in length.
  85. 103
    104. An assay panel for enriching cfDNA molecules for cancer diagnosis, the assay panel comprising:at least 500 different pairs of polynucleotide probes, wherein each pair of the at least 500 pairs of probes (i) comprises two different probes configured to overlap with each other by an overlapping sequence of 30 or more nucleotides and (ii) is configured to hybridize to a modified fragment obtained from processing of the cfDNA molecules, wherein each of the cfDNA molecules corresponds to or is derived from one or more genomic regions, wherein each of the one or more genomic regions comprises at least five methylation sites and has an anomalous methylation pattern in cancerous training samples relative to non-cancerous training samples.
  86. 104
    105. The assay panel of claim 104, wherein the overlapping sequence comprises at least 40, 50, 75, or 100 nucleotides.
  87. 105
    106. The assay panel of any one of claims 104-105, comprising at least 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, or 25,000 pairs of probes.
  88. 106
    107. An assay panel for enriching cfDNA molecules for cancer diagnosis, the assay panel comprising:at least 1,000 polynucleotide probes, wherein each of the at least 1,000 probes is configured to hybridize to a modified polynucleotide obtained from processing of the cfDNA molecules, wherein each of the cfDNA molecules corresponds to or is derived from, one or more genomic regions, wherein each of the one or more genomic regions comprises at least five methylation sites, and has an anomalous methylation pattern in cancerous training samples relative to non-cancerous samples.
  89. 107
    108. The assay panel of any one of claims 104-107, wherein the processing of the cfDNA molecules comprises converting unmethylated C (cytosine) to U (uracil) in the cfDNA molecules.
  90. 108
    109. The assay panel of any one of claims 104-108, wherein each of the polynucleotide probes on the panel is conjugated to an affinity moiety.
  91. 109
    110. The assay panel of claim 109, wherein the affinity moiety is a biotin moiety.
  92. 110
    111. The assay panel of any one of claims 104-110, wherein each of the one or more genomic regions is either hypermethylated or hypomethylated in the cancerous training samples relative to non-cancerous reference samples.
  93. 111
    112. The assay panel of any one of claims 104-111, wherein at least 80%, 85%, 90%, 92%, 95%, or 98% of the probes on the panel have exclusively either CpG or CpA on CpG detection sites.
  94. 112
    113. The assay panel of any one of claims 104-112, wherein each of the probes on the panel comprises less than 20, 15, 10, 8, or 6 CpG detection sites.
  95. 113
    114. The assay panel of any one of claims 104-113, wherein each of the probes on the panel is designed to have fewer than 20, 15, 10, or 8 off-target genomic regions.
  96. 114
    115. The assay panel of claim 114, wherein the fewer than 20 off-target genomic regions are identified using a k-mer seeding strategy.
  97. 115
    116. The assay panel of claim 115, wherein the fewer than 20 off-target genomic regions are identified using k-mer seeding strategy combined to local alignment at seed locations.
  98. 116
    117. The assay panel of any one of claims 104-116, comprising at least 1,000, 2,000, 2,500, 5,000, 10,000, 12,000, 15,000, 20,000, or 25,000 probes.
  99. 117
    118. The assay panel of any one of claims 104-117, wherein the at least 500 pairs of probes or the at least 1,000 probes together comprise at least 0.2 million, 0.4 million, 0.6 million, 0.8 million, 1 million, 2 million, 4 million, or 6 million nucleotides.
  100. 118
    119. The assay panel of any one of claims 104-118, wherein each of the probes on the panel comprises at least 50, 75, 100, or 120 nucleotides.
  101. 119
    120. The assay panel of any one of claims 104-119, wherein each of the probes on the panel comprises less than 300, 250, 200, or 150 nucleotides.
  102. 120
    121. The assay panel of any one of claims 104-120, wherein each of the probes on the panel comprises 100-150 nucleotides.
  103. 121
    122. The assay panel of any one of claims 104-121, wherein at least 30% of the genomic regions are in exons or introns.
  104. 122
    123. The assay panel of any one of claims 104-122, wherein at least 15% of the genomic regions are in exons.
  105. 123
    124. The assay panel of any one of claims 104-123, wherein at least 20% of the genomic regions are in exons.
  106. 124
    125. The assay panel of any one of claims 104-124, wherein less than 10% of the genomic regions are in intergenic regions.
  107. 125
    126. The assay panel of any one claims 104-125, wherein each of the one or more genomic regions is selected from one of Lists 1-8.
  108. 126
    127. The assay panel of any one claims 104-126, wherein an entirety of probes on the panel together are configured to hybridize to modified fragments obtained from the cfDNA molecules corresponding to or derived from at least 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in one or more of Lists 1-8 .
  109. 127
    128. The assay panel of any one of claims 104-127, an entirety of probes on the panel together are configured to hybridize to modified fragments obtained from the cfDNA molecules corresponding to or derived from at least 500, 1,000, 5000, 10,000 or 15,000 genomic regions in one or more of Lists 1-8.
  110. 128
    129. An assay panel for enriching cfDNA molecules for cancer diagnosis, comprising a plurality of polynucleotide probes, wherein each of the polynucleotide probes is configured to hybridize to a modified fragment obtained from processing of the cfDNA molecules, wherein each of the cfDNA molecules corresponds to or is derived from one or more genomic regions selected from one or more of Lists 1-8.
  111. 129
    130. The assay panel of claim 129, wherein each of the cfDNA molecules corresponds to or is derived from one or more genomic regions selected from one or more of Lists 1-8.
  112. 132
    133. The assay panel of any one of claims 129-132, wherein the processing of the cfDNA molecules comprises converting unmethylated C (cytosine) to U (uracil) in the cfDNA molecules.
  113. 133
    134. The assay panel of claim 133, wherein each of probes on the panel is conjugated to an affinity moiety.
  114. 134
    135. The assay panel of claim 134, wherein the affinity moiety is biotin.
  115. 135
    136. The assay panel of any of claims 129-135, wherein at least 80%, 85%, 90%, 92%, 95%, or 98% of the probes on the panel have exclusively either CpG or CpA on CpG detection sites.
  116. 136
    137. A method of providing sequence information informative of a presence or absence of cancer, the method comprising the steps of:(a) obtaining a test sample comprising a plurality of cfDNA test molecules;(b) processing the cfDNA test molecules, thereby obtaining converted test fragments;(c) contacting the converted test fragments with an assay panel, thereby enriching a subset of the converted test fragments by hybridization capture;and (d) sequencing the subset of the converted test fragments, thereby obtaining a set of sequence reads.
  117. 137
    138. The method of claim 137, wherein the converted test fragments are bisulfiteconverted test fragments.
  118. 139
    140. The method of any one of claims 137-139, further comprising determining a cancer classification by evaluating the set of sequence reads, wherein the cancer classification is (a) a presence or absence of cancer;(b) a stage of cancer;(c) a presence or absence of a type of cancer.
  119. 140
    141. The method of claim 140, wherein the cancer classification is a presence or absence of cancer.
  120. 141
    142. The method of any one of claims 140 and 141, wherein the step of determining a cancer classification comprises:(a) generating a test feature vector based on the set of sequence reads;and (b) applying the test feature vector to a classifier.
  121. 142
    143. The method of claim 142, wherein the classifier comprises a model that is trained by a training process with a cancer set of fragments from one or more training subjects with cancer and a non-cancer set of fragments from one or more training subjects without cancer, wherein both cancer set of fragments and the non-cancer set of fragments comprise a plurality of training fragments.
  122. 146
    147. The method of claim 146, wherein the stage of cancer is selected from Stage I, Stage II, Stage III, and Stage IV.
  123. 148
    149. The method of claim 148, wherein the step of determining a cancer classification comprises:(a) generating a test feature vector based on the set of sequence reads;and (b) applying the test feature vector to a classifier.
  124. 149
    150. The method of claim 149, wherein the classifier comprises a model that is trained by a training process with a cancer set of fragments from one or more training subjects with cancer and a non-cancer set of fragments from one or more training subjects without cancer, wherein both the cancer set of fragments and the non-cancer set of fragments comprise a plurality of training fragments.
  125. 150
    151. The method of any one of claims 148-150, wherein the type of cancer is selected from the group consisting of head and neck cancer, liver/bile duct cancer, upper GI cancer, pancreatic/gallbladder cancer;colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoid neoplasms, melanoma, sarcoma, breast cancer, and uterine cancer.
  126. 151
    152. The method of any one of claims 149-150, wherein the type of cancer is head and neck cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 85%, or at least 87%.
  127. 152
    153. The method of any one of claims 149-150, wherein the type of cancer is liver/bile duct cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 60%, at least 65%, at least 70%, or at least 73%.
  128. 153
    154. The method of any one of claims 149-150, wherein the type of cancer is an upper GI tract cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 75%, at least 80%, or at least 85%.
  129. 154
    155. The method of any one of claims 149-150, wherein the type of cancer is a pancreatic/gallbladder, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 85%, or at least 90%.
  130. 155
    156. The method of any one of claims 149-150, wherein the type of cancer is colorectal cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, at least 95%, or at least 98%.
  131. 156
    157. The method of any one of claims 149-150, wherein the type of cancer is ovarian cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 85%, or at least 87%.
  132. 157
    158. The method of any one of claims 149-150, wherein the type of cancer is lung cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, at least 95%, or at least 97%.
  133. 158
    159. The method of any one of claims 149-150, wherein the type of cancer is multiple myeloma, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 85%, or at least 90% or at least 93%.
  134. 159
    160. The method of any one of claims 149-150, wherein the type of cancer is a lymphoid neoplasm, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, or at least 95% or at least 98%.
  135. 160
    161. The method of any one of claims 149-150, wherein the type of cancer is a melanoma, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, or at least 95% or at least 98%.
  136. 161
    162. The method of any one of claims 149-150, wherein the type of cancer is a sarcoma, and the classifier, at 99.4% specificity, has a sensitivity of at least 35%, at least 40%, at least 45%, or at least 50%.
  137. 162
    163. The method of any one of claims 149-150, wherein the type of cancer is breast cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, or at least 95% or at least 98%.
  138. 163
    164. The method of any one of claims 149-150, wherein the type of cancer is uterine cancer, and the classifier, at 99.4% specificity, has a sensitivity of at least 70%, at least 80%, at least 90%, or at least 95% or at least 97%.
  139. 164
    165. The method of claim any one of claims 140-164, wherein the polynucleotide probes together are configured to hybridize to converted fragments obtained from the cfDNA molecules corresponding to or derived from at least 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in any one of Lists 1-8.
  140. 165
    166. The method of any one of claims 140-165, wherein the step of determining a cancer classification is performed by the method comprising the steps of:(a) generating a test feature vector based on the set of sequence reads;and (b) applying the test feature vector to a model obtained by a training process with a cancer set of fragments from one or more training subjects with a cancer and a non-cancer set of fragments from one or more training subjects without cancer, wherein both the cancer set of fragments and the non-cancer set of fragments comprise a plurality of training fragments.
  141. 166
    167. The method of claim 166, wherein the training process comprises:(a) obtaining sequence information of training fragments from a plurality of training subjects;(b) for each training fragment, determining whether that training fragment is hypomethylated or hypermethylated, wherein each of the hypomethylated and hypermethylated training fragments comprises at least a threshold number of CpG sites with at least a threshold percentage of the CpG sites being unmethylated or methylated, respectively, (c) for each training subject, generating a training feature vector based on the hypomethylated training fragments and a training feature vector based on the hypermethylated training fragments, and (d) training the model with the training feature vectors from the one or more training subjects without cancer and the training feature vectors from the one or more training subjects with cancer.
  142. 168
    169. The method of any one of claims 166-168, wherein the model comprises one of a kernel logistic regression classifier, a random forest classifier, a mixture model, a convolutional neural network, and an autoencoder model.
  143. 169
    170. The method of any one of claims 166-169, further comprising the steps of:(a) obtaining a cancer probability for the test sample based on the model;and (b) comparing the cancer probability to a threshold probability to determine whether the test sample is from a subject with cancer or without cancer.
  144. 170
    171. The method of claim 170, further comprising administering an anti-cancer agent to the subject.
  145. 171
    172. A method of treating a cancer patient, the method comprising:administering an anti-cancer agent to a subject who has been identified as a cancer subject by the method of claim 170.
  146. 173
    174. A method comprising the steps of:(a) obtaining a set of sequence reads of modified test fragments, wherein the modified test fragments are or have been obtained by processing a set of nucleic acid fragments from a test subject, wherein each of the nucleic acid fragments corresponds to or is derived from a plurality of genomic regions selected from any one of Lists 1-8;and (b) applying the set of sequence reads or a test feature vector obtained based on the set of sequence reads to a model obtained by a training process with a cancer set of fragments from one or more training subjects with cancer and a non-cancer set of fragments from one or more training subjects without cancer, wherein both cancer set of fragments and the non-cancer set of fragments comprise a plurality of training fragments.
  147. 174
    175. The method of claim 174, further comprising the step of obtaining the test feature vector comprising:(a) for each of the nucleic acid fragments, determining whether the nucleic acid fragment is hypomethylated or hypermethylated, wherein each of the hypomethylated and hypermethylated nucleic acid fragments comprises at least a threshold number of CpG sites with at least a threshold percentage of the CpG sites being unmethylated or methylated, respectively;(b) for each of a plurality of CpG sites in a reference genome: quantifying a count of hypomethylated nucleic acid fragments which overlap the CpG site and a count of hypermethylated nucleic acid fragments which overlap the CpG site;and generating a hypomethylation score and a hypermethylation score based on the count of hypomethylated nucleic acid fragments and hypermethylated nucleic acid fragments;(c) for each nucleic acid fragment, generating an aggregate hypomethylation score based on the hypomethylation score of the CpG sites in the nucleic acid fragment and an aggregate hypermethylation score based on the hypermethylation score of the CpG sites in the nucleic acid fragment;(d) ranking the plurality of nucleic acid fragments based on aggregate hypomethylation score and ranking the plurality of nucleic fragments based on aggregate hypermethylation score;and (e) generating the test feature vector based on the ranking of the nucleic acid fragments.
  148. 175
    176. The method of any one of claims 174-175, wherein the training process comprises:(a) for each training fragment, determining whether that training fragment is hypomethylated or hypermethylated, wherein each of the hypomethylated and hypermethylated training fragments comprises at least a threshold number of CpG sites with at least a threshold percentage of the CpG sites being unmethylated or methylated, respectively, (b) for each training subject, generating a training feature vector based on the hypomethylated training fragments and a training feature vector based on the hypermethylated training fragments, and (c) training the model with the training feature vectors from the one or more training subjects without cancer and the feature vectors from the one or more training subjects with cancer.
  149. 176
    177. The method of any one of claims 174-175, wherein the training process comprises:(a) for each training fragment, determining whether that training fragment is hypomethylated or hypermethylated, wherein each of the hypomethylated and hypermethylated training fragments comprises at least a threshold number of CpG sites with at least a threshold percentage of the CpG sites being unmethylated or methylated, respectively, (b) for each of a plurality of CpG sites in a reference genome: quantifying a count of hypomethylated training fragments which overlap the CpG site and a count of hypermethylated training fragments which overlap the CpG site;and generating a hypomethylation score and a hypermethylation score based on the count of hypomethylated training fragments and hypermethylated training fragments;(c) for each training fragment, generating an aggregate hypomethylation score based on the hypomethylation score of the CpG sites in the training fragment and an aggregate hypermethylation score based on the hypermethylation score of the CpG sites in the training fragment;(d) for each training subject: ranking the plurality of training fragments based on aggregate hypomethylation score and ranking the plurality of training fragments based on aggregate hypermethylation score;and generating a feature vector based on the ranking of the training fragments;(e) obtaining training feature vectors for one or more training subjects without cancer and training feature vectors for the one or more training subjects with cancer;and (f) training the model with the feature vectors for the one or more training subjects without cancer and the feature vectors for the one or more training subjects with cancer.
  150. 177
    178. The method of claim 177, wherein for each CpG site in a reference genome, quantifying a count of hypomethylated training fragments which overlap that CpG site and a count of hypermethylated training fragments which overlap that CpG site further comprises:(a) quantifying a cancer count of hypomethylated training fragments from the one or more training subjects with cancer that overlap that CpG site and a non-cancer count of hypomethylated training fragments from the one or more training subjects without cancer that overlap that CpG site;and (b) quantifying a cancer count of hypermethylated training fragments from the one or more training subjects with cancer that overlap that CpG site and a non-cancer count of hypermethylated training fragments from the one or more training subjects without cancer that overlap that CpG site.
  151. 178
    179. The method of claim 178, wherein for each CpG site in a reference genome, generating a hypomethylation score and a hypermethylation score based on the count of hypomethylated training fragments and hypermethylated training fragments further comprises:(a) for generating the hypomethylation score, calculating a hypomethylation ratio of the cancer count of hypomethylated training fragments over a hypomethylation sum of the cancer count of hypomethylated training fragments and the noncancer count of hypomethylated training fragments;and (b) for generating the hypermethylation score, calculating a hypermethylation ratio of the cancer count of hypermethylated training fragments over a hypermethylation sum of the cancer count of hypermethylated training fragments and the non-cancer count of hypermethylated training fragments.
  152. 179
    180. The method of any one of claims 174-179, wherein the model comprises one of a kernel logistic regression classifier, a random forest classifier, a mixture model, a convolutional neural network, and an autoencoder model.
  153. 180
    181. The method of any one of claims 174-180, wherein the set of sequence reads is obtained by using the assay panel of any one of claims 104-136.
  154. 181
    182. A method of designing an assay panel for cancer diagnosis, comprising the steps of:(a) identifying a plurality of genomic regions, wherein each of the plurality of genomic regions (i) comprises at least 30 nucleotides, and (ii) comprises at least five methylation sites, (b) selecting a subset of the genomic regions, wherein the selection is made when cfDNA molecules corresponding to or derived from each of the genomic regions in cancer training samples have an anomalous methylation pattern, wherein the anomalous methylation pattern comprises at least five methylation sites known to be, or identified as, either hypomethylated or hypermethylated, and (c) designing the assay panel comprising a plurality of probes, wherein each of the probes is configured to hybridize to a modified fragment obtained from processing cfDNA molecules corresponding to or derived from one or more of the subset of the genomic regions.
  155. 182
    183. The method of claim 182, wherein the processing of the cfDNA molecules comprises converting unmethylated C (cytosine) to U (uracil) in the cfDNA molecules.
  156. 183
    184. A cancer assay panel, comprising:at least 500 pairs of probes, wherein each pair of the at least 500 pairs comprises two probes configured to overlap each other by an overlapping sequence, wherein the overlapping sequence comprises a 30-nucleotide sequence, and wherein the 30-nucleotide sequence is configured to have sequence complementarity with one or more genomic regions, wherein the one or more genomic regions have at least five methylation sites, and wherein the at least five methylation sites have an abnormal methylation pattern in non-cancerous samples or cancerous samples.
  157. 184
    185. The cancer assay panel of claim 184, wherein the overlapping sequence comprises at least 40, 50, 75, or 100 nucleotides.
  158. 185
    186. The cancer assay panel of any one of claims 184-185, comprising at least 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000 or 25,000 pairs of probes.
  159. 186
    187. A cancer assay panel, comprising:at least 1,000 probes, wherein each of the probes is designed as a hybridization probe complementary to one or more genomic regions, wherein each of the genomic regions comprises: (i) at least 30 nucleotides, and (ii) at least five methylation sites, wherein the at least five methylation sites have an abnormal methylation pattern and are either hypomethylated or hypermethylated in cancerous samples or non-cancerous samples.
  160. 187
    188. The cancer assay panel of any one or claims 184-187, wherein the abnormal methylation pattern has at least a threshold p-value rarity in the non-cancerous samples.
  161. 188
    189. The cancer assay panel of any one of claims 184-188, wherein each of the probes is designed to have less than 20 off-target genomic regions.
  162. 189
    190. The cancer assay panel of claim 189, wherein the less than 20 off-target genomic regions are identified using a k-mer seeding strategy.
  163. 190
    191. The cancer assay panel of claim 190, wherein the less than 20 off-target genomic regions are identified using k-mer seeding strategy combined to local alignment at seed locations.
  164. 191
    192. The cancer assay panel of any one of claims 184-191, wherein each of the genomic regions was selected based on criteria comprising:(a) a number (Ncancer) of the cancerous samples including at least one cfDNA fragment having the abnormal methylation pattern;and (b) a number (Nnon-cancer) of the non-cancerous samples including at least one cfDNA fragment having the abnormal methylation pattern.
  165. 192
    193. The cancer assay panel of claim 192, wherein each of the genomic regions was selected based on criteria positively correlated to Ncancer and inversely correlated to the sum of Ncancer and N non-cancer.
  166. 193
    194. The cancer assay panel of any one of claims 184-193, comprising at least 1,000, 2,000, 2,500, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, or 50,000 probes.
  167. 194
    195. The cancer assay panel of any one of claims 184-194, wherein the at least 500 pairs of probes or the at least 1,000 probes together comprise at least 0.2 million, 0.4 million, 0.6 million, 0.8 million, 1 million, 2 million, 3 million, 4 million, 5 million, or 6 million nucleotides.
  168. 195
    196. The cancer assay panel of any one of claims 184-195, wherein each of the probes comprises at least 50, 75, 100, or 120 nucleotides.
  169. 196
    197. The cancer assay panel of any one of claims 184-196, wherein each of the probes comprises less than 300, 250, 200, or 150 nucleotides.
  170. 197
    198. The cancer assay panel of any one of claims 184-197, wherein each of the probes comprises 100-150 nucleotides.
  171. 198
    199. The cancer assay panel of any one of claims 184-198, wherein each of the probes comprises less than 20, 15, 10, 8, or 6 methylation sites.
  172. 199
    200. The cancer assay panel of any one of claims 184-199, wherein at least 80%, 85%, 90%, 92%, 95%, or 98% of the at least five methylation sites are either methylated or unmethylated in the cancerous samples.
  173. 200
    201. The cancer assay panel of any one of claims 184-200, wherein each of the probes is configured to have less than 20, 15, 10, or 8 off-target genomic regions.
  174. 201
    202. The cancer assay panel of any one of claims 184-201, wherein at least 30% of the genomic regions are in exons or introns.
  175. 202
    203. The cancer assay panel of any one of claims 184-202, wherein at least 15% of the genomic regions are in exons.
  176. 203
    204. The cancer assay panel of any one of claims 184-203, wherein at least 20% of the genomic regions are in exons.
  177. 204
    205. The cancer assay panel of any one of claims 184-204, wherein less than 10% of the genomic regions are in intergenic regions.
  178. 205
    206. The cancer assay panel of any one of claims 184-205, wherein the genomic regions are selected from any one of Lists 1-8.
  179. 206
    207. The cancer assay panel of any one of claims 184-206, wherein the genomic regions are selected from List 3.
  180. 207
    208. The cancer assay panel of any one of claims 184-207, wherein the genomic regions are selected from List 5.
  181. 208
    209. The cancer assay panel of any one of claims 184-208, wherein the genomic regions are selected from List 8.
  182. 209
    210. The cancer assay panel of any one of claims 184-209, wherein the genomic regions comprise at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in any one of Lists 1-8.
  183. 210
    211. The cancer assay panel of any of any one of claims 184-210, wherein the genomic regions comprise at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in List 3.
  184. 211
    212. The cancer assay panel of any one of claims 184-211, wherein the genomic regions comprise at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in List 5.
  185. 212
    213. The cancer assay panel of any one of claims 184-212, wherein the genomic regions comprise at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in List 8.
  186. 213
    214. The cancer assay panel of any one of claims 184-213, wherein the at least 1,000 or at least 2,000 probes are configured to be complementary to at least 500, 1,000, 5000, 10,000 or 15,000 genomic regions in any one of Lists 1-8.
  187. 214
    215. The cancer assay panel of any one of claims 184-214, wherein the at least 1,000 or at least 2,000 probes are configured to be complementary to at least 500, 1,000, 5000, 10,000 or 15,000 genomic regions in List 3.
  188. 215
    216. The cancer assay panel of any one of claims 184-215, wherein the at least 1,000 or at least 2,000 probes are configured to be complementary to at least 500, 1,000, 5000, 10,000 or 15,000 genomic regions in List 5.
  189. 216
    217. The cancer assay panel of any one of claims 184-216, wherein the at least 1,000 or at least 2,000 probes are configured to be complementary to at least 1,000, 5000, 10,000 or 15,000 genomic regions in List 8.
  190. 217
    218. The cancer assay panel of any one of claims 184-217, wherein the 30nucleotide sequence comprises at least five CpG detection sites, wherein at least 80% of the at least five CpG detection sites comprise CpG, UpG, or CpA.
  191. 218
    219. A cancer assay method comprising:receiving a sample comprising a plurality of nucleic acid fragments;treating the plurality of nucleic acid fragments to convert unmethylated cytosine to uracil, thereby obtaining a plurality of converted nucleic acid fragments;hybridizing the plurality of converted nucleic acid fragments with the probes on the cancer assay panel of any of above claims;enriching a subset of the plurality of converted nucleic acid fragments;and sequencing the enriched subset of the converted nucleic acid fragments, thereby providing a set of sequence reads.
  192. 219
    220. The method of claim 219, further comprising the step of:determining a health condition by evaluating the set of sequence reads, wherein the health condition is (i) a presence or absence of cancer, or (ii) a stage of cancer.
  193. 220
    221. The method of any of claims 219-220, wherein the set of nucleic acid fragments is obtained from a human subject.
  194. 221
    222. A method for diagnosing cancer, comprising the steps of:(a) obtaining a set of sequence reads by sequencing a set of nucleic acid fragments from a subject;(b) determining methylation status of a plurality of genomic regions, the plurality of genomic regions comprise genomic regions selected from the genomic regions in any one of Lists 1-8;and (c) determining a health condition of the subject by evaluating the methylation status, wherein the health condition is (i) a presence or absence of cancer;or (ii) a stage of cancer.
  195. 222
    223. The method of claim 222, wherein the genomic regions are selected from List 3.
  196. 223
    224. The method of claim 223, wherein the genomic regions are selected from List 5.
  197. 224
    225. The method of claim 224, wherein the genomic regions are selected from List 8.
  198. 225
    226. A method for diagnosing cancer, comprising the steps of:(a) obtaining a set of sequence reads by sequencing a set of nucleic acid fragments from a subject;(b) determining methylation status of a plurality of genomic regions, the plurality of genomic regions comprise at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in any one of Lists 1-8;and (c) determining a health condition of the subject by evaluating the methylation status, wherein the health condition is (i) a presence or absence of cancer;or (ii) a stage of cancer.
  199. 226
    227. The method of claim 226, wherein the genomic regions comprise at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% of the genomic regions in List 3.
  200. 229
    230. A method for diagnosing cancer, comprising the steps of:(a) obtaining a set of sequence reads by sequencing a set of nucleic acid fragments from a subject;(b) determining methylation status of a plurality of at least 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000 or 25,000 genomic regions among genomic regions in any one of Lists 1-8;and (c) determining a health condition of the subject by evaluating the methylation status, wherein the health condition is (i) a presence or absence of cancer;or (ii) a stage of cancer.
  201. 230
    231. The method of claim 230, wherein the at least 1,000, 2,000 probes are configured to be complementary to at least 500, 1,000, 5000, 10,000 or 15,000 genomic regions in List 3.
  202. 233
    234. A method of designing a cancer assay panel comprising the steps of:identifying a plurality of genomic regions, wherein each of the plurality of genomic regions (i) comprises at least 30 nucleotides, and (ii) comprises at least five methylation sites, wherein the at least five methylation sites are either hypomethylated or hypermethylated, comparing methylation status of the at least five methylation sites in each of the plurality of genomic regions between cancerous samples and non-cancerous samples, selecting a subset of the genomic regions, wherein at least five methylation sites of the subset of the genomic regions have an abnormal methylation pattern in cancerous samples relative to non-cancerous samples, and designing a cancer assay panel comprising a plurality of probe sets, wherein each of the plurality of probe sets comprises at least a pair of probes configured to target one of the subset of the genomic regions.
  203. 234
    235. The method of claim 234, wherein the abnormal methylation pattern matches that of a cfDNA fragment from the cancerous samples overlapping at least one of the at least five methylation sites, wherein the cfDNA has at least a threshold p-value rarity relative to a training data set of the non-cancerous samples.
  204. 235
    236. The method of any of claims 234-235, wherein the step of selecting is performed based on criteria comprising:(a) a number Ncancer of the cancerous samples including cfDNA fragments having the abnormal methylation pattern;and (b) a number Nnon-cancer of the non-cancerous samples including cfDNA fragments having the abnormal methylation pattern.
  205. 236
    237. The method of claim 236, wherein the step of selecting is based on criteria positively correlated to Ncancer and inversely correlated to the sum of Ncancer and Nnon-cancer.
  206. 237
    238. The method of any of claims 234-237, wherein each of the plurality of probes has less than 20, 15, 10 or 8 off-target genomic regions.
  207. 238
    239. The cancer assay panel of claim 238, wherein the less than 20, 15, 10, or 8 offtarget genomic regions are identified using a k-mer seeding strategy.
  208. 239
    240. The cancer assay panel of claim 239, wherein the less than 20, 15, 10 or 8 offtarget genomic regions are identified using k-mer seeding strategy combined to local alignment at seed locations.
  209. 240
    241. The method of any of claims 234-240, further comprising the step of making the cancer assay panel comprising the plurality of probes.
  210. 241
    242. A cancer assay panel made by the method of claim 241.
  211. 242
    243. The cancer assay panel of claim 242, wherein the subset of genomic regions comprises genomic regions of any one of Lists 1-8.
  212. 254
    255. A cancer assay panel comprising a plurality of probes, wherein each of the plurality of probes is configured to overlap with one of the genomic regions in any one of Lists 1-8, and the plurality of probes together overlap with at least 90%, 95% or 100% of the genomic regions in any one of Lists 1-8.
  213. 255
    256. A cancer assay panel comprising a plurality of probes, wherein each of the plurality of probes is configured to overlap with one of the genomic regions in any one of Lists 1-8, and the plurality of probes together overlap with at least 500, 1,000, 5000, 10,000 or 15,000 genomic regions in any one of Lists 1-8.
Independent claims213