Artificial intelligence-based generation of sequencing metadata
14 claims: 2 independent, 12 dependent
- 1クラスターメタデータ判定タスクのためのニューラルネットワークベースのテンプレート生成器を訓練するために、グラウンドトゥルース訓練データを生成するコンピュータ実装の方法であって、配列決定実行中に生成された一連の画像セットにアクセスすることであって、前記一連の各画像セットは、配列決定実行のそれぞれの配列決定サイクル中に生成され、前記一連の各画像が、クラスター及びそれらの周囲の背景を描き、前記一連の各画像は、ピクセルドメイン内にピクセルを含み、前記ピクセルのそれぞれは、サブピクセルドメイン内の複数のサブピクセルに分割される、アクセスすることと、ベースコーラーから前記サブピクセルの各々を4つの塩基(A、C、T、及びG)のうちの1つと分類するベースコールを取得し、それにより、前記配列決定 実行 の複数の配列決定サイクルにわたって、前記サブピクセルのそれぞれについてベースコールシーケンスを生成することと、実質的に一致するベースコールシーケンスを共有する隣接するサブピクセルの不連続領域として前記クラスターを識別するクラスターマップを生成することと、クラスターマップ内の前記不連続領域に基づいてクラスターメタデータを決定することであって、前記クラスターメタデータが、クラスター中心、クラスター形状、クラスターサイズ、クラスター背景、及び/又はクラスター境界を含む、決定することと、クラスターメタデータを使用して、クラスターメタデータ判定タスクのためのニューラルネットワークベースのテンプレート生成器を訓練するために、グラウンドトゥルース訓練データを生成することであって、グラウンドトゥルース訓練データが、減衰マップ、三元マップ、又はバイナリマップを含み、前記ニューラルネットワークベースのテンプレート生成器が、前記グラウンドトゥルース訓練データに基づいて、前記減衰マップ、前記三元マップ、又は前記バイナリマップを出力として生成するように訓練される、生成することと、を含む、コンピュータ実装の方法。
- 2ハイスループット核酸配列決定技術におけるスループットを増加させるために 、ニ ューラルネットワークベースのテンプレート生成器による出力として生成された、前記減衰マップ、前記三元マップ、又は前記バイナリマップから導き出された前記クラスターメタデータを 、ニューラルネットワークベースのベースコーラーによってベースコールするために 使用すること、を更に含む、請求項1に記載のコンピュータ実装の方法。
- 3前記不連続領域のいずれにも属さないサブピクセルを背景として識別することによって、前記クラスターマップを生成すること、を更に含む、請求項1又は2に記載のコンピュータ実装の方法。
- 4前記クラスターマップが、前記ベースコールシーケンスが実質的に一致しない2つの連続するサブピクセル間のクラスター境界部分を識別する、請求項1から3のいずれか一項に記載のコンピュータ実装の方法。
- 5前記クラスターマップが、ベースコーラーによって決定された前記クラスターの予備中心座標における原点サブピクセルを特定することと、前記原点サブピクセルから開始し、連続的に連続した非原点サブピクセルを継続することによって、実質的に一致するベースコールシーケンスを幅優先探索する、請求項1から4のいずれか一項に記載のコンピュータ実装の方法。
- 6前記クラスターマップの前記不連続領域の質量の中心を、前記不連続領域を形成するそれぞれの連続するサブピクセルの座標の平均として計算することによって、前記クラスターの超位置中心座標を決定することと、前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、メモリ内の前記クラスターの前記超位置中心座標を記憶することと、を更に含む、請求項1から5のいずれか一項に記載のコンピュータ実装の方法。
- 7前記クラスターの前記超位置中心座標における前記クラスターマップの非接合領域内の質量サブピクセルの中心を特定することと、補間を使用して前記クラスターマップをアップサンプリングし、前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、前記メモリ内に前記アップサンプリングされたクラスターマップを記憶することと、アップサンプリングされたクラスターマップでは、隣接するサブピクセルが属する不連続領域内の質量サブピクセルの中心からの隣接するサブピクセルの距離に比例する減衰係数に基づいて、前記不連続領域内の各連続サブピクセルに値を割り当てることと、を更に含む、請求項6に記載のコンピュータ実装の方法。
- 8前記不連続領域内の前記連続するサブピクセルを表し、前記サブピクセルが前記割り当てられた値に基づいて前記背景として特定される、前記アップサンプリングされたクラスターマップから前記減衰マップを生成することと、前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、前記メモリに前記減衰マップを記憶することと、を更に含む、請求項7に記載のコンピュータ実装の方法。
- 9前記アップサンプリングされたクラスターマップにおいて、前記クラスターごとに、前記非接合領域内の前記連続するサブピクセルを、同じクラスターに属するクラスター内部サブピクセルとして分類することと、クラスターの中心サブピクセルとしての質量サブピクセルの中心と、クラスター境界部分を含むサブピクセルと、背景サブピクセルとして前記背景として特定されたサブピクセルとを分類することと、前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、前記メモリに前記分類を記憶することと、を更に含む、請求項8に記載のコンピュータ実装の方法。
- 10前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、 前記クラスターごとに、クラスター内部サブピクセル の座標 、クラスター中心サブピクセル の座標 、境界サブピクセル の座標 、及び背景サブピクセル の座標 をメモリ内 に記 憶することと、前記クラスターマップをアップサンプリングするために使用される因子によって座標をダウンスケールすることと、クラスターごとに、前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、前記メモリに前記ダウンスケールされた座標を記憶することと、を更に含む、請求項1から9のいずれか一項に記載のコンピュータ実装の方法。
- 11フローセルの複数のタイルのクラスターマップを生成することと、前記クラスターマップをメモリに記憶し、前記クラスター中心、前記クラスター形状、前記クラスターサイズ、前記クラスター背景、及び/又は前記クラスター境界を含む、前記クラスターマップに基づいて、前記クラスター内のクラスターのクラスターメタデータを決定することと、前記タイル内の前記クラスターのアップサンプリングされたクラスターマップにおいて、クラスターごとにサブピクセルをクラスターごとに分類することと、同じクラスターに属するクラスター内部サブピクセルとしてのサブピクセル、クラスター中心サブピクセル、境界サブピクセル、及び背景サブピクセルに分類することと、前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、前記メモリに前記分類を記憶することと、 前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、 前記クラスターにクラスターごとに、前記クラスター内部サブピクセルの座標、前記クラスター中心サブピクセル の座標 、前記境界サブピクセル の座標 、及び前記背景サブピクセル の座標を前 記メモリ内 に記 憶することと、前記クラスターマップをアップサンプリングするために使用される因子によって前記座標をダウンスケールすることと、前記タイルにわたるクラスターごとに、前記ニューラルネットワークベースのテンプレート生成器を訓練するための前記グラウンドトゥルース訓練データとして使用するために、前記メモリ内の前記ダウンスケールされた座標を記憶することと、を更に含む、請求項1から10のいずれか一項に記載のコンピュータ実装の方法。
- 12前記ベースコールシーケンスが、ベースコールの所定の部分が、順序位置ごとに一致するときに実質的に一致する、請求項1から11のいずれか一項に記載のコンピュータ実装の方法。
- 13前記クラスターマップが、不連続領域のための所定の最小数のサブピクセルに基づいて生成される、請求項1から12のいずれか一項に記載のコンピュータ実装の方法。
- 14フローセルが、前記クラスターを占有するウェルのアレイを有する少なくとも1つのパターン化表面を有し 、 前記ウェルのうちの どの 1つが、少なくとも1つのクラスターによって実質的に占有され ているか 、前記ウェルのうちの どの 1つが最小限に占有され ているか 、 および 前記 ウェルのうちの どの 1つ が 、複数の集団によって共占有され ているか、 をクラスターの決定された形状及びサイズに基づいて決定することを更に含む、 請求項1から13のいずれか一項に記載のコンピュータ実装の方法。
Independent claims14
868 paragraphs, as filed
(PRIORITY APPLICATION) This application claims priority to or the benefit of the following applications:
U.S. Provisional Patent Application No. 62/821,602, entitled Training Data Generation for Artificial Intelligence-Based Sequencing, filed March 21, 2019 (Attorney Docket No. ILLM1008-1/IP-1693-PRV);
U.S. Provisional Patent Application No. 62/821,618, entitled Artificial Intelligence-Based Generation of Sequencing Metadata, filed March 21, 2019 (Attorney Docket No. ILLM1008-3/IP-1741-PRV);
U.S. Provisional Patent Application No. 62/821,681, entitled Artificial Intelligence-Based Base Calling, filed March 21, 2019 (Attorney Docket No. ILLM1008-4/IP-1744-PRV);
U.S. Provisional Patent Application No. 62/821,724, entitled Artificial Intelligence-Based Quality Scoring, filed March 21, 2019 (Attorney Docket No. ILLM1008-7/IP-1747-PRV);
U.S. Provisional Patent Application No. 62/821,766, entitled Artificial Intelligence-Based Sequencing, filed March 21, 2019 (Attorney Docket No. ILLM1008-9/IP-1752-PRV);
Dutch Patent Application No. 2023310, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing", filed on June 14, 2019 (Attorney Reference No. ILLM1008-11/IP-1693-NL);
Dutch patent application No. 2023311, entitled "Artificial Intelligence-Based Generation of Sequencing Metadata", filed on June 14, 2019 (Attorney Docket No. ILLM1008-12/IP-1741-NL);
Dutch patent application No. 2023312, entitled "Artificial Intelligence-Based Base Calling", filed on June 14, 2019 (Attorney Docket No. ILLM1008-13/IP-1744-NL);
Dutch patent application No. 2023314, entitled "Artificial Intelligence-Based Quality Scoring", filed on June 14, 2019 (Attorney Docket No. ILLM1008-14/IP-1747-NL);
Dutch Patent Application No. 2023316, entitled "Artificial Intelligence-Based Sequencing", filed on June 14, 2019 (Attorney Docket No. ILLM1008-15/IP-1752-NL), and
U.S. patent application Ser. No. 16/825,987, entitled Training Data Generation for Artificial Intelligence-Based Sequencing, filed on March 20, 2020 (Attorney Docket No. ILLM1008-16/IP-1693-US);
U.S. patent application Ser. No. 16/825,991, entitled Training Data Generation for Artificial Intelligence-Based Sequencing, filed on March 20, 2020 (Attorney Docket No. ILLM1008-17/IP-1741-US);
U.S. patent application Ser. No. 16/826,126, entitled Artificial Intelligence-Based Base Calling, filed on March 20, 2020 (Attorney Docket No. ILLM1008-18/IP-1744-US);
U.S. patent application Ser. No. 16/826,134, entitled Artificial Intelligence-Based Quality Scoring, filed on March 20, 2020 (Attorney Docket No. ILLM1008-19/IP-1747-US);
U.S. patent application Ser. No. 16/826,168, entitled Artificial Intelligence-Based Sequencing, filed on March 21, 2020 (Attorney Docket No. ILLM1008-20/IP-1752-PRV);
PCT Patent Application No. PCT__________, entitled "Artificial Intelligence Based Generation of Sequencing Metadata," filed concurrently herewith and subsequently published as PCT International Publication No. WO____________ (Attorney Docket No. ILLM1008-22/IP-1741-PCT);
PCT Patent Application No. PCT___________ entitled "Artificial Intelligence-Based Base Calling," filed concurrently herewith, and subsequently published as PCT International Publication No. WO____________ (Attorney Docket No. ILLM1008-23/IP-1744-PCT);
PCT Patent Application No. PCT__________ entitled Artificial Intelligence-Based Quality Scoring, filed concurrently herewith, and subsequently published as PCT International Publication No. WO____________ (Attorney Docket No. ILLM1008-24/IP-1747-PCT); and
PCT Patent Application No. PCT___________, entitled "Artificial Intelligence-Based Sequencing," filed concurrently herewith, and subsequently published as PCT International Publication No. WO____________ (Attorney Docket No. ILLM1008-25/IP-1752-PCT).
The priority application is incorporated herein by reference for all purposes as if fully set forth herein.
(Built-in)
The following are incorporated by reference for all purposes as if fully set forth herein:
U.S. Provisional Patent Application No. 62/849,091, entitled Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing, filed on May 16, 2019 (Attorney Docket No. ILLM1011-1/IP-1750-PRV);
U.S. Provisional Patent Application No. 62/849,132, entitled Base Calling Using Convolutions, filed May 16, 2019 (Attorney Docket No. ILLM1011-2/IP-1750-PR2);
U.S. Provisional Patent Application No. 62/849,133, entitled Base Calling Using Compact Convolutions, filed on May 16, 2019 (Attorney Docket No. ILLM1011-3/IP-1750-PR3);
U.S. Provisional Patent Application No. 62/979,384, entitled Artificial Intelligence-Based Base Calling of Index Sequences, filed on February 20, 2020 (Attorney Docket No. ILLM1015-1/IP-1857-PRV);
U.S. Provisional Patent Application No. 62/979,414, entitled Artificial Intelligence-Based Many-To-Many Base Calling, filed on February 20, 2020 (Attorney Docket No. ILLM1016-1/IP-1858-PRV);
U.S. Provisional Patent Application No. 62/979,385, entitled Knowledge Distillation-Based Compression of Artificial Intelligence-Based Base Caller, filed on February 20, 2020 (Attorney Docket No. ILLM1017-1/IP-1859-PRV);
U.S. Provisional Patent Application No. 62/979,412, entitled Multi-Cycle Cluster Based Real Time Analysis System, filed on February 20, 2020 (Attorney Docket No. ILLM1020-1/IP-1866-PRV);
U.S. Provisional Patent Application No. 62/979,411, entitled Data Compression for Artificial Intelligence-Based Base Calling, filed on February 20, 2020 (Attorney Docket No. ILLM1029-1/IP-1964-PRV);
U.S. Provisional Patent Application No. 62/979,399, entitled Squeezing Layer for Artificial Intelligence-Based Base Calling, filed on February 20, 2020 (Attorney Docket No. ILLM1030-1/IP-1982-PRV);
Liu P, Hemani A, Paul K, Weis C, Jung M, Wehn N. 3D-Stacked Many-Core Architecture for Biological Sequence Analysis Problems.Int J Parallel Prog.2017, 45(6):1420-60,
Z. Wu, K. Hammad, R. Mittmann, S. Magierowski, E. Ghafar-Zadeh, and X. Zhong, "FPGA-Based DNA Basecalling Hardware Acceleration", in Proc. IEEE 61st Int. Midwest Symp. Circuits Syst. ,Aug.2018, pp.1098-1101,
Z. Wu, K. Hammad, E. Ghafar-Zadeh, and S. Magierowski, "FPGA-Accelerated 3rd Generation DNA Sequencing," in IEEE Transactions on Biomedical Circuits and Systems, Volume 14, Issue 1, Feb. 2020, pp. 65-74,
Prabhakar et al., "Plasticine: A Reconfigurable Architecture for Parallel Patterns", ISCA'17, June 24-28, 2017, Toronto, ON, Canada.
M.Lin,Q.Chen,and S.Yan, Network in Network, in Proc. of ICLR,2014,
L.Sifre, Rigid-motion Scattering for Image Classification,Ph.D.thesis,2014,
L. Sifre and S. Mallat, "Rotation, Scaling and Deformation Invariant Scattering for Texture Discrimination", in Proc. of CVPR, 2013,
F.Chollet, Xception:Deep Learning with Depthwise Separable Convolutions, in Proc. of CVPR,2017,
X. Zhang,
K. He, X. Zhang, S. Ren, and J. Sun, "Deep Residual Learning for Image Recognition", in Proc. of CVPR, 2016,
S. Xie, R. Girshick, P. Dollar, Z. Tu, and K. He, "Aggregated Residual Transformation For Deep NeuroNetworks", Proc. of CVPR, 2017,
AG Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, "Mobilenets: Efficient Convolutional Neural Networks for Mobile Vision Applications", in arXiv:1704.04861, 2017 ,
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L. Chen, "MobileNetV2: Inverted Residuals and Linear Bottlenecks", in arXiv:1801.04381v3, 2018,
Z. Qin, Z. Zhang, X. Chen and Y. Peng, "FD-MobileNet:Improved MobileNet with a Fast Downsampling Strategy", in arXiv:1802.03750,2018,
Liang-Chieh Chen,George Papandreou,Florian Schroff,and Hartwig Adam.Rethinking atrous convolution for semantic image segmentation.CoRR, abs/1706.055887,2017,
J.Huang,V.Rathod,C.Sun,M.Zhu,A.Korattikara,A.Fathi,I.Fischer,Z.Wojna,Y.Song,S.Guadarrama,et al.Speed/accuracy trade-offs for modern convolutional object detectors.arXiv preprint arXiv:1611.10012,2016,
S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, "WAVENET: A GENERATIVE MODEL FOR RAW AUDIO", arXiv:1609.03499, 2016,
SOArik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta and M. Shoeybi, DEEP VOICE: REAL-TIME NEURAL TEXT-TO-SPEECH, arXiv:1702.07825,2017,
F.Yu and V.Koltun, "MULTI-SCALE CONTEXT AGGREGATION BY DILATED CONVOLUTIONS", arXiv:1511.07122,2016,
K. He, X. Zhang, S. Ren, and J. Sun, DEEP RESIDUAL LEARNING FOR IMAGE RECOGNITION, arXiv:1512.03385, 2015,
RKSrivastava, K. Greff, and J. Schmidhuber, HIGHWAY NETWORKS, arXiv:1505.00387, 2015,
G. Huang, Z. Liu, L. van der Maaten and KQWeinberger, "DENTILY CONNECTED CONVOLUTIONAL NETWORKS", arXiv:1608.06993, 2017,
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, GOING DEEPER WITH CONVOLUTIONS, arXiv:1409.4842, 2014,
S. Ioffe and C. Szegedy, BATCH NORMALIZATION:ACCELERATING DEEP NETWORK TRAINING BY REDUCING INTERNAL COVARIATE SHIFT, arXiv:1502.03167,2015,
JMWolterink, T. Leiner, MAViergever, and 1. Isgum, DILATED CONVOLUTIONAL NEURAL NETWORKS FOR CARDIOVASCULAR MR SEGMENTATION IN CONGENITAL HEART DISEASE, arXiv:1704.03669, 2017,
LCPiqueras, AUTOREGRESSIVE MODEL BASED ON A DEEP CONVOLUTIONAL NEURAL NETWORK FOR AUDIO GENERATION, Tampere University of Technology, 2016,
J. Wu, "Introduction to Convolutional Neural Networks", Nanjing University, 2017,
"Illumina CMOS Chip and One-Channel SBS Chemistry", Illumina,Inc.2018,2 pages,
"skikit-image/peak.py at master", GitHub, 5 pages, [Retrieved 2018-11-16]. Retrieved from Internet <URL:https://github.com/scikit-image/scikit-image/blob/master/skimage/feature/peak.py#L25>,
"3.3.9.11. Watershed and random walker for segmentation", Scipy lecture notes, 2 pages, [Retrieved 2018-11-13]. Retrieved from Internet <URL:http://scipy-lectures.org/packages/scikit-image/auto_examples/plot_segmentations.html>,
Mordvintsev, Alexander and Revision, Abid K., "Image Segmentation with Watershed Algorithm", Revision 43532856, 2013, 6 pages [Retrieved 2018-11-13]. Retrieved from Internet: <URL:https://opencv-python-tutroals.readthedocs.io/en/latest/py_tutorials/py_imgproc/py_watershed/py_watershed.html>,
Mzur, "Watershed.py", 25 October 2017, 3 pages, [Retrieved 2018-11-13]. Retrieved from Internet: <URL:https://github.com/mzur/watershed/blob/master/Watershed.py>.
Thakur, Pratibha, et.al. "A Survey of Image Segmentation Techniques", International Journal of Research in Computer Applications and Robotics, Vol. 2, Issue. 4, April 2014, Pg.: 158-165,
Long, Jonathan, et.al., Fully Convolutional Networks for Semantic Segmentation,: IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 39, Issue 4,1 April 2017, 10 pages,
Ronneberger, Olaf, et.al., U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, 18 May 2015, 8 pages,
Xie, W., et.al., Microscopy cell counting and detection with fully convolutional regression networks, Computer methods in biomechanics and biomedical engineering: Imaging & Visualization, 6(3), pp.283-292, 2018.
Xie, Yuanpu, et al., "Beyond classification: structured regression for robust cell detection using convolutional neural network", International Conference on Medical Image Computing and Computer-Assisted Intervention. October 2015, 12 pages,
Snuverink, IAF, "Deep Learning for Pixelwise Classification of Hyperspectral Images", Master of Science Thesis, Delft University of Technology, 23 November 2017, 19 pages,
Shevchenko, A., "Keras weighted categorical_crossentropy", 1 page, [Retrieved 2019-01-15]. Retrieved from Internet: <URL:https://gist.github.com/skeeet/cad06d584548fb45eece1d4e28cfa98b>,
van den Assem, DCF, "Predicting periodic And chaotic signals using Wavenets", Master of Science Thesis, Delft University Of Technology, 18 August 2017, Pages 3-38,
IJ Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio, CONVOLUTIONAL NETWORKS, Deep Learning, MIT Press, 2016, and
J. Gu, Z. Wang, J. Kuen, L. Ma, A. Shahroudy, B. Shuai, T. Liu, X. Wang, and G. Wang, RECENT ADVANCES IN CONVOLUTIONAL NEURAL NETWORKS, arXiv:1512.07108, 2017.
FIELD OF THE DISCLOSURE The technology disclosed relates to artificial intelligence computers and digital data processing systems and corresponding data processing methods and products for emulating intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using deep neural networks, such as deep convolutional neural networks, to analyze data.
The subject matter described in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or related to the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which as such may also correspond to the implementation of the claimed technology.
Deep neural networks are a class of artificial neural networks that use multiple nonlinear and complex transformation layers to model high-level functions in a continuous manner. Deep neural networks provide feedback via backpropagation, which communicates the difference between observed and predicted outputs to adjust parameters. Deep neural networks have evolved with the availability of large training datasets, the power of parallel distributed computing, and advanced training algorithms. Deep neural networks have driven major advances in many domains, such as computer vision, speech recognition, and natural language processing.
Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are components of deep neural networks. Convolutional neural networks have been successful in image recognition, especially in structures that include convolutional layers, nonlinear layers, and pooling layers. Recurrent neural networks are designed to exploit the continuous information of input data with periodic connections between constituent units such as perceptrons, long short-term memory units, and gated recurrent units. In addition, many other emerging deep neural networks have been proposed for limited situations, such as deep spatiotemporal neural networks, multidimensional recurrent neural networks, and convolutional autoencoders.
The goal of training a deep neural network is the optimization of the weight parameters in each layer, which gradually combines simpler features into complex ones so that a better hierarchical representation can be learned from the data. A single cycle of the optimization process consists of: First, given a training dataset, a forward pass sequentially computes the outputs in each layer and propagates the feature signals forward through the network. At the final output layer, an objective loss function measures the error between the inferred output and the given label. To minimize the training error, a backward pass backpropagates the error signal using the chain rule and computes the gradients for all weights in the entire neural network. Finally, the stochastic parameters are updated using an optimization algorithm based on stochastic gradient descent. While batch gradient descent updates parameters for each complete dataset, stochastic gradient descent provides a stochastic approximation by performing updates for each small set of data examples. Several optimization algorithms are derived from stochastic gradient descent, e.g., Adagrad and The Adam training algorithm performs stochastic gradient descent while adaptively modifying the learning rate based on the update frequency of each parameter and the momentum of the gradient, respectively.
Another core element in training deep neural networks is regularization, which refers to a strategy that intends to avoid overfitting and thus achieve good generalization performance. For example, weight decay adds a penalty term to the objective loss function so that the weight parameters converge to smaller absolute values. Dropout randomly removes hidden units from the neural network during training and can be seen as an ensemble of possible sub-networks. To improve the capabilities of dropout, a new activation function, maxout, and a variant of dropout for recurrent neural networks called rnnDrop are proposed. Furthermore, batch normalization provides a new regularization method via scalar feature normalization for each activation in a mini-batch, learning the mean and variance of each as parameters.
Given that sequence data are multi- and high-dimensional, deep neural networks hold considerable promise for bioinformatics research due to their broad applicability and enhanced predictive capabilities. Convolutional neural networks have been employed to solve sequence-based problems in genomics, such as motif discovery, pathogenic variant identification, and gene expression inference. Convolutional neural networks use a weight-sharing strategy that is particularly useful for studying DNA, which can capture short sequence motifs that recapitulate local patterns in DNA that are presumed to have significant biological functions. A notable feature of convolutional neural networks is the use of convolutional filters.
Unlike traditional classification approaches based on carefully designed and manually crafted features, convolutional filters perform adaptive learning of features similar to the process of mapping raw input data to an information representation of knowledge. In this sense, convolutional filters act as a set of motif scanners, since a set of such filters can recognize relevant patterns in the input and update itself during the training procedure. Recurrent neural networks are able to capture long-range dependencies in continuous data of various lengths, such as protein or DNA sequences.
Therefore, an opportunity arises to use a coherent deep learning-based framework for template generation and base calling.
In the era of high-throughput technologies, accumulating the highest yield of interpretable data at the lowest cost per effort remains a significant challenge. Cluster-based methods of nucleic acid sequencing, such as those that utilize bridge amplification for cluster formation, have made a valuable contribution to the goal of increasing the throughput of nucleic acid sequencing. These cluster-based methods rely on sequencing a dense population of nucleic acids immobilized on a solid support and typically involve the use of image analysis software to suppress the optical signal generated during the simultaneous sequencing of multiple clusters located at distinct locations on the solid support.
However, such solid-phase nucleic acid cluster-based sequencing techniques face considerable obstacles that limit the amount of throughput that can be achieved.For example, determining the nucleic acid sequence of two or more clusters that are too physically close to each other to be spatially resolved, or that actually overlap physically on a solid support, can pose obstacles for cluster-based sequencing methods.For example, current image analysis software can require valuable time and computational resources to determine which of two overlapping clusters an optical signal originates from.As a result, compromises are inevitable for various detection platforms with respect to the amount and/or quality of nucleic acid sequence information that can be obtained.
High density nucleic acid aggregate-based genomics methods extend to other areas of genome analysis as well. For example, nucleic acid cluster-based genomics can be used in sequencing applications, diagnostics and screening, gene expression analysis, epigenetic analysis, genetic analysis of polymorphisms, etc. Each of these nucleic acid cluster-based genomics techniques is limited by the inability to resolve data generated from closely adjacent or spatially overlapping nucleic acid clusters.
Clearly, there is a need to improve the quality and quantity of nucleic acid sequence data that can be obtained rapidly and cost-effectively for a variety of applications, including genomics (e.g., for genomic characterization of any and all animal, plant, microbial, or other biological species or populations), pharmacogenomics, transcriptomics, diagnostics, prognosis, biomedical risk assessment, clinical and research genetics, personalized medicine, drug efficacy and drug interaction assessment, veterinary medicine, agricultural, evolutionary, and biological research, aquatic culture, forestry, marine exploration, ecological and environmental management, and other purposes.
The disclosed technology provides neural network-based methods and systems that address these and similar needs, including increasing levels of throughput in high-throughput nucleic acid sequencing technologies, and offers other related advantages.
The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings may also be available in PAIR via the Supplemental Content tab.
In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings, in which:
<figref num="1">1 illustrates one embodiment of a processing pipeline for determining cluster metadata using subpixel base calls.</figref><figref num="2">1 shows one embodiment of a flow cell containing clusters within its tiles.</figref><figref num="3">An example of an Illumina GA-IIx flow cell with eight lanes is shown.</figref><figref num="4">An image set of four channel chemical sequence images is depicted, i.e. the image set has four sequence images captured using four different wavelength bands (image/imaging channels) in the pixel domain.</figref><figref num="5">1 is an embodiment of dividing a sequence image into sub-pixels (or sub-pixel regions).</figref><figref num="6">During subpixel base calling, the preliminary center coordinates of the clusters identified by the base caller are shown.</figref><figref num="7">An example is shown of merging sub-pixel base calls generated over multiple sequencing cycles to generate a so-called "cluster map" that contains cluster metadata.</figref><figref num="8a">1 shows an example of a cluster map generated by merging subpixel base calls.</figref><figref num="8b">1 illustrates one embodiment of a subpixel base call.</figref><figref num="9">13 illustrates another example of a cluster map that identifies cluster metadata.</figref><figref num="10">We show how the center of mass (COM) of discontinuous regions in a cluster map is calculated.</figref><figref num="11">13 illustrates one embodiment of a calculation of a weighted attenuation coefficient based on the Euclidean distance from a subpixel of a discontinuous region to a COM of the discontinuous region.</figref><figref num="12">1 illustrates one implementation of an exemplary ground truth attenuation map derived from an exemplary cluster map generated by sub-pixel base calling.</figref><figref num="13">1 illustrates one embodiment for deriving a ternary map from a cluster map.</figref><figref num="14">1 illustrates one embodiment for deriving a binary map from a cluster map.</figref><figref num="15">FIG. 2 is a block diagram illustrating one embodiment for generating training data used to train the neural network-based template generator and the neural network-based basecaller.</figref><figref num="16">1 illustrates the characteristics of the disclosed training examples used to train the neural network-based template generator and the neural network-based base caller.</figref><figref num="17">1 illustrates one embodiment of processing input image data through the disclosed neural network based template generator to generate output values for each unit in an array. In one embodiment, the array is an attenuation map. In another embodiment, the array is a ternary map. In yet another embodiment, the array is a binary map.</figref><figref num="18">FIG. 1 illustrates one embodiment of a post-processing technique applied to attenuation maps, ternary maps, or binary maps generated by a neural network-based template generator to derive cluster metadata including cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and/or cluster boundaries.</figref><figref num="19">1 illustrates one embodiment for extracting cluster intensities in the pixel domain.</figref><figref num="20">1 illustrates one embodiment for extracting cluster intensities in the sub-pixel domain.</figref><figref num="21a">1 shows three different implementations of a neural network-based template generator.</figref><figref num="21b">15 illustrates one embodiment of input image data provided as input to the neural network-based template generator 1512. The input image data includes a series of image sets having sequence images generated during a particular number of initial sequencing cycles of a sequencing run.</figref><figref num="22">FIG. 21B illustrates one embodiment of extracting patches from the sequence of image sets of FIG. 21b to generate a sequence of "downsized" image sets that form the input image data.</figref><figref num="23">FIG. 21B illustrates one embodiment of upsampling the sequence of image sets of FIG. 21b to generate a sequence of "upsampled" image sets that form the input image data.</figref><figref num="24">FIG. 24 illustrates one embodiment for extracting patches from the series of upsampled image sets of FIG. 23 to generate a series of "upsampled and downsized" image sets that form the input image data.</figref><figref num="25">1 illustrates one embodiment of an overall exemplary process for generating ground truth data for training a neural network-based template generator.</figref><figref num="26">1 illustrates one embodiment of a regression model.</figref><figref num="27">1 illustrates one embodiment of generating a ground truth attenuation map from a cluster map, which is used as ground truth data for training a regression model.</figref><figref num="28">1 is an embodiment of a method for training a regression model using a backpropagation-based gradient update technique.</figref><figref num="29">1 is an embodiment of template generation by a regression model during inference.</figref><figref num="30">1 illustrates one embodiment of post-processing the attenuation map to identify cluster metadata.</figref><figref num="31">1 illustrates one embodiment of a watershed segmentation technique that identifies non-overlapping groups of adjacent clusters/inter-cluster sub-pixels that characterize clusters.</figref><figref num="32">1 is a table showing an exemplary U-Net structure for a regression model.</figref><figref num="33">We present a different approach to extract cluster intensities using cluster shape information identified in a template image.</figref><figref num="34">1 shows different approaches to base calling using the output of a regression model.</figref><figref num="35">Figure 1 shows the difference in base calling performance when the RTA base caller uses ground truth center of mass (COM) positions as cluster centers as opposed to using non-COM positions as cluster centers. The results show that using COMs improves base calling.</figref><figref num="36">On the left, we show an example attenuation map that generated the regression model. On the right, Fig. 36 also shows an example ground truth attenuation map that the regression model approximates during training.</figref><figref num="37">1 illustrates one embodiment of a peak locator that identifies cluster centers in an attenuation map by detecting peaks.</figref><figref num="38">Peaks detected by the peak locator in the attenuation map generated by the regression model are compared to peaks in the corresponding ground truth attenuation map.</figref><figref num="39">Demonstrate the performance of your regression model using precision and recall statistics.</figref><figref num="40">Compare the performance of the RTA base caller and the regression model for a library concentration of 20 pM (normal operation).</figref><figref num="41">Compare the performance of the RTA base caller and the regression model for a library concentration of 30 pM (high density run).</figref><figref num="42">The number of non-overlapping proper read pairs, i.e., the number of pairs of reads where neither read is aligned within a reasonable distance from the others detected by the regression model, is compared to those detected by RTA base calling.</figref><figref num="43">On the right, the first attenuation map generated by the regression model is shown. On the left, Figure 43 shows the second attenuation map generated by the regression model.</figref><figref num="44">Compare the performance of the RTA base caller and the regression model for 40 pM library concentration (high density run).</figref><figref num="45">The first attenuation map generated by the regression model is shown on the left. On the right, Fig. 45 shows the results of thresholding, peak location processing and watershed division techniques applied to the first attenuation map.</figref><figref num="46">1 illustrates one embodiment of a binary classification model.</figref><figref num="47">1 is an implementation of training a binary classification model using a backpropagation based gradient update technique with softmax scoring.</figref><figref num="48">1 is another embodiment of training a binary classification model using a backpropagation based gradient update technique with sigmoid scores.</figref><figref num="49">1 illustrates another embodiment of input image data provided to a binary classification model and corresponding class labels used to train the binary classification model.</figref><figref num="50">1 is one implementation of template generation with a binary classification model during inference.</figref><figref num="51">1 illustrates one embodiment in which the binary map is subjected to peak detection to identify cluster centers.</figref><figref num="52a">An example binary map generated by a binary classification model is shown on the left. Figure 52a also shows an example ground truth binary map to which the binary classification model is proximate during training on the right.</figref><figref num="52b">Use accuracy statistics to indicate the performance of a binary classification model.</figref><figref num="53">1 is a table illustrating an example structure of a binary classification model.</figref><figref num="54">1 illustrates one embodiment of a three-way classification model.</figref><figref num="55">1 is an implementation of training a ternary classification model using a backpropagation-based gradient update technique.</figref><figref num="56">1 illustrates another implementation of input image data provided to a ternary classification model and the corresponding class labels used to train the ternary classification model.</figref><figref num="57">1 is a table illustrating an exemplary structure of a three-way classification model.</figref><figref num="58">1 is an embodiment of template generation with a three-way classification model during inference.</figref><figref num="59">4 shows a ternary map generated by a ternary classification model.</figref><figref num="60">The unit array generated by the ternary classification model 5400 is shown along with the output values for each unit.</figref><figref num="61">We present one embodiment in which the ternary map is subjected to post-processing to identify cluster centers, cluster backgrounds, and cluster interiors.</figref><figref num="62a">1 shows an exemplary prediction of a three-way classification model.</figref><figref num="62b">13 shows another exemplary prediction of a three-way classification model.</figref><figref num="62c">13 illustrates yet another exemplary prediction of a three-way classification model.</figref><figref num="63">FIG. 62b shows one embodiment for deriving cluster centers and shapes from the output of the ternary classification model of FIG. 62a.</figref><figref num="64">Compare the base calling performance of a binary classification model, a regression model, and an RTA base caller.</figref><figref num="65">We compare the performance of the ternary classification model with that of the RTA-based caller under three conditions, five sequence metrics, and two driving densities.</figref><figref num="66">We compare the performance of the regression model with that of the RTA-based caller under three conditions, five sequence metrics, and two driving densities considered in Figure 65.</figref><figref num="67">We focus on the penultimate layer of the neural network-based template generator.</figref><figref num="68">Visualize what the penultimate layer of the neural network-based template generator has learned as a result of backpropagation-based gradient update training. The illustrated embodiment visualizes 24 out of 32 trained convolution filters in the penultimate layer shown in Figure 67.</figref><figref num="69">Cluster center predictions from a binary classification model (in blue) are overlaid on the RTA base calls (in pink).</figref><figref num="70">We overlay cluster center predictions produced by RTA-based color (in pink) on a visualization of the trained convolutional filters in the penultimate layer of a binary classification model (in pink).</figref><figref num="71">1 illustrates one embodiment of training data used to train a neural network-based template generator.</figref><figref num="72">13 is an embodiment of using beads for image registration based on cluster center prediction of a neural network based template generator.</figref><figref num="73">1 illustrates one embodiment of cluster statistics for clusters identified by a neural network-based template generator.</figref><figref num="74">We show how the ability of the neural network-based template generator to distinguish between adjacent clusters improves as the number of initial sequencing cycles for which the input image data is used increases from 5 to 7.</figref><figref num="75">Figure 2 shows the difference in base calling performance when the RTA base caller uses ground truth center of mass (COM) positions as cluster centers as opposed to when non-COM positions are used as cluster centers.</figref><figref num="76">We show the performance of the neural network-based template generator on additional detected clusters.</figref><figref num="77">1 shows different datasets used to train the neural network-based template generator.</figref><figref num="78A">1 illustrates one embodiment of a sequence system, the sequence system including a configurable processor.</figref><figref num="78B">1 illustrates one embodiment of a sequence system, the sequence system including a configurable processor.</figref><figref num="79">FIG. 1 is a simplified block diagram of a system for analysis of sensor data from a sequencing system, such as base call sensor output.</figref><figref num="80">FIG. 1 is a simplified diagram illustrating aspects of a base call operation, including functions of a run-time program executed by a host processor.</figref><figref num="81">FIG. 79 is a simplified diagram of a configuration of a configurable processor such as that shown in FIG.</figref><figref num="82">78B is a computer system that can be used by the sequencing system of FIG. 78A to implement the techniques disclosed herein.</figref>
The following description is presented to enable any person skilled in the art to make and use the disclosed technology, and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
(introduction)
Base calling from digital images is massively parallel and computationally intensive, which presents a number of technical challenges that were identified prior to implementing our novel technique.
The signal from the image set being evaluated becomes increasingly weak as the classification of the base progresses periodically, especially over longer and longer strands of the base. As the classification of the base extends over the length of the strand, the signal-to-noise ratio decreases and the reliability decreases. The updated estimate of reliability is projected as the estimated reliability of the change in the classification of the base.
A digital image is captured from the amplified clusters of sample strands. The sample is amplified by replicating the strands using various physical structures and chemicals. During sequencing by synthesis, the tags are chemically bound in cycles and stimulated to glow. Digital sensors collect photons from the tags which are read out as pixels to generate an image.
Interpreting the digital image to classify the bases requires resolving the positional uncertainty, which is hampered by limited image resolution. At resolutions higher than those collected during base calling, it is clear that the imaged clusters have irregular shapes and uncertain center positions. Cluster positions are not mechanically controlled, so the cluster centers are not aligned with pixel centers. The pixel center can be an integer coordinate assigned to the pixel. In other embodiments, it can be the upper left corner of the pixel. In yet other embodiments, it can be the center of gravity or center of mass of the pixel. Amplification does not produce uniform cluster shapes. Thus, the distribution of cluster signals in the digital image is a statistical distribution rather than a regular pattern. We determine this positional uncertainty.
One of the signal classes does not produce a detectable signal and can be classified to a specific location based on the "dark" signal. Therefore, a template is needed to classify during the dark cycle. Template generation resolves the initial location uncertainty using multiple imaging cycles to avoid missing dark signals.
Tradeoffs in image sensor size, magnification, and stepper design lead to pixel sizes that are too large to process cluster centers to coincide with sensor pixel centers. This disclosure uses pixel in two senses: A physical sensor pixel is an area of a photosensor that reports detected photons. A logical pixel, simply called a pixel, is the data corresponding to at least one physical pixel, read out from the sensor pixel. A pixel may be subdivided or "upsampled" into sub-pixels (e.g., 4x4 sub-pixels). Sub-pixels may be assigned values by interpolation, such as bilinear interpolation or area weighting, to account for the possibility that all photons hit one side of a physical pixel and not the other. Interpolation or bilinear interpolation is also applied when a pixel is reframed by applying an affine transformation to the data from the physical pixel.
Larger physical pixels are more sensitive to weak signals than smaller pixels. Digital sensors improve over time, but the physical limitations of collector surface area are inevitable. Considering the design tradeoffs, legacy systems are designed to collect and analyze image data from a 3x3 patch of sensor pixels, where the center of that cluster is located in the center pixel of the patch.
High-resolution sensors capture only a portion of the imaged medium at a time. The sensor steps above the imaged medium, covering the entire field of view. Thousands of digital images can be collected during one processing cycle.
The sensor and illumination design are combined to distinguish at least four illumination response values used to classify the bases. If a conventional RGB camera with a Bayer color filter array was used, the four sensor pixels would be combined into a single RGB value. This would reduce the effective sensor resolution by a factor of four. Alternatively, it can be collected in a single location using different illumination wavelengths and/or different filters rotated to a location between the imaged medium and the sensor. The number of images required to distinguish between the four base classes varies between systems. Some systems use one image with four intensity levels for different classes of bases. Other systems use two images with different illumination wavelengths (e.g., red and green) and/or filters with a kind of truth table to classify the bases. The system can also use four images with different illumination wavelengths and/or filters tuned to a particular base class.
Highly parallel processing of digital images is in fact necessary to align relatively short strands, on the order of 30-2000 base pairs, into sequences that are much longer in length, potentially millions or even billions in length. Since redundant samples are desirable on the imaged medium, parts of the sequence may be covered by many sample reads. Millions or at least hundreds of thousands of sample clusters are imaged from a single imaged medium. Large-scale processing of many such clusters increases sequencing capacity while decreasing costs.
Sequencing capacity is increasing at a pace that replicates Moore's law. First sequencing costs a billion dollars, but 2018 services such as Illumina provide results for hundreds of dollars. As sequencing becomes mainstream and unit costs fall, less computing power is available for classification, which increases the challenge of near real-time classification. With these technical challenges in mind, the inventors turn to the disclosed technology.
The disclosed techniques improve both the process during template generation to resolve position uncertainty and during base classification of clusters at resolved positions. Applying the disclosed techniques can use cheaper hardware to reduce machine costs. Near real-time analysis can be cost-effective and reduce the delay between image collection and base classification.
The disclosed technique can use upsampled images generated by interpolating sensor pixels to subpixels, then generate templates that resolve positional uncertainties. The resulting subpixels are submitted to a base caller for classification, which treats the subpixel as if it were the center of a cluster. Clusters are identified from groups of adjacent subpixels that repeatedly receive the same base classification. This aspect of the technique can leverage existing base calling techniques to identify cluster shapes and super-search for cluster centers at subpixel resolution.
Another aspect of the disclosed technology is to create ground truth, pairing images with reliable identified cluster centers and/or cluster shapes. Deep learning systems and other machine learning approaches require substantial training sets. Human-curated data is expensive to compile. The disclosed technology can be used to generate large sets of sensitively classified training data in non-standard modes of operation without the intervention or expense of a human curator. The training data correlates raw images with cluster centers and/or cluster shapes available from existing classifiers in non-standard modes of operation, such as CNN-based deep learning systems. One training image can be rotated and reflected to generate additional equally valid examples. Training examples can be focused on regions of a given size within the entire image. The context evaluated during base calling determines the size of the example training regions, not the size of the image or the entire imaged medium.
The disclosed technique can generate different kinds of maps that can be used as training data or as templates for base classification, which correlate cluster centers and/or cluster shapes with the digital image. First, sub-pixels can be classified as cluster centers, thereby localizing the cluster centers within the physical sensor pixel. Second, the cluster center can be calculated as the centroid of the cluster shape. This location can be reported to a selected numerical precision. Third, the cluster center can be reported at the surrounding sub-pixels in an attenuation map, either at sub-pixel or pixel resolution. The attenuation map reduces the weight given to photons detected within a region as the separation of the region from the cluster center increases, attenuating signals from more distant locations. Fourth, a binary or ternary classification can be applied to sub-pixels or pixels within a cluster of neighboring regions. In a binary classification, a region is classified as belonging to a cluster center or as background. In a ternary classification, a third class type is assigned to regions that include cluster interiors but are not cluster centers. The sub-pixel classification of cluster center locations can be substituted for real-valued cluster center coordinates within a larger optical pixel.
Alternative map styles can be generated initially as a ground truth data set or can be generated using a neural network with training. For example, clusters can be depicted as discontinuous regions of adjacent sub-pixels with appropriate classification. The intensities of the mapped clusters from the neural network can be post-processed by a peak detector filter to calculate cluster centers if the centers have not already been determined. Adjacent regions can be assigned to separate clusters by applying a so-called watershed analysis. When generated by a neural network inference engine, the map can be used as a template to evaluate sequences of digital images and classify bases over cycles of base calling.
(Neural network based template generation)
The first step in template generation is to identify cluster metadata, which identifies the spatial distribution of clusters, including their centers, shapes, sizes, backgrounds, and/or boundaries.
(Identifying cluster metadata)
FIG. 1 illustrates one embodiment of a processing pipeline for identifying cluster metadata using subpixel base calls.
Figure 2 shows one embodiment of a flow cell containing clusters within its tiles. The flow cell is divided into lanes. The lanes are further divided into non-overlapping regions called "tiles." During the sequencing procedure, the populations on the tiles and their surrounding background are imaged.
Figure 3 shows an exemplary Illumina GA-IIx flow cell with eight lanes. Figure 3 also shows a close-up of one tile and its clusters and their surrounding background.
FIG. 4 depicts an image set of a four-channel chemical sequence image, i.e., the image set has four sequence images captured using four different wavelength bands (image/imaging channels) in the pixel domain. Each image in the image set covers a tile of a flow cell and shows the intensity emission of clusters on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. In one embodiment, each imaging channel corresponds to one of a plurality of filter wavelength bands. In another embodiment, each imaging channel corresponds to one of a plurality of imaging events in a sequencing cycle. In yet another embodiment, each imaging channel corresponds to a combination of illumination with a particular laser and imaging through a particular optical filter. The intensity emission of the clusters comprises a signal detected from the analyte that can be used to classify a base associated with the analyte. For example, the intensity emission can be a signal indicative of photons emitted by a tag chemically attached to the analyte during a cycle in which the tag is stimulated and can be detected by one or more digital sensors.
FIG. 5 is an embodiment of dividing a sequence image into subpixels (or subpixel regions). In another embodiment shown, quarter (0.25) subpixels are used, whereby each pixel in the sequence image is divided into 16 subpixels. Assuming that the illustrated sequence image has a resolution of 20×20 pixels, i.e., 400 pixels, the division produces 6400 subpixels. Each of the subpixels is treated by a base caller as a region center for subpixel base calling. In some embodiments, the base caller does not use neural network-based processing. In other embodiments, the base caller is a neural network-based base caller.
For a given sequencing cycle and a particular subpixel, the base caller is configured with logic to perform image processing steps to generate base calls for the particular subpixel for the given sequencing cycle by extracting the intensity data of the subpixel from the corresponding image set of the sequencing cycle. This is done for each of the subpixels and for each of the multiple sequencing cycles. Experiments were also performed using a 1/4 subpixel division of an Illumina MiSeq sequencer's 1800x1800 pixel resolution tile image. Subpixel base calls were performed for 50 sequencing cycles and 10 tile lanes.
Figure 6 shows the preliminary center coordinates of clusters identified by the base caller during subpixel base calling. Figure 6 also shows the "origin subpixel" or "center subpixel" that contains the preliminary center coordinates.
7 shows an example of merging sub-pixel base calls generated over multiple sequencing cycles to generate a so-called "cluster map" that contains cluster metadata. In the illustrated embodiment, the sub-pixel base calls are merged using a breadth-first search approach.
Figure 8a shows an example of a cluster map generated by merging subpixel base calls. Figure 8b shows an example of a subpixel base call. Figure 8b also shows an embodiment in which the cluster map is generated by analyzing the base call sequences for each subpixel generated from the subpixel base calls.
(Sequencing image)
Cluster metadata determination involves analyzing image data generated by a sequencing device 102 (e.g., Illumina's iSeq, HiSeqX, HiSeq3000, HiSeq4000, HiSeq2500, NovaSeq 6000, NextSeq, NextSeqDx, MiSeq, and MiSeqDx). The following description outlines how image data is generated and what it describes, according to one embodiment.
Base calling is the process by which the raw signal of the sequencing instrument 102, i.e., the intensity data extracted from the image, is decoded into DNA sequence and quality scores. In one embodiment, the Illumina platform employs cyclic reversible termination (CRT) chemistry for base calling. This process relies on growing emergent DNA strands complementary to the template DNA strand with modified nucleotides, tracking the emission signal of each newly added nucleotide. The modified nucleotides have a 3' removable block that anchors the fluorophore signal of the nucleotide type.
Sequencing is performed in repeated cycles, each of which includes three steps: (a) extending the transstrand by adding modified nucleotides; (b) exciting the fluorophore using one or more lasers of the optical system 104 and imaging through different filters of the optical system 104 to generate a sequence image 108; and (c) cleaving the fluorophore and removing the 3' block in preparation for the next sequencing cycle. The incorporation and imaging cycle is repeated for a specified number of sequencing cycles to define the read length of the entire population. Using this approach, each cycle queries a new position along the template strand.
The trément power of the Illumina platform stems from its ability to simultaneously run and sense millions or even billions of clusters undergoing CRT reactions. The sequencing process is carried out in a flow cell 202, a small glass slide that holds the input DNA fragments during the sequencing process. The flow cell 202 is connected to a high-throughput optical system 104 that includes a microscope image, an excitation laser, and a fluorescence filter. The flow cell 202 contains multiple chambers called lanes 204. The lanes 204 are physically separated from each other and may contain different tagged sequencing libraries, which are distinguishable without sample cross-contamination. An imaging device 106 (e.g., a solid-state imaging device such as a charge-coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) sensor) takes snapshots at multiple locations along the lane 204 in a series of non-overlapping regions called tiles 206.
For example, there are 100 tiles per lane on an Illumina Genome Analyzer II and 68 tiles per lane in an Illumina HiSeq2000. A tile 206 holds hundreds of thousands to millions of clusters. An image generated from a tile with clusters shown as bright spots is shown at 208. A cluster 302 contains about a thousand identical copies of a template molecule, but the clusters differ in size and shape. Clusters are grown from the template molecules by bridge amplification of the input library before sequencing runs. The purpose of the amplification and cluster growth is to increase the intensity of the emitted signal, since the imaging device 106 cannot reliably sense a single fluorophore. However, because the physical distance of the DNA fragments in a cluster 302 is small, the imaging device 106 perceives the cluster of fragments as a single spot 302.
The output of the sequencing operation is a sequence image 108 that shows the intensity emission of clusters on a tile in the pixel domain, each for a particular combination of lane, tile, sequencing cycle, and fluorophore (208A, 208C, 208T, 208G).
In one embodiment, the biosensor comprises an array of optical sensors. The optical sensors are configured to sense information from corresponding pixel regions (e.g., reaction sites/wells/nanocells) on the detection surface of the biosensor. Analytes disposed within pixel regions are said to be associated with the pixel regions, i.e., associated analytes. In a sequencing cycle, the optical sensors corresponding to the pixel regions are configured to detect/capture/sense luminescence/photons from the associated analytes and accordingly generate pixel signals for each imaged channel. In one embodiment, each imaging channel corresponds to one of a plurality of filter wavelength bands. In another embodiment, each imaging channel corresponds to one of a plurality of imaging events in a sequencing cycle. In yet another embodiment, each imaging channel corresponds to a combination of illumination with a particular laser and imaging through a particular optical filter.
The pixel signals from the photosensors are communicated (e.g., via a communications port) to a signal processor coupled to the biosensor. For each sequencing cycle and each imaging channel, the signal processor generates an image in which the pixels respectively depict/contain/show/represent/characterize the pixel signal obtained from the corresponding photosensor. In this manner, a pixel in the image corresponds to (i) the photosensor of the biosensor that generated the pixel signal represented by the pixel, (ii) the relevant analyte whose radiation was detected by the corresponding photosensor and converted into a pixel signal, and (iii) the pixel area on the detection surface of the biosensor that holds the relevant analyte.
For example, consider a sequencing operation uses two different imaging channels: a red channel and a green channel. Then, in each sequencing cycle, the signal processor generates a red image and a green image. In this way, for a series of k sequencing cycles of a sequencing run, a sequence having k pairs of red and green images is generated as output.
Pixels in red and green images (i.e., different imaging channels) have a one-to-one correspondence within a sequencing cycle. This means that corresponding pixels in a pair of red and green images show intensity data of the same associated analyte in different imaging channels. Similarly, pixels across a pair of red and green images have a one-to-one correspondence between sequencing cycles. This means that corresponding pixels in different pairs of red and green images show intensity data of the same associated analyte for different acquisition events/time steps (sequencing cycles) of a sequencing run.
Corresponding pixels in the red and green images (i.e., different imaging channels) can be considered as pixels of an "image per cycle" that represent intensity data in a first red channel and a second green channel. A cycle-by-cycle image whose pixels depict pixel signals of a subset of the pixel area, i.e., an area (tile) of the sensing surface of the biosensor, is called a "tile-by-cycle image." A patch extracted from a cycle-by-cycle tile image is called a "image-by-cycle patch." In one embodiment, patch extraction is performed by an input preparer.
The image data includes a series of per-cycle image patches generated for a series of k-sequence cycles of a sequencing run. Pixels in the per-cycle image patch include intensity data for an associated analyte, the intensity data being acquired for one or more imaging channels (e.g., red and green channels) by corresponding photosensors configured to detect emissions from the associated analytes. In one embodiment, when based on a single target cluster, the per-cycle image patch is centered on a central pixel that includes intensity data for the target-associated analyte and non-central pixels, the non-central pixels in the per-cycle image patch include intensity data for associated analytes adjacent to the target-associated analyte. In one embodiment, the image data is prepared by an input preparer.
(Subpixel base call)
The disclosed technique accesses a series of image sets generated during a sequencing run. The image sets include sequence images 108. Each successive image set is captured during each sequencing cycle of a sequencing run. The series of images (or sequence images) captures the clusters on the tiles of the flow cell and their surrounding background.
In one embodiment, the sequencing run utilizes four channel chemistry and each image set has four images. In another embodiment, the sequencing run utilizes two channel chemistry and each image set has two images. In yet another embodiment, the sequencing run utilizes one channel chemistry and each image set has two images. In yet another embodiment, each image set has only one image.
The sequence image 108 in the pixel domain is first converted to the sub-pixel domain by the sub-pixel addresser 110 to generate the sequence image 112 in the sub-pixel domain. In one embodiment, each pixel in the sequence image 108 is divided into 16 sub-pixels 502. Thus, in one embodiment, the sub-pixels 502 are quarter sub-pixels. In another embodiment, the sub-pixels 502 are half sub-pixels. As a result, each of the sequence images 112 in the sub-pixel domain has a number of sub-pixels 502.
The sub-pixels are then separately fed as inputs to the base caller 114 to obtain base calls from the base caller 114 that classify each of the sub-pixels as one of four bases (A, C, T, and G). This generates a base call sequence 116 for each of the sub-pixels across multiple sequencing cycles of a sequencing run. In one embodiment, the sub-pixels 502 are identified to the base caller 114 based on their integer or non-integer coordinates. By tracking the emission signals from the sub-pixels 502 across a set of images generated during multiple sequencing cycles, the base caller 114 recovers the underlying DNA sequence of each sub-pixel. An example of this is shown in FIG. 8b.
In other embodiments, the disclosed technology classifies each of the subpixels as one of five bases (A, C, T, G, and N) from the base caller 114. In such embodiments, the N base calls represent undetermined base calls, typically resulting from low levels of extracted intensity.
Some examples of base callers 114 include non-neural network based Illumina offerings such as Real Time Analysis (RTA), the Firecrest program of the Genome Analyzer Analysis Pipeline, the Integrated Primary Analysis and Reporting (Ipar) machine, and Off-Line Basecaller (OLB). For example, the base caller 114 generates base call sequences by interpolating sub-pixel intensities, including at least one of the following: nearest neighbor intensity extraction, Gaussian-based intensity extraction, intensity extraction based on average 2x2 sub-pixel area, intensity extraction based on brightest test of 2x2 sub-pixel area, intensity extraction based on average 3x3 sub-pixel area, bilinear intensity extraction, bicubic intensity extraction, and/or intensity extraction based on weighted area coverage. These techniques are described in detail in the Appendix entitled "Intensity Extraction Methods".
In other embodiments, the base caller 114 can be a neural network-based base caller, such as the neural network-based base caller 1514 disclosed herein.
The base call sequences 116 for each subpixel are then provided as input to a searcher 118, which searches for substantially matching base call sequences of consecutive subpixels. The base call sequences of consecutive subpixels "substantially match" when a predetermined portion of the base calls match a criterion for each ordinal position (e.g., 41 matches in >=45 cycles, 4 mismatches in <=45 cycles, 4 mismatches in <=50 cycles, or 2 mismatches in <=34 cycles).
The searcher 118 then generates a cluster map 802 that identifies clusters, such as 804a-d, of adjacent subpixels that share substantially matching base call sequences. This application uses "disjoint," "disjoint," and "non-overlapping" interchangeably. The search involves calling subpixels that include part of a cluster and allowing them to link the called subpixels to adjacent subpixels that share substantially matching base call sequences. In some implementations, the searcher 118 requires that at least some of the discontinuous regions have a predetermined minimum number of subpixels (e.g., more than 4, 6, or 10 subpixels) to be treated as a cluster.
In some implementations, the base caller 114 also identifies preliminary center coordinates of the cluster. The subpixel that contains the preliminary center coordinate is referred to as the origin subpixel. Some example preliminary center coordinates (604a-c) identified by the base caller 114 and corresponding origin subpixels (606a-c) are shown in FIG. 6. However, as described below, the identification of the origin subpixel (preliminary center coordinate of the cluster) is not required. In some implementations, the searcher 118 uses a breadth-first search starting from the origin subpixels 606a-c and continuing through successive consecutive non-origin subpixels 702a-c to identify substantially matching base call sequences of subpixels. This is optional, as described below.
(Cluster map)
FIG. 8a shows an example of a cluster map 802 generated by merging subpixel base calls. The cluster map identifies multiple discontinuous regions (indicated by different colors in FIG. 8a). Each discontinuous region includes non-overlapping groups of contiguous subpixels (from which sequence images and from which the cluster map is generated via subpixel base calls) that represent a respective cluster on the tile. The regions between the discontinuous regions represent the background on the tile. Subpixels within the background regions are referred to as "background subpixels." Subpixels within the discontinuous regions are referred to as "cluster subpixels" or "interior cluster subpixels." In this description, the origin subpixel is the subpixel in which the preliminary center cluster coordinates, as determined by the RTA or another base caller, are located.
The origin subpixel contains the preliminary center cluster coordinate, meaning that the area covered by the origin subpixel contains a coordinate location that coincides with the preliminary center cluster coordinate location. Because the cluster map 802 is an image of logical subpixels, the origin subpixel is a portion of the subpixels in the cluster map.
The search to identify clusters having substantially matching base call sequences of subpixels does not need to begin with identifying an origin subpixel (a preliminary center coordinate of the cluster) because the search can be performed for all subpixels and can begin at any subpixel (e.g., the 0,0 subpixel or any random subpixel). Thus, the search can begin at any subpixel because the search is not dependent on the origin subpixel, as each subpixel is evaluated to determine whether it shares a substantially matching base call sequence with another contiguous subpixel.
Regardless of whether the origin subpixel is used, certain clusters that do not include the origin subpixel (initial center coordinate of the cluster) predicted by the base caller 114 are identified. Some examples of clusters that are identified by merging of subpixel base calls and do not include the origin subpixel are clusters 812a, 812b, 812c, 812d, and 812e in FIG. 8a. Thus, the disclosed techniques identify additional or extra clusters whose centers may not have been identified by the base caller 114. Thus, the use of the base caller 114 to identify the origin subpixel (initial center coordinate of the cluster) is optional and not required to search for substantially matching base call sequences of consecutive subpixels.
In one embodiment, the origin subpixels (initial center coordinates of the clusters) identified by the base caller 114 are first used to identify a first set of clusters (by identifying substantially matching base call sequences of consecutive subpixels). Subpixels that are not part of the first set of clusters are then used to identify a second set of clusters (by identifying substantially matching base call sequences of consecutive subpixels). This enables the disclosed techniques to identify additional or extra clusters whose centers are not identified by the base caller 114. Finally, subpixels that are not part of the first and second sets of clusters are identified as background subpixels.
An example of sub-pixel base calling is shown in Figure 8b, where each sequencing cycle has an image set with four different images (i.e., A, C, T, G images) captured using four different wavelength bands (images/imaging channels) and four different fluorescent dyes (one for each base).
In this example, a pixel in an image is divided into 16 subpixels. The subpixels are then base called separately by the base caller 114 in each sequencing cycle. To base call a given subpixel in a particular sequencing cycle, the base caller 114 uses the intensity of the given subpixel in each of the four A, C, T, G images. For example, the intensity of the image area covered by subpixel 1 in each of the four A, C, T, G images in cycle 1 is used to base call subpixel 1 in cycle 1. For subpixel 1, these image areas include the top left 1/16 area of the top left pixel in each of the four A, C, T, G images in cycle 1. Similarly, the intensity of the image area covered by subpixel m in each of the four A, C, T, G images in cycle n is used to determine subpixel m in cycle n. For subpixel m, these image areas include the bottom right 1/16 area of the bottom right pixel in each of the four A, C, T, G images in cycle 1.
This process generates base call sequences 116 for each subpixel over multiple sequencing cycles. The searcher 118 then evaluates pairs of consecutive subpixels to determine whether they have substantially matching base call sequences. If yes, the pair of subpixels is stored in the cluster map 802 as belonging to the same cluster in a discontinuous region. If no, the pair of subpixels is stored in the cluster map 802 as not belonging to the same discontinuous region. Thus, the cluster map 802 identifies consecutive sets of subpixels where the base calls for the subpixels substantially match over multiple cycles. The cluster map 802 thus uses information from the multiple clusters to provide a plurality of clusters, where each cluster of the multiple clusters has a high degree of confidence in providing sequence data for a single DNA strand.
The cluster metadata generator 122 then processes the cluster map 802 to determine cluster metadata, including determining the spatial distribution of the clusters, including their centers (810a), shapes, sizes, backgrounds, and/or boundaries (FIG. 9).
In some implementations, the cluster metadata generator 122 identifies subpixels in the cluster map 802 as background that do not belong to any of the disjoint regions and therefore do not contribute to any clusters. Such subpixels are referred to as background subpixels 806a-c.
In some embodiments, the cluster map 802 identifies cluster boundaries 808a-c between two consecutive subpixels where the base call sequences do not substantially match.
The cluster map is stored in memory (e.g., cluster map data store 120) for use as ground truth for training classifiers such as neural network-based template generator 1512 and neural network-based base caller 1514. Cluster metadata may also be stored in memory (e.g., cluster metadata data store 124).
FIG. 9 illustrates another example of a cluster map that identifies cluster metadata including the spatial distribution of clusters, along with cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and/or cluster boundaries.
(Center of Mass (COM))
Figure 10 shows how the center of mass (COM) of a discontinuous region in a cluster map is calculated. The COM can be used as the "corrected" or "improved" center of the corresponding cluster in downstream processing.
In some implementations, for each cluster, a center of mass calculator 1004 determines the cluster's hyperlocation center coordinates 1006 by calculating the center of mass of the discontinuous regions of the cluster map as the average of the coordinates of each contiguous subpixel that forms the discontinuous region. The hyperlocation center coordinates of the cluster are then stored for each cluster in memory for use as ground truth for training the classifier.
In some implementations, the subpixel classifier identifies, for each cluster, a center of mass subpixel 1008 within the discontinuous regions 804a-d of the cluster map 802 at the superlocation center coordinates 1006 of the cluster.
In another alternative embodiment, the cluster map is upsampled using interpolation, and the upsampled cluster map is stored in memory for use as ground truth for training the classifier.
(Attenuation coefficient and attenuation map)
FIG. 11 illustrates one embodiment of a calculation of a weighted attenuation coefficient for a subpixel based on the Euclidean distance from the subpixel to the center of mass (COM) of the discontinuous region to which the subpixel belongs. In another illustrated embodiment, the weighted attenuation coefficient is given the highest value to the subpixel that includes the COM and decreases for subpixels further away from the COM. The weighted attenuation coefficient is used to derive a ground truth attenuation map 1204 from the cluster map generated from the subpixel base calls described above. The ground truth attenuation map 1204 includes a unit array and assigns at least one output value to each unit in the array. In some implementations, the units are subpixels and each subpixel is assigned an output value based on the weighted attenuation coefficient. The ground truth attenuation map 1204 is then used as ground truth for training the disclosed neural network based template generator 1512. In some implementations, information from the ground truth attenuation map 1204 is also used to prepare the input of the disclosed neural network based base caller 1514.
12 illustrates one implementation of an exemplary ground truth attenuation map 1204 derived from an exemplary cluster map generated by subpixel base calling as described above. In some implementations, in the cluster-by-cluster upsampled cluster map, for each cluster, a value is assigned to each contiguous subpixel in a discontinuous region based on an attenuation coefficient 1102 that is proportional to the distance 1106 of the contiguous subpixel from the center of mass subpixel 1104 in the discontinuous region to which the adjacent subpixel belongs.
12 illustrates a ground truth attenuation map 1204. In one implementation, the subpixel values are normalized intensity values between zero and one. In another implementation, all subpixels identified as background in the upsampled cluster map are assigned the same predefined value. In some implementations, the predefined value is a zero intensity value.
In some implementations, the ground truth attenuation map 1204 is generated by the ground truth attenuation map generator 1202 from the upsampled cluster map representing contiguous subpixels in discontinuous regions based on their assigned values. The ground truth attenuation map 1204 is stored in memory for use as ground truth to train the classifier. In one implementation, each subpixel in the ground truth attenuation map 1204 has a normalized value between zero and one.
(Triple (3-class) map)
FIG. 13 illustrates one embodiment of deriving a ground truth ternary map 1304 from a cluster map. The ground truth ternary map 1304 includes a unit array and assigns at least one output value to each unit in the array. By name, the ternary map embodiment of the ground truth ternary map 1304 assigns three output values to each unit in the array such that for each unit, the first output value corresponds to the classification label or score of the background class, the second output value corresponds to the classification label or score of the cluster center class, and the third output value corresponds to the classification label or score of the cluster/cluster inner class. The ground truth ternary map 1304 is used as ground truth data for training the neural network-based template generator 1512. In some embodiments, information from the ground truth ternary map 1304 is also used to prepare the input of the neural network-based base caller 1514.
13 illustrates an exemplary ground truth ternary map 1304. In another implementation, in the upsampled cluster map, contiguous subpixels in discontinuous regions are classified by cluster by the ground truth ternary map generator 1302 as cluster interior subpixels that belong to the same cluster, center of mass subpixels as cluster center subpixels, and background subpixels as subpixels that do not belong to any cluster. In some implementations, the classifications are stored in a ground truth ternary map 1304. These classifications and the ground truth ternary map 1304 are stored in memory for use as ground truth for training a classifier.
In another alternative embodiment, for each cluster, the coordinates of the cluster interior subpixel, the cluster center subpixel, and the background subpixel are stored in memory for use as ground truth for training the classifier. The coordinates are then downscaled by the factor used to upsample the cluster map. The downscaled coordinates for each cluster are then stored in memory for use as ground truth for training the classifier.
In yet another embodiment, the ground truth ternary map generator 1302 uses the cluster map to generate ternary ground truth data 1304 from the upsampled cluster map. The ternary ground truth data 1304 labels background subpixels as belonging to a background class, cluster center subpixels as belonging to a cluster center class, and cluster interior subpixels as belonging to a cluster interior class. In some visualization embodiments, color coding can be used to depict and distinguish the different class labels. The ternary ground truth data 1304 is stored in memory for use as ground truth to train a classifier.
(Binary (2-class) map)
FIG. 14 illustrates one embodiment of deriving a ground truth binary map 1404 from a cluster map. The binary map 1404 includes an array of units and assigns at least one output value to each unit in the array. By name, the binary map assigns two output values to each unit in the array such that for each unit, the first output value corresponds to the classification label or score of the cluster center class and the second output value corresponds to the classification label or score of the non-center class. The binary map is used as ground truth data for training the neural network-based template generator 1512. In some embodiments, information from the binary map is also used to prepare input for the neural network-based base caller 1514.
14 illustrates a ground truth binary map 1404. A ground truth binary map generator 1402 uses the cluster map 120 to generate binary ground truth data 1404 from the upsampled cluster map. The binary ground truth data 1404 labels cluster center subpixels as belonging to a cluster center class and labels all other subpixels as belonging to a non-center class. The binary ground truth data 1404 is stored in memory for use as ground truth to train a classifier.
In some implementations, the disclosed techniques generate cluster maps 120 for multiple tiles of a flow cell, store the cluster maps in memory, and determine the spatial distribution of clusters within the tile based on the cluster maps 120 including their shapes and sizes. The disclosed techniques then classify the subpixels for each cluster in the upsampled cluster map 120 of clusters within the tile into cluster interior subpixels, cluster center subpixels, and background subpixels that belong to the same cluster. The disclosed techniques then store the classification in memory for use as ground truth for training a classifier, and for each cluster in the cluster, store the coordinates of the cluster interior subpixels, cluster center subpixels, and background subpixels in memory for use as ground truth for training a classifier. The disclosed techniques then downscale the coordinates by the factor used to upsample the cluster map, and store the downscaled coordinates for each cluster in memory for use as ground truth for training a classifier.
In some embodiments, the flow cell has at least one patterned surface with an array of wells that occupy clusters. In such embodiments, based on the determined shapes and sizes of the clusters, the disclosed technology identifies (1) which one of the wells is substantially occupied by at least one group, (2) which one of the wells is minimally occupied, and (3) which one of the wells is co-occupied by multiple groups. This allows for the determination of metadata for each of multiple clusters that co-occupy the same well, i.e., the center, shape, and size of two or more clusters that share the same well.
In some embodiments, the solid support on which the sample is amplified into clusters comprises a patterned surface. A "patterned surface" refers to an arrangement of different regions in or on an exposed layer of a solid support. For example, one or more regions can be features in which one or more amplification primers are present. The features can be separated by interstitial regions in which no amplification primers are present. In some embodiments, the pattern can be an xy format of features in rows and columns. In some embodiments, the pattern can be a repeating sequence of features and/or interstitial regions. In some embodiments, the pattern can be a random sequence of features and/or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Pat. No. 8,778,849, U.S. Pat. No. 9,079,148, U.S. Pat. No. 8,778,848, and U.S. Patent Application Publication No. 2014/0243224, each of which is incorporated herein by reference.
In some embodiments, the solid support comprises an array of wells or depressions on its surface, which can be fabricated as commonly known in the art using a variety of techniques, including but not limited to photolithography, stamping techniques, molding techniques, and microetching techniques. As will be understood in the art, the technique used will depend on the composition and shape of the array substrate.
The features within the patterned surface may be wells in an array of wells (e.g., microwells or nanowells) on glass, silicon, plastic, or other suitable solid support with a patterned covalently attached gel, such as poly(N-(5-azidoacetamylpentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication Nos. 2013/184796, WO 2016/066586, and 2015-002813, each of which is incorporated by reference in its entirety). This process creates a gel pad used for sequencing, which can be stable over a large number of cycles and across a sequencing run. Covalently attaching a polymer to the wells is useful to maintain the gel in the structured features throughout the life of the structured substrate during various applications. However, in many embodiments, the gel does not need to be covalently attached to the wells. For example, in some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Pat. No. 8,563,477, which is incorporated by reference in its entirety) that is not covalently bonded to any portion of the structured substrate can be used as the gel material.
In certain other embodiments, a structured substrate can be made by patterning a solid support material with wells (e.g., microwells or nanocells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or chemically modified variants thereof), such as an azido version of SFA (azido-SFA), and polishing the gel-coated support, for example by chemical or mechanical polishing, thereby retaining the gel in the wells but removing or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the wells. Primer nucleic acids can then be attached to the gel material. A solution of target nucleic acids (e.g., a fragmented human genome) can then be contacted with the polished substrate such that individual target nucleic acids seed individual wells through interaction with primers bound to the gel material, but the target nucleic acids do not occupy the intervening regions due to inactivity or inactivity of the gel material. Amplification of the target nucleic acids will be confined to the wells because the absence or inactivity of the gel in the intervening regions prevents outward migration of growing nucleic acid colonies. The process is conveniently manufacturable and scalable, utilizing micro- or nano-fabrication methods.
As used herein, the term "flow cell" refers to a chamber that includes a solid surface through which one or more fluidic reagents can flow. Examples of flow cells and related fluidic systems and detection platforms that can be easily used in the method of the present disclosure are described in, for example, Bentley et al., Nature 456:53-59 (2008), WO 04/018497, U.S. Patent No. 7,057,026, WO 91/06678, U.S. Patent No. 07/123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281, and U.S. Patent No. 2008/0108082, each of which is incorporated herein by reference.
Throughout this disclosure, the terms "P5" and "P7" are used when referring to amplification primers. It will be understood that any suitable amplification primers can be used in the methods presented herein, and the use of P5 and P7 is only an exemplary implementation. The use of amplification primers such as P5 and P7 on flow cells is known in the art, as exemplified by the disclosures of WO 2007/010251, WO 2006/064199, WO 2005/065814, WO 2015/106941, WO 1998/044151 and WO 2000/018957, which are incorporated herein by reference in their entirety. For example, any suitable forward amplification primer can be useful in the methods presented herein for amplification of complementary sequences and sequences, whether immobilized or in solution. Similarly, any suitable reverse amplification primer can be useful in the methods provided herein for complementary sequences and amplification of sequences, whether immobilized or in solution. One of skill in the art will know how to design and use suitable primer sequences for capture and amplification of nucleic acids as provided herein.
In some embodiments, the flow cell has at least one non-patterned surface, and the clusters are scattered non-uniformly on the non-patterned surface.
In some embodiments, the density of the clusters is about 100,000 clusters/mm<sup>2</sup>~Approximately 1,000,000 clusters/mm<sup>2</sup>In another embodiment, the density of the clusters is in the range of about 1,000,000 clusters/mm<sup>2</sup>~Approximately 10,000,000 clusters/mm<sup>2</sup>The range is.
In one embodiment, the preliminary center coordinates of the clusters determined by the base caller are defined within a template image of the tile, hi some embodiments, the pixel resolution, image coordinate system, and measurement scale of the image coordinate system are the same as the template image and the image.
In another embodiment, the disclosed technology relates to determining metadata about clusters on a tile of a flow cell. First, the disclosed technology accesses (1) a set of images of the tile captured during a sequencing run, and (2) preliminary center coordinates of the clusters determined by a base caller.
Then, for each image set, the disclosed technique obtains one of four bases: (1) an origin subpixel with a preliminary center coordinate; and (2) a predefined neighborhood of consecutive subpixels that are contiguously adjacent to each of the origin subpixels. This generates a base call sequence for each of the origin subpixels and each of the predefined neighborhood of consecutive subpixels. The predefined neighborhood of consecutive subpixels can be an m×n subpixel patch centered on a subpixel that includes the origin subpixel. In one embodiment, the subpixel patch is 3×3 subpixels. In other embodiments, the image patch can be any size, such as 5×5, 15×15, 20×20, etc. In other embodiments, the predefined neighborhood of consecutive subpixels can be an n connected subpixel neighborhood centered on a subpixel that includes the origin subpixel.
In one embodiment, the disclosed technique identifies as background those sub-pixels in the cluster map that do not belong to any of the disjoint regions.
The disclosed techniques then generate a cluster map that identifies clusters as discontinuous regions of adjacent subpixels that (a) are contiguous to at least a portion of a corresponding one of the origin subpixels and (b) share a substantially matching base call sequence of one of the four bases with at least a portion of a corresponding one of the origin subpixels.
The disclosed technique then stores the cluster map in memory and determines the shape and size of the clusters based on discontinuous regions in the cluster map. In another embodiment, the centers of the clusters are also determined.
(Generating training data for the template generator)
FIG. 15 is a block diagram illustrating one embodiment for generating the training data used to train the neural network-based template generator 1512 and the neural network-based basecaller 1514.
16 illustrates characteristics of the disclosed training examples used to train the neural network based template generator 1512 and the neural network based base caller 1514. Each training example corresponds to a tile and is labeled with a corresponding ground truth data representation. In some implementations, the ground truth data representation is a ground truth mask or map that identifies ground truth cluster metadata in the form of a ground truth attenuation map 1204, a ground truth ternary map 1304, or a ground truth binary map 1404. In some implementations, multiple training examples correspond to the same tile.
In one embodiment, the disclosed technology relates to generating training data 1504 for neural network based template generation and base calling. First, the disclosed technology accesses a number of images 108 of a flow cell 202 captured over multiple cycles of a sequencing run. The flow cell 202 has a number of tiles. In the number of images 108, each of the tiles has a series of image sets generated over multiple cycles. Each image in the sequence of image sets 108 shows the intensity emission of one particular cluster 302 of tiles and their surrounding background 304 at one particular cycle.
The training set constructor 1502 then constructs a training set 1504 having a number of training examples. As shown in FIG. 16, each training example corresponds to a particular one of the tiles and includes image data from at least some of the image sets in the sequence of image sets 1602 of the particular one of the tiles. In one implementation, the image data includes images in at least some of the image sets in the sequence of image sets 1602 of the particular one of the tiles. For example, the images may have a resolution of 1800×1800. In other implementations, the images may be of any resolution, such as 100×100, 3000×3000, 10000×10000, etc. In yet other implementations, the image data includes at least one image patch from each of the images. In one implementation, the image patch covers a portion of a particular one of the tiles. In one example, the image patch may have a resolution of 20×20. In other implementations, the image patches can have any resolution, such as 50x50, 70x70, 90x90, 100x100, 3000x3000, 10000x10000, etc.
In some implementations, the image data includes an upsampled representation of the image patch. The upsampled representation may have a resolution of, for example, 80 x 80. In other implementations, the upsampled representation may have any resolution, such as 50 x 50, 70 x 70, 90 x 90, 100 x 100, 3000 x 3000, 10000 x 10000, etc.
In some implementations, the multiple training examples correspond to the same particular one of the tiles and each include a different image patch from a respective image of at least some of the image sets in the sequence of image sets 1602 for the same particular one of the tiles. In such implementations, at least some of the different image patches overlap one another.
The ground truth generator 1506 then generates at least one ground truth data representation for each of the training examples, which identifies the spatial distribution of the clusters and at least one of their surrounding background, as represented by the image data, including at least one of the cluster shapes, cluster sizes, and/or cluster boundaries, and/or cluster centers.
In one implementation, the ground truth data representation identifies clusters as discontinuous regions of adjacent sub-pixels, with the centers of the clusters being the centers of the mass sub-pixels in corresponding ones of the discontinuous regions, and their surrounding background.
In one implementation, the ground truth data representation has an upsampled resolution of 80 x 80. In other implementations, the ground truth data representation can have any resolution, such as 50 x 50, 70 x 70, 90 x 90, 100 x 100, 3000 x 3000, 10000 x 10000, etc.
In one embodiment, the ground truth data representation identifies each sub-pixel as being either a cluster center or non-center, hi another embodiment, the ground truth data representation identifies each sub-pixel as being either a cluster interior, a cluster center, or the surrounding background.
In some implementations, the disclosed technology stores the training set 1504 and associated ground truth data 1508 in memory as training data 1504 for training a neural network based template generator 1512 and a neural network based base caller 1514. The training is handled by a trainer 1510.
In some embodiments, the disclosed techniques generate training data for a variety of flow cells, sequencing instruments, sequencing protocols, sequencing chemistries, sequencing reagents, and cluster densities.
(Neural network based template generator)
In an inference or production embodiment, the disclosed technique uses peak detection and segmentation to determine cluster metadata. The disclosed technique processes input image data 1702 derived from a series of image sets 1602 via a neural network 1706 to generate an alternative representation 1708 of the input image data 1702. For example, an image set may be for a particular sequence cycle and may include four images, one for each image channel A, C, T, and G. Thus, for a sequence that runs for 50 sequence cycles, there will be 50 such image sets, for a total of 200 images. When arranged in time, the image sets with four image patches per image set form a series of image sets 1602. In some implementations, image patches of a particular size are extracted from each image in the 50 image sets to form four image patch sets per image patch set, which in one implementation is the input image data 1702. In other implementations, the input image data 1702 includes image patch sets having four image patches per image patch set for image patch sets having fewer than 50 sequencing cycles, i.e., fewer than 1, 2, 3, 15, or 20 sequencing cycles.
17 illustrates one implementation of processing input image data 1702 through a neural network based template generator 1512 to generate an output value for each unit in an array. In one implementation, the array is an attenuation map 1716. In another implementation, the array is a ternary map 1718. In yet another implementation, the array is a binary map 1720. Thus, the array may represent one or more characteristics of each of a plurality of locations represented in the input image data 1702.
Unlike training the template generator using the structure of the previous figure, including the ground truth attenuation map 1204, the ground truth ternary map 1304, and the ground truth binary map 1404, the attenuation map 1716, the ternary map 1718, and/or the binary map 1720 are generated by forward propagation of the trained neural network based template generator 1512. The forward propagation can be during training or during inference. During training, by backpropagation based gradient updates, the attenuation map 1716, the ternary map 1718, and the binary map 1720 (i.e., cumulatively the output 1714) progressively match or approach the ground truth attenuation map 1204, the ground truth ternary map 1304, and the ground truth binary map 1404, respectively.
The size of the image array analyzed during inference depends on the size of the input image data 1702 according to one embodiment (e.g., the same or an upscaled or downscaled version). Each unit can represent a pixel, a subpixel, or a superpixel. The output value per unit of the array can characterize/represent/show an attenuation map 1716, a ternary map 1718, or a binary map 1720. In some implementations, the input image data 1702 is also a pixel-, subpixel-, or superpixel-resolution array of units. In another such implementation, the neural network-based template generator 1512 uses semantic segmentation techniques to generate output values for each unit in the input array. Further details regarding the input image data 1702 can be found in Figures 21b, 22, 23, and 24 and their discussions.
In some implementations, the neural network-based template generator 1512 is a fully convolutional network such as that described in J. Long, E. Shelhamer, and T. Darrell, "Fully convolutional networks for semantic segmentation," CVPR, (2015), which is incorporated by reference herein. In other alternative implementations, the neural network-based template generator 1512 is a fully convolutional network such as that described in Ronneberger O, Fischer P, Brox T, "U-net: Convolutional networks for biomedical image segmentation," Med. Image, available at http://link.springer.com/chapter/10.1007/978-3-319-24574-4_28, which is incorporated by reference herein. The neural network-based template generator 1512 is a U-Net network with skip connections between the decoder and the encoder, such as that described in "Comput. Comput. Assist. Interv. (2015)". The U-Net structure resembles an autoencoder with two main sub-structures: 1) an encoder that takes an input image and reduces its spatial resolution through multiple convolutional layers to generate an encoding; and 2) a decoder that takes the representation and encodes and increases the spatial resolution to generate a reconstructed image as output. U-Net introduces two innovations to this structure: first, the objective function is set to reconstruct the segmentation mask using a loss function, and second, the convolutional layers of the encoder are connected to the corresponding layers of the same resolution in the decoder using skip connections. In yet a further embodiment, the neural network-based template generator 1512 is a deep fully convolutional segmentation neural network with an encoder sub-network and a corresponding decoder network. In another such implementation, the encoder sub-network includes a hierarchy of encoders and the decoder sub-network includes a hierarchy of decoders that map the low-resolution encoder feature maps to the full input resolution feature maps. Further details regarding segmentation networks can be found in the Appendix entitled "Segmentation Networks."
In one embodiment, the neural network-based template generator 1512 is a convolutional neural network. In another embodiment, the neural network-based template generator 1512 is a recurrent neural network. In yet another embodiment, the neural network-based template generator 1512 is a residual neural network having residual Boch and residual connections. In a further embodiment, the neural network-based template generator 1512 is a combination of a convolutional neural network and a recurrent neural network.
It will be appreciated that the neural network-based template generator 1512 (i.e., the neural network 1706 and/or the output layer 1710) can use various padding and striding configurations. It can use different output functions (e.g., classification or regression) and may or may not include one or more fully connected layers. It can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or asexual convolution, transposed convolution, depth-separable convolution, 1×1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It can use one or more loss functions such as logistic regression/logarithmic loss, multiclass cross-entropy/softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It can use any parallel, efficient, and compression schemes such as TFRecords, compression encoding (e.g. PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallel, data parallel, and synchronous/asynchronous SGD. It includes nonlinear transformation functions such as upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, Pehjoll connections, activation functions (e.g. nonlinear transformation functions are rectified linear unit (ReLU), leaky ReLU, exponential linear unit (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g. max or average pooling), global average pooling layers, and attention mechanisms.
In some implementations, each image in the sequence of image sets 1602 covers a tile and shows the intensity emission of clusters on the tile and their surrounding background captured for a particular imaging channel at a particular one of multiple sequencing cycles of a sequencing run performed on the flow cell. In one implementation, the input image data 1702 includes at least one image patch from each of the images in the sequence of image sets 1602. In another such implementation, the image patch covers a portion of a tile. In one example, the image patch has a resolution of 20×20. In other cases, the resolution of the image patch may range from 20×20 to 10000×10000. In another implementation, the input image data 1702 includes upsampled subpixel resolution representations of image patches from each of the images in the sequence of image sets 1602. In one example, the upsampled subpixel representation has a resolution of 80×80. In other cases, the resolution of the upsampled subpixel representation may range from 80×80 to 10000×10000.
The input image data 1702 has an array of units 1704 that describe the clusters and their surrounding background. For example, an image set may be for a particular sequence cycle and may include four images, one for each image channel A, C, T, and G. Thus, for a sequence that runs for 50 sequence cycles, there will be 50 such image sets, for a total of 200 images. When arranged in time, the image sets form a sequence of image sets 1602 with four image patches per image set. In some implementations, image patches of a particular size are extracted from each image in the 50 image sets to form 50 image patch sets with four image patches per image patch set, which in one implementation is the input image data 1702. In other implementations, the input image data 1702 includes image patch sets with four image patches per image patch set for less than 50 sequencing cycles, i.e., less than 1, 2, 3, 15, 20 sequencing cycles. An alternative representation is a feature map. The feature maps may be convolutional features or convolutional representations when the neural network is a convolutional neural network, and may be hidden state features or hidden state representations when the neural network is a recurrent neural network.
The disclosed technique then processes the alternative representation 1708 through an output layer 1710 to generate an output 1714 having an output value 1712 for each unit in the array 1704. The output layer may be a classification layer, such as a softmax or sigmoid, that generates an output value per unit. In one implementation, the output layer is a ReLU layer or any other activation function layer that generates an output value per unit.
In one embodiment, the units in the input image data 1702 are pixels, and therefore an output value 1712 for each pixel is generated at the output 1714. In another embodiment, the units in the input image data 1702 are sub-pixels, and therefore an output value 1712 for each sub-pixel is generated at the output 1714. In yet another embodiment, the units in the input image data 1702 are super-pixels, and therefore an output value 1712 for each super-pixel is generated at the output 1714.
Deriving cluster metadata from attenuation maps, ternary maps and/or binary maps
18 illustrates one embodiment of a post-processing technique applied to the attenuation map 1716, ternary map 1718, or binary map 1720 generated by the neural network based template generator 1512 to derive cluster metadata including cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and/or cluster boundaries. In some embodiments, the post-processing technique is applied by a post-processor 1814, which further includes a threshold holder 1802, a peak locator 1806, and a divider 1810.
The input to the thresholder 1802 is an attenuation map 1716, a ternary map 1718, or a binary map 1720 generated by a template generator 1512, such as the disclosed neural network-based template generator. In one embodiment, the thresholder 1802 applies a threshold to values in the attenuation map, ternary map, or binary map to identify background units 1804 (i.e., subpixels that characterize non-cluster background) and non-background units. In other words, once the output 1714 is generated, the thresholder 1802 applies a threshold to the output values of the units 1712 to classify or reclassify a first subset of the units 1712 as "background units" 1804 that depict the background around the cluster and "non-background units" that represent units that may belong to the cluster. The threshold applied by the thresholder 1802 may be preset.
The input to the peak locator 1806 is also the attenuation map 1716, the ternary map 1718, or the binary map 1720, generated by the neural network-based template generator 1512. In one embodiment, the peak locator 1806 applies peak detection of values in the attenuation map 1716 to the ternary map 1718, or the binary map 1720 to identify center units 1808 (i.e., center subpixels that characterize cluster centers). In other words, the peak locator 1806 processes the output values of the units 1712 in the output 1714 and classifies a second subset of the units 1712 as "center units" 1808 that contain the centers of the clusters. In some embodiments, the centers of the clusters detected by the peak locator 1806 are also the center of mass of the clusters. The center units 1808 are then provided to a divider 1810. Further details regarding the peak locator 1806 can be found in the appendix entitled "Peak Detection".
Thresholding and peak detection can occur in parallel or after the other, i.e., they are not dependent on each other.
The input to the divider 1810 is also the attenuation map 1716, ternary map 1718, or binary map 1720 generated by the neural network based template generator 1512. Additional supplemental inputs to the divider 1810 include the thresholded units (background, non-background) 1804 identified by the thresholder 1802 and the center unit 1808 identified by the peak locator 1806. The divider 1810 uses the background, non-background 1804, and center unit 1808 to identify discontinuous regions 1812 (i.e., non-overlapping groups of adjacent clusters/inter-cluster sub-pixels that characterize a cluster). In other words, the divider 1810 processes the output values of the units 1712 in the output 1714 and uses the background and non-background units 1804, as well as the center unit 1808, to determine the shape 1812 of the cluster as a non-overlapping region of contiguous units separated by the background unit 1804 and centered on the center unit 1808. The output of the divider 1810 is cluster metadata 1812. The cluster metadata 1812 identifies cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and/or cluster boundaries.
In one embodiment, divider 1810 starts with central unit 1808 and determines, for each central unit, a set of consecutively consecutive units that represent the same cluster whose center of mass is contained in the central unit. In one embodiment, divider 1810 uses a so-called "watershed" segmentation technique to subdivide consecutive clusters into multiple adjacent clusters at valleys in intensity. Further details regarding the watershed segmentation technique and other segmentation techniques can be found in the Appendix entitled "Watershed Segmentation."
In one implementation, the output values of the units 1712 in the output 1714 are continuous values such as those coded in the ground truth attenuation map 1204. In another implementation, the output values are softmax scores such as those coded in the ground truth ternary map 1304 and the ground truth binary map 1404. In one implementation of the ground truth attenuation map 1204, successive units in corresponding ones of the non-overlapping regions have output values weighted according to the distance of the successive units from the central unit in the non-overlapping region to which the adjacent units belong. In such an embodiment, the central unit has the highest output value in each one of the non-overlapping regions. As described above, during training, the backpropagation based gradient update causes the attenuation map 1716, the ternary map 1718, and the binary map 1720 (i.e., cumulatively the output 1714) to progressively match or approach the ground truth ternary map 1304 and the ground truth binary map 1404 of the ground truth attenuation map 1204, respectively.
(Pixel domain - intensity extraction from regular cluster shapes)
The discussion now turns to how the cluster shapes determined by the disclosed techniques can be used to extract the intensities of the clusters. Because clusters typically have irregular shapes and contours, the disclosed techniques can be used to identify which sub-pixels contribute to discontinuous regions of irregular shape that represent the cluster shape.
19 illustrates one embodiment of extracting cluster intensities in the pixel domain. A "template image" or "template" can refer to a data structure that includes or identifies cluster metadata 1812 derived from the attenuation map 1716, the ternary map 1718, and/or the binary map 1718. The cluster metadata 1812 identifies cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and/or cluster boundaries.
In some implementations, the template image is in the upsampled sub-pixel domain to distinguish cluster boundaries at a fine level. However, the sequence image 108 containing the cluster and background intensity data is typically in the pixel domain. Thus, the disclosed technology proposes two approaches to extract the intensities of irregularly shaped clusters from the optical pixel resolution sequence image using the cluster shape information encoded in the template image in the upsampled sub-pixel resolution. In the first approach shown in FIG. 19, non-overlapping groups of contiguous sub-pixels identified in the template image are located in the pixel resolution sequence image and their intensities are extracted by interpolation. Further details regarding this intensity extraction technique can be found in FIG. 33 and its discussion.
In one implementation, when the non-overlapping regions have irregular contours and the units are sub-pixels, the cluster intensity 1912 of a given cluster is determined by the intensity extractor 1902 as follows.
First, the subpixel locator 1904 identifies the subpixels that contribute to the cluster intensity of a given cluster based on corresponding non-overlapping regions of adjacent subpixels that identify the shape of the given cluster.
The subpixel locator 1904 then locates the identified subpixel within one or more optical pixel resolution images 1918 generated for the one or more imaging channels in the current sequencing cycle. In one embodiment, integer or non-integer coordinates (e.g., floating points) are located within the optical resolution image, the pixel resolution image, after downscaling based on a downscaling factor that matches the upsampling factor used to create the subpixel domain.
An interpolator and subpixel intensity combiner 1906 for the identified subpixels in the processed images then combines the interpolated intensities and normalizes the combined interpolated intensities to generate a cluster intensity per image for a given cluster in each of the images. The normalization is performed by normalizer 1908 and is based on a normalization factor. In one embodiment, the normalization factor is the number of identified subpixels. This is done to normalize/account for different cluster sizes and non-uniform illumination that clusters receive depending on their location on the flow cell.
Finally, a cross-channel sub-pixel intensity accumulator 1910 combines the per-image cluster intensities for each of the images to determine a cluster intensity 1912 for a given cluster in the current sequence cycle.
The given cluster is then base called based on the cluster strengths 1912 in the current sequencing cycle by any one of the base calls discussed in this application to generate a base call 1916.
However, in some implementations, when the cluster size is large enough, the outputs of the neural network based base caller 1514, i.e., the attenuation map 1716, the ternary map 1718 and the binary map 1720, are in the optical pixel domain. Thus, in such implementations, the template image is also in the optical pixel domain.
(Sub-pixel domain - intensity extraction from regular cluster shapes)
FIG. 20 shows a second approach to extract cluster intensities in the sub-pixel domain. In this second approach, the sequence images are optically upsampled from pixel resolution to sub-pixel resolution. This results in a correspondence between the cluster shapes describing the sub-pixels in the template image and the cluster intensities representing the sub-pixels in the upsampled sequence images. Cluster intensities are then extracted based on the correspondence. Further details regarding this intensity extraction technique can be found in FIG. 33 and its discussion.
In one implementation, when the non-overlapping regions have irregular contours and the units are sub-pixels, the cluster intensity 2012 of a given cluster is determined by the intensity extractor 2002 as follows.
First, the subpixel locator 2004 identifies the subpixels that contribute to the cluster intensity of a given cluster based on corresponding non-overlapping regions of adjacent subpixels that define the shape of the given cluster.
The subpixel locator 2004 then locates the identified subpixels in one or more subpixel resolution images 2018 that are upsampled from the corresponding optical pixel resolution images 1918 generated for one or more imaging channels in the current sequencing cycle. The upsampling may be performed by nearest neighbor intensity extraction, Gaussian-based intensity extraction, intensity extraction based on average 2x2 subpixel area, intensity extraction based on brightest test of 2x2 subpixel area, average 3x3 subpixel area, bilinear intensity extraction, bilinear intensity extraction, and/or intensity extraction based on weighted area coverage. These techniques are described in detail in the appendix entitled "Intensity Extraction Methods." The template image may serve as a mask for intensity extraction in some implementations.
A subpixel intensity combiner 2006 in each of the upsampled images then combines the intensities of the identified subpixels and normalizes the combined intensity to generate a cluster intensity per image for a given cluster in each of the upsampled images. The normalization is performed by normalizer 2008 and is based on a normalization factor. In one embodiment, the normalization factor is the number of identified subpixels. This is done to normalize/account for different cluster sizes and non-uniform illumination that clusters receive depending on their location on the flow cell.
Finally, a cross-channel sub-pixel intensity accumulator 2010 combines the cluster intensities per image for each of the upsampled images to determine a cluster intensity 2012 for a given cluster in the current sequence cycle.
The given cluster is then called based on the cluster strengths 2012 in the current sequencing cycle by any one of the base calls discussed in this application to generate a base call 2016.
(Type of neural network-based template generator)
This discussion details three different implementations of the neural network-based template generator 1512, shown in Figure 21a and including: (1) an attenuation map-based template generator 2600 (also referred to as a regression model), (2) a binary map-based template generator 4600 (also referred to as a binary classification model), and (3) a ternary map-based template generator 5400 (also referred to as a ternary classification model).
In one embodiment, regression model 2600 is a fully convolutional network. In another embodiment, regression model 2600 is a U-Net network with skip connections between the decoder and the encoder. In one implementation, binary classification model 4600 is a fully convolutional network. In another embodiment, binary classification model 4600 is a U-Net network with skip connections between the decoder and the encoder. In one implementation, ternary classification model 5400 is a fully convolutional network. In another embodiment, ternary classification model 5400 is a U-Net network with skip connections between the decoder and the encoder.
(Input image data)
21b illustrates one embodiment of input image data 1702 provided as input to the neural network-based template generator 1512. The input image data 1702 includes a series of image sets 2100 having sequence images 108 generated during a particular number of initial sequencing cycles of a sequencing sequence (e.g., the first 2-7 sequencing cycles).
In some embodiments, the intensities of the sequence images 108 are corrected for background and/or aligned with each other using affinity transformation. In one embodiment, the sequencing run utilizes four channel chemistry and each image set has four images. In another embodiment, the sequencing run utilizes two channel chemistry and each image set has two images. In yet another embodiment, the sequencing run utilizes one channel chemistry and each image set has two images. In yet another embodiment, each image set has only one image. These and other different embodiments are described in Appendices 6 and 9.
Each image 2116 in the series of image sets 2100 covers a tile 2104 of the flow cell 2102 and shows the intensity emission of a cluster 2106 on the tile 2104 and their surrounding background captured for a particular image channel at a particular one of multiple sequencing cycles of a sequencing run. In one example, for cycle t1, the image set includes four images 2112A, 2112C, 2112T, 2112G, including one image for each base A, C, T, and G labeled with a corresponding fluorescent dye and imaged in a corresponding wavelength band (image/imaging channel).
For illustrative purposes, in image 2112G, FIG. 21b shows cluster intensity emission as 2108 and background intensity emission as 2110. In another embodiment, for cycle tn, the image set also includes four images 2114A, 2114C, 2114T, 2114G, including one image for each base A, C, T, and G labeled with a corresponding fluorescent dye and imaged in a corresponding wavelength band (image/imaging channel). Also for illustrative purposes, in image 2114A, FIG. 21b shows cluster intensity emission as 2118 and in image 2114T shows background intensity emission as 2120.
The input image data 1702 is encoded using an intensity channel (also called an imaging channel). For each c-image acquired from the sequencer for a particular sequencing cycle, a separate imaging channel is used to encode its intensity signal data. For example, consider that the sequencing uses a two-channel chemistry that produces a red image and a green image in each sequencing cycle. In such a case, the input data 2632 includes (i) a first red imaging channel having w×h pixels that shows the intensity emission of one or more clusters and their surrounding background captured in the red image, and (ii) a second green imaging channel having w×h pixels that shows the intensity emission of one or more clusters and their surrounding background captured in the green image.
(Non-image data)
In another embodiment, the input data to the neural network based template generator 1512 and the neural network based base caller 1514 is based on pH changes induced by the release of hydrogen ions during molecular elongation, which are detected and converted to a voltage change proportional to the number of bases incorporated (e.g., in the case of Ion Torrent).
In yet another embodiment, the input data is constructed from nanopore sensing using a biosensor to measure a disruption in electrical current as an analyte passes through or near the opening of the nanopore. For example, Oxford Nanopore ONT sequencing is based on the following concept: a single strand of DNA (or RNA) is passed through a membrane via a nanopore and a potential difference is applied across the membrane. Nucleotides present within the pore affect the electrical resistance of the pore, so that current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal ("squished" due to its appearance when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer data acquisition (DAC) values taken at a 4 kHz frequency (for example). With a DNA strand speed of ~450 base pairs per second, this gives, on average, about 9 raw observations per base. This signal is then processed to identify breaks in the pore signal that correspond to individual reads. Extensions of these raw signals are called bases, a process that converts DAC values into sequences of DNA bases. In some embodiments, the input data includes normalized or scaled DAC values.
In another embodiment, image data is not used as input to the neural network based template generator 1512 or the neural network based base caller 1514. Instead, the input to the neural network based template generator 1512 and the neural network based base caller 1514 is based on pH changes induced by the release of hydrogen ions during molecular elongation. The pH changes are detected and converted to voltage changes proportional to the number of bases incorporated (e.g., in the case of Ion Torrent).
In yet another embodiment, the input to the neural network based template generator 1512 and the neural network based base caller 1514 is constructed from nanopore sensing using a biosensor to measure a disruption in electrical current when an analyte passes through or near the opening of the nanopore. ONT sequencing is based on the following concept: a single strand of DNA (or RNA) is passed through a membrane via a nanopore and a potential difference is applied across the membrane. Nucleotides present within the pore affect the electrical resistance of the pore, so that current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal ("squished" due to its appearance when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer data acquisition (DAC) values taken at a 4 kHz frequency (for example). With a DNA strand speed of ~450 base pairs per second, this gives, on average, about 9 raw observations per base. This signal is then processed to identify breaks in the pore signal that correspond to individual reads. Extensions of these raw signals are base-called, a process that converts the DAC values into sequences of DNA bases. In some embodiments, the input data 2632 includes normalized or scaled DAC values.
(Patch Extraction)
Figure 22 illustrates one embodiment of extracting patches from the sequence of image sets 2100 of Figure 21b to generate a sequence of "downsized" image sets that form the input image data 1702. In another embodiment shown, the sequence images 108 in the sequence of image sets 2100 are of size LxL (e.g., 2000x2000). In other embodiments, L is any number ranging from 1 to 10,000.
In one embodiment, patch extractor 2202 extracts patches from sequence images 108 in the series of image sets 2100 to generate a series of downsized image sets 2206, 2208, 2210, and 2212. Each image in the series of downsized image sets is a patch of size M×M (e.g., 20×20) extracted from a corresponding sequence determined image in the series of image sets 2100. The size of the patch can be preset. In other alternative embodiments, M is any number in the range of 1 to 1000.
In FIG. 22, four exemplary series of downsized image sets are shown. A first exemplary series of downsized image sets 2206 is extracted from coordinates 0,0 to 20,20 in the sequence image 108 in the series of image sets 2100. A second exemplary series of downsized image sets 2208 is extracted from coordinates 20,20 to 40,40 in the sequence image 108 in the series of image sets 2100. A third exemplary series of downsized image sets 2210 is extracted from coordinates 40,40 to 60,60 in the sequence image 108 in the series of image sets 2100. A fourth exemplary series of downsized image sets 2212 is extracted from coordinates 60,60 to 80,80 in the sequence image 108 in the series of image sets 2100.
In some implementations, the series of downsized image sets forms the input image data 1702 that is provided as an input to the neural network based template generator 1512. Multiple series of downsized image sets can be provided simultaneously as an input batch, and a separate output can be generated for each series in the input batch.
(Upsampling)
FIG. 23 illustrates one implementation of upsampling the sequence of image sets 2100 of FIG. 21 b to generate a sequence of upsampled image sets 2300 that form the input image data 1702 .
In one embodiment, the upsampler 2302 upsamples the sequence images 108 in the series of image sets 2100 by an upsampling factor (eg, 4x) and generates a series of upsampled image sets 2300 .
In another illustrated embodiment, a sequence image 108 in the series of image sets 2100 is of size L×L (e.g., 2000×2000) and is upsampled by an upsampling factor of 4 to generate an upsampled image of size U×U (e.g., 8000×8000) in the series of upsampled image sets 2300.
In one embodiment, the sequence images 108 in the series of images set 2100 are directly fed to the neural network based template generator 1512, and the upsampling is performed by an initial layer of the neural network based template generator 1512. That is, the upsampler 2302 is part of the neural network based template generator 1512 and operates as a first layer to upsample the sequence images 108 in the series of images set 2100 and generate the series of upsampled images set 2300.
In some implementations, the series of upsampled image sets 2300 form the input image data 1702 that is provided as an input to the neural network based template generator 1512 .
FIG. 24 illustrates one embodiment for extracting patches from the series of upsampled image sets 2300 of FIG. 23 to generate a series of upsampled and downsized image sets 2406 , 2408 , 2410 and 2412 that form the input image data 1702 .
In one embodiment, the patch extractor 2202 extracts patches from the upsampled images in the series of upsampled image sets 2300 to generate a series of upsampled image sets 2406, 2408, 2410 and a downsized image set 2412. Each upsampled image in the series of upsampled image sets and the downsized image set is a patch of size M×M (e.g., 80×80) extracted from a corresponding upsampled image in the series of upsampled image sets 2300. The size of the patch can be preset. In another alternative embodiment, M is any number in the range of 1 to 1000.
In Fig. 24, four exemplary series of upsampled and downsized image sets are shown. A first series of examples of upsampled and downsized image sets 2406 are extracted from coordinates 0,0 to 80,80 in upsampled images in the series of upsampled image sets 2300. A second series of examples of upsampled and downsized images 2408 are extracted from coordinates 80,80 to 160,160 in upsampled images in the series of upsampled image sets 2300. A third series of examples of upsampled and downsized images 2410 are extracted from coordinates 160,160 to 240,240 in upsampled images in the series of upsampled image sets 2300. A fourth series of examples of upsampled and downsized images 2412 are extracted from coordinates 240,240 to 320,320 in upsampled images in the series of upsampled image sets 2300.
In some implementations, the series of upsampled and downsized image sets form the input image data 1702 that is provided as an input to the neural network based template generator 1512. Multiple series of upsampled and downsized image sets can be provided simultaneously as an input batch, and a separate output can be generated for each series in the input batch.
(output)
The three models are trained to produce different outputs. This is achieved by using different types of ground truth data representations as training labels. The regression model 2600 is trained to produce outputs that characterize/represent the so-called "attenuation map" 1716. The binary classification model 4600 is trained to produce outputs that characterize/represent the so-called "binary map" 1720. The ternary classification model 5400 is trained to produce outputs that characterize/represent the so-called "ternary map" 1718.
The output 1714 of each type of model includes a unit array 1712. The unit 1712 can be a pixel, a subpixel, or a superpixel. The output of each type of model includes output values per unit such that the output values of the unit array jointly characterize/represent/represent an attenuation map 1716 in the case of the regression model 2600, a binary map 1720 in the case of the binary classification model 4600, and a ternary map 1718 in the case of the ternary classification model 5400. The details are as follows:
(Ground truth data generation)
25 illustrates one implementation of an overall exemplary process for generating ground truth data for training the neural network based template generator 1512. For the regression model 2600, the ground truth data may be the attenuation map 1204. For the binary classification model 4600, the ground truth data may be the binary map 1404. For the ternary classification model 5400, the ground truth data may be the ternary map 1304. The ground truth data is generated from cluster metadata. The cluster metadata is generated by the cluster metadata generator 122. The ground truth data is generated by the ground truth data generator 1506.
In another embodiment shown, ground truth data is generated for tile A on lane A of flow cell A. The ground truth data is generated from sequence images 108 of tile A captured during a sequencing run. The sequence images 108 of tile A are in the pixel domain. In one example with a four-channel chemistry generating four sequence images per sequencing cycle, 200 sequence images 108 for 50 sequencing cycles are accessed. Each of the 200 sequence images 108 shows the intensity emission of clusters on tile A and their surrounding background captured in a particular image channel at a particular sequencing cycle.
The sub-pixel addresser 110 converts the sequence image 108 into the sub-pixel domain (eg, by dividing each pixel into multiple sub-pixels) and generates a sequence image 112 in the sub-pixel domain.
A base caller 114 (e.g., RTA) then processes the sequence image 112 in the sub-pixel domain and generates a base call for each sub-pixel and each of the 50 sequencing cycles, referred to herein as a "sub-pixel base call."
The subpixel base calls 116 are then merged to generate a base call sequence for each subpixel over the 50 sequencing cycles. Each subpixel base call sequence has 50 base calls, i.e., one base call for each of the 50 sequencing cycles.
The searcher 118 evaluates the base call sequences of successive subpixels on a pairwise basis. The search involves evaluating each subpixel to determine which of the successive subpixels share substantially matching base call sequences. The base call sequences of successive subpixels "substantially match" when a predetermined portion of the base calls match a per-ordinal position criterion (e.g., 41 matches in >=45 cycles, 4 mismatches in <=45 cycles, 4 mismatches in <=50 cycles, or 2 mismatches in <=34 cycles).
In some implementations, the base caller 114 also identifies preliminary center coordinates of the cluster. The subpixel containing the preliminary center coordinate is referred to as the center or origin subpixel. Some example preliminary center coordinates (604a-c) identified by the base caller 114 and corresponding origin subpixels (606a-c) are shown in FIG. 6. However, as described below, the identification of the origin subpixel (preliminary center coordinate of the cluster) is not required. In some implementations, the searcher 118 uses a breadth-first search to identify substantially matching base call sequences of subpixels, starting from the origin subpixels 606a-c and continuing through successive non-origin subpixels 702a-c. This is optional, as described below.
The search for a substantially matching base call sequence for a subpixel does not require identification of an origin subpixel (initial center coordinate of a cluster) because the search can be performed for all subpixels and the search need not start at the origin subpixel, but instead at any subpixel (e.g., the 0,0 subpixel or any random subpixel). Thus, the search can start at any subpixel, without having to utilize the origin subpixel, because each subpixel is evaluated to determine whether it shares a substantially matching base call sequence with another contiguous subpixel.
Regardless of whether the origin subpixel is used, certain clusters that do not include the origin subpixel (initial center coordinate of the cluster) are identified as predicted by base caller 114. Some examples of clusters that are identified by merging subpixel base calls and do not include the origin subpixel are clusters 812a, 812b, 812c, 812d, and 812e in Figure 8a. Thus, the use of base caller 114 to identify the origin subpixel (initial center coordinate of the cluster) is optional and not required for the search for substantially matching base call sequences of subpixels.
Searcher 118: (1) identifies contiguous subpixels having substantially matching base call sequences as so-called "discontiguous regions", (2) further evaluates the base call sequences of these subpixels that do not belong to any of the non-joint regions already identified in (1) to obtain additional discontiguous regions, and (3) then identifies background subpixels as subpixels that do not belong to any of the discontiguous regions already identified in (1) and (2). Action (2) enables the disclosed techniques to identify additional or additional clusters whose centers are not identified by base caller 114.
The results of searcher 118 are encoded in a so-called "cluster map" for Tile A and stored in cluster map data store 120. In the cluster map, each of the clusters on Tile A is identified by respective discontinuous regions of adjacent sub-pixels, and background sub-pixels separate the isolated regions to identify the surrounding background on Tile A.
Center of mass (COM) calculator 1004 determines the center of each of the clusters on Tile A by calculating the COM of each of the discontinuous regions as the average of the coordinates of each contiguous sub-pixel that forms the discontinuous region. The centers of mass of the clusters are stored as COM data 2502.
Subpixel classifier 2504 uses cluster map and COM data 2502 to generate subpixel classifications 2506. Subpixel classifier 2506 classifies (1) background subpixels, (2) COM subpixels (one COM subpixel for each discontinuous region including the COM of each discontinuous region), and (3) cluster/intra-cluster subpixels that form each discontinuous region. That is, each subpixel in the cluster map is assigned one of three categories.
Based on the subpixel classification 2506 in some embodiments, (i) a ground truth attenuation map 1204 is generated by the ground truth attenuation map generator 1202, (ii) a ground truth binary map 1304 is generated by the ground truth binary map generator 1302, and (iii) a ground truth ternary map 1404 is generated by the ground truth ternary map generator 1402.
1. (Regression model)
26 illustrates one embodiment of a regression model 2600. In another embodiment shown, the regression model 2600 is a fully convolutional network 2602 that processes input image data 1702 through an encoder sub-network and a corresponding decoder sub-network. The encoder sub-network includes a hierarchy of encoders. The decoder sub-network includes a hierarchy of decoders that map the low resolution encoder feature maps to the full input resolution attenuation maps 1716. In another embodiment, the regression model 2600 is a U-Net network 2604 with skip connections between the decoders and the encoders. Further details regarding segmentation networks can be found in the appendix entitled "Segmentation Networks."
(Attenuation map)
27 illustrates one implementation of generating a ground truth attenuation map 1204 from a cluster map 2702. The ground truth attenuation map 1204 is used as ground truth data for training the regression model 2600. In the ground truth attenuation map 1204, the ground truth attenuation map generator 1202 assigns a weighted attenuation value to each neighboring subpixel based on a weighted attenuation coefficient. The weighted attenuation value is proportional to the Euclidean distance of the neighboring subpixel from the center of mass (COM) subpixel in the discontinuous region to which the neighboring subpixel belongs, such that the weighted attenuation value is highest (e.g., 1 or 100) for COM subpixels and decreases for subpixels further away from the COM subpixel. In some implementations, the weighted attenuation value is multiplied by a preset coefficient, such as 100.
Furthermore, the ground truth attenuation map generator 1202 assigns all background sub-pixels the same pre-determined value (eg, the minimum background value).
The ground truth attenuation map 1204 represents contiguous subpixels in discontinuous regions and background subpixels based on assigned values. The ground truth attenuation map 1204 also stores the assigned values in a unit array, with each unit in the array representing a corresponding subpixel in the input.
(Training)
FIG. 28 is one implementation of training 2800 of a regression model 2600 using a back-propagation based gradient update technique to modify the parameters of the regression model 2600 until the attenuation map 1716 generated by the regression model 2600 as a training output during training 2800 progressively approaches or matches the ground truth attenuation map 1204 of the ground.
Training 2800 includes iteratively optimizing to minimize an error 2806 between the attenuation map 1716 and the ground truth attenuation map 1204 and updating parameters of the regression model 2600 based on the error 2806. In one implementation, the loss function is the mean squared error, which is minimized for each subpixel between the weighted attenuation values of corresponding subpixels in the attenuation map 1716 and the ground truth attenuation map 1204.
Training 2800 includes hundreds, thousands, and/or millions of forward propagations 2808 and backward propagations 2810, including parallelogram techniques such as batching. Training data 1504 includes a series of upsampled and downsized image sets as input image data 1702. Training data 1504 is annotated with ground truth labels by annotator 2806. Training 2800 can be operated on by trainer 1510 using a stochastic gradient update algorithm such as Adam.
(inference)
29 is an implementation of template generation by regression model 2600 during inference 2900 in which attenuation map 1716 is generated by regression model 2600 as an inference output during inference 2900. An example of attenuation map 1716 is disclosed in the Appendix entitled "Regression_Model_Ouput." The Appendix includes unit weighted attenuation output values 2910 that together represent attenuation map 1716.
Inference 2900 includes hundreds, thousands, and/or millions of forward propagations 2904, including parallelogram techniques such as batching. Inference 2900 is performed on inference data 2908, which includes a series of upsampled and downsized image sets as input image data 1702. Inference 2900 can be acted upon by a tester 2906.
(Watershed separation)
30 includes (i) thresholding the attenuation map 1716 to identify background sub-pixels that characterize cluster backgrounds, and (ii) peak detection to identify center sub-pixels that characterize cluster centers. Thresholding is performed by a threshold keeper 1802 that uses a local threshold binary to generate a binarized output. Peak detection is performed by a peak locator 1806 to identify cluster centers. Further details regarding the peak locator can be found in the Appendix entitled "Peak Detection."
31 shows one implementation of the watershed segmentation technique that takes as input the background subpixels and the center subpixels, each identified by a thresholder 1802, and a peak locator 1806 finds the intensity valleys between adjacent clusters and outputs non-overlapping groups of adjacent cluster/intra-cluster subpixels that characterize the clusters. Further details regarding the watershed segmentation technique can be found in the Appendix entitled "Watershed Segmentation."
In one implementation, the watershed divider 3102 takes as input (1) the attenuation map 1716, (2) the negated output values 1802, and (3) the cluster centers identified by the peak locator 1806 as input (1) minus the output values 2910. Based on the input, the watershed divider 3102 then generates an output 3104. In the output 3104, each cluster center is identified as a unique set/group of sub-pixels that belong to the cluster center (as long as the sub-pixels are "1" in the binary output, i.e., are not background sub-pixels). Further, the clusters are filtered based on containing at least four sub-pixels. The watershed divider 3102 can be part of the divider 1810, which in turn is part of the post-processor 1814.
(Network Structure)
32 is a table showing an example U-Net structure for regression model 2600, including details of the layers, the dimensionality of the layer outputs, the model parameter magnitudes, and the interconnections between layers for regression model 2600. Similar details are disclosed in the file entitled "Regression_Model_Example_Architecture" submitted as an appendix to this application.
(Cluster Intensity Extraction)
33 shows a different approach to extracting cluster intensities using cluster shape information identified in a template image. As described above, the template image identifies the cluster shape information in the upsampled sub-pixel resolution. However, the cluster intensity information is in the sequence image 108, which is typically optical resolution.
According to the first approach, the coordinates of the sub-pixels are located in the sequence image 108 and their respective intensities are extracted using bilinear interpolation and normalized based on the count of the sub-pixels contributing to the cluster.
The second approach uses a weighted area coverage technique to modulate the intensity of a pixel according to the number of sub-pixels that contribute to the pixel. Again, the modulated pixel intensity is normalized by a sub-pixel count parameter.
The third technique uses quadratic interpolation to upsample the sequence images to the sub-pixel domain, sums the intensities of the upsampled pixels that belong to a cluster, and normalizes the summed intensity based on the count of the upsampled pixels that belong to the cluster.
(Experimental results and discussion)
Figure 34 shows different approaches to base calling using the output of the regression model 2600. In the first approach, the cluster centers identified from the output of the neural network based template generator 1512 in the template image are fed into a base caller (e.g., Illumina's Time Analysis software, referred to herein as "RTA base calling") for base calling.
In the second approach, cluster intensities extracted from sequence images based on cluster shape information in a template image instead of cluster centers are fed to the RTA base caller for base calling.
Figure 35 shows the difference in base calling performance when RTA base calling uses ground truth center of mass (COM) positions as cluster centers as opposed to using non-COM positions as cluster centers. The results show that using COM improves base calling.
(Example of model output)
Figure 36 shows, on the left, an example attenuation map 1716 generated by the regression model 2600. Figure 36 also shows, on the right, an example ground truth attenuation map 1204 that the regression model 2600 approximates during training.
Both the attenuation map 1716 and the ground truth attenuation map 1204 depict clusters as discontinuous regions of adjacent sub-pixels, with the centers of the clusters shown as central sub-pixels at the center of mass of corresponding regions of the discontinuous regions, and the clusters as their surrounding background.
Also, consecutive sub-pixels in corresponding ones of the discontinuous regions have values weighted according to the distance of the consecutive sub-pixels from a central sub-pixel in the discontinuous region to which the adjacent sub-pixels belong. In one implementation, the central sub-pixel has the highest value in the corresponding one of the discontinuous regions. In one implementation, all background sub-pixels have the same minimum background value in the attenuation map.
37 shows one embodiment of a peak locator 1806 that identifies cluster centers in an attenuation map by detecting peaks 3702. Further details regarding the peak locator can be found in the Appendix entitled "Peak Detection."
38 compares the peaks detected by the peak locator 1806 in the attenuation map 1716 generated by the regression model 2600 with the peaks in the corresponding ground truth attenuation map 1204. The red markers are the peaks predicted by the regression model 2600 as cluster centers, and the green markers are the ground truth centers of cluster agglomerations.
(Further experimental results and considerations)
39 shows the performance of the regression model 2600 using the accuracy and recalibration statistics. The accuracy and recalibration statistics demonstrate that the regression model 2600 is good at recovering all the identified cluster centers.
Figure 40 compares the performance of regression model 2600 with the RTA base caller for a library concentration of 20 pM (normal operation). By running the RTA base caller, regression model 2600 identifies 34,323 (4.46%) clusters in a higher cluster density environment (i.e., 988,884 clusters).
Figure 40 also shows the results of other sequencing metrics such as the number of clusters that pass the chasity filter ("%PF" (pass filter)), the number of aligned reads ("% aligned"), the number of overlapping reads ("% duplicates"), all reads that aligned to the reference sequence ("% mismatch"), the number of reads that do not match the reference sequence for a quality score of 30 and the bases referred to above ("%Q30 bases").
Figure 41 compares the performance of regression model 2600 with the RTA base caller for a 30 pM library concentration (high density run). Running with the RTA base caller, regression model 2600 identifies 34,323 (6.27%) more clusters in a much higher cluster density environment (i.e., 1,351,588 clusters).
Figure 41 also shows the results of other sequencing metrics such as the number of clusters that pass the chasity filter ("%PF" (pass filter)), the number of aligned reads ("% aligned"), the number of overlapping reads ("% duplicates"), all reads that aligned to the reference sequence ("% mismatch"), the number of reads that do not match the reference sequence for a quality score of 30 and the bases referred to above ("%Q30 bases").
Figure 42 compares the number of two non-overlapping (unique or overlapping duplicate) correct read pairs, i.e., the number of paired reads where both reads are aligned within a reasonable distance from each other, detected by regression model 2600 with those detected by the RTA-based collar. The comparison is done for both normal runs at 20 pM and high density runs at 30 pM.
More importantly, Figure 42 shows that the disclosed neural network-based template generator can detect more clusters with fewer sequencing cycles of input to template generation. With only four sequencing cycles, regression model 2600 identifies 11% more non-overlapping correct read pairs than RTA base caller in a 20 pM regular run and 33% more correct read pairs than RTA base caller in a 30 pM high density run. With seven sequencing cycles, regression model 2600 identifies 4.5% more non-overlapping correct read pairs than RTA base caller in a 20 pM regular run and 6.3% more correct read pairs than RTA base caller in a 30 pM high density run.
Figure 43 shows on the right a first attenuation map generated by regression model 2600. The first attenuation map identifies the clusters and their surrounding background imaged during normal operation at 20 pM, along with their spatial distribution showing the cluster shapes, cluster sizes, and cluster centers.
On the left, Figure 43 shows a second attenuation map generated by regression model 2600. The second attenuation map identifies the clusters imaged during the 30 pM high density run and their surrounding background, along with their spatial distribution showing cluster shape, cluster size, and cluster centers.
Figure 44 compares the performance of regression model 2600 with the RTA base caller for a library concentration of 40 pM (high density run). Regression model 2600 generated 89,441,688 more aligned bases than the RTA base caller in a much higher cluster density environment (i.e., 1,509,395 clusters).
Figure 44 also shows the results of other sequencing metrics such as the number of clusters that pass the chasity filter ("%PF" (pass filter)), the number of aligned reads ("% aligned"), the number of overlapping reads ("% duplicates"), all reads that aligned to the reference sequence ("% mismatch"), the number of reads that mismatch the reference sequence for a quality score of 30 and the bases referred to above ("%Q30 bases").
(Further examples of model output)
Figure 45 shows, on the left, a first attenuation map generated by regression model 2600. The first attenuation map identifies the clusters imaged during normal operation at 40 pM and their surrounding background, along with their spatial distribution showing cluster shape, cluster size, and cluster centers.
At the top right, Fig. 45 shows the results of a threshold and peak location applied to the first attenuation map to distinguish each cluster from each other and from the background and to identify their respective cluster centers. In some implementations, the intensity of each cluster is identified and a chassis filter (or pass filter) is specified that is applied to reduce the mismatch rate.
2. (Binary Classification Model)
FIG. 46 illustrates one example of a binary classification model 4600. In another embodiment shown, the binary classification model 4600 is a deep fully convolutional segmentation neural network that processes input image data 1702 through an encoder sub-network and a corresponding decoder sub-network. The encoder sub-network includes a hierarchy of encoders. The decoder sub-network includes a hierarchy of decoders that map the low-resolution encoder feature map to the full input resolution binary map 1720. In another embodiment, the binary classification model 4600 is a U-Net network with skip connections between the decoders and encoders. Further details regarding segmentation networks can be found in the appendix entitled "Segmentation Networks."
(Binary Map)
The final output layer of the binary classification model 4600 is a unit-wise classification layer that generates a classification label for each unit in the output array. In some implementations, the unit-wise partitioning layer is a sub-pixel-wise classification layer that generates a softmax classification score distribution for each sub-pixel in the binary map 1720 across two classes, i.e., the cluster center class and the non-cluster class, and the classification label for a given sub-pixel are determined from the corresponding softmax classification score distribution.
In another alternative embodiment, the unit-wise classification layer is a sub-pixel-wise classification layer that generates a sigmoid classification score for each sub-pixel in the binary map 1720 such that the activation of the unit is interpreted as the probability that the unit belongs to a first class, and conversely, one minus one gives the probability of belonging to a second class.
The binary map 1720 represents each subpixel based on its predicted classification score. The binary map 1720 also stores the predicted classification scores in a unit array, with each unit in the array representing a corresponding subpixel in the input.
(Training)
FIG. 47 is one implementation of training 4700 of a binary classification model 4600 using a backpropagation-based gradient update technique that modifies parameters of the binary classification model 4600 until the binary map 1720 of the binary classification model 4600 progressively approaches or matches the ground truth binary map 1404.
In the illustrated implementation, the final output layer of the binary classification model 4600 is a softmax-based per-subpixel classification layer. In another implementation of softmax, the ground truth binary map generator 1402 assigns each ground truth subpixel to either (i) a cluster center value pair (e.g., [1, 0]) or (ii) a non-center value pair (e.g., [0, 1]).
In a cluster-centered pair of values [1,0], the first value [1] represents the cluster-centered class label and the second value [0] represents the non-centered class label. In a non-centered pair of values [0,1], the first value [0] represents the cluster-centered class label and the second value [1] represents the non-centered class label.
The ground truth binary map 1404 represents each subpixel based on an assigned value pair/value. The ground truth binary map 1404 also stores the assigned value pairs/values in a unit array, with each unit in the array representing a corresponding subpixel in the input.
Training involves iteratively optimizing a loss function that minimizes an error 4706 (e.g., a softmax error) between the binary map 1720 and the ground truth binary map 1404, and updating parameters of the binary classification model 4600 based on the error 4706.
In one implementation, the loss function is a custom weighted binary cross-entropy loss, and the error 4706 is minimized for each subpixel between the predicted classification score (e.g., softmax score) and the labeled class score (e.g., softmax score) of the corresponding subpixel in the binary map 1720 and the ground truth binary map 1404, as shown in FIG.
The custom-weighted loss function gives more weight to the COM subpixel each time it is misclassified, multiplied by the corresponding reward (or penalty) weight specified in the reward (or penalty) matrix. Further details regarding the custom-weighted loss function can be found in the Appendix entitled "Custom-Weighted Loss Function".
Training 4700 includes hundreds, thousands, and/or millions of forward propagations 4708 and backward propagations 4710, including parallelogram techniques such as batching. Training data 1504 includes a series of upsampled and downsized image sets as input image data 1702. Training data 1504 is annotated with ground truth labels by annotator 2806. Training 2800 can be operated on by trainer 1510 using a stochastic gradient update algorithm such as Adam.
FIG. 48 is another embodiment of training 4800 of a binary classification model 4600, where the final output layer of the binary classification model 4600 is a sigmoid-based sub-pixel-wise classification layer.
In a sigmoid alternative implementation, the ground truth binary map generator 1302 assigns each ground truth subpixel either (i) a cluster center value (e.g., [1]) or (ii) a non-center value (e.g., [0]). The COM subpixel is assigned the cluster center value pair/value, and all other subpixels are assigned the non-center value pair/value.
For cluster central values, values above a threshold midpoint between 0 and 1 (e.g., values above 0.5) represent central class labels. For non-central values, values below a threshold midpoint between 0 and 1 (e.g., values below 0.5) represent non-central class labels.
The ground truth binary map 1404 represents each subpixel based on an assigned value pair/value. The ground truth binary map 1404 also stores the assigned value pairs/values in a unit array, with each unit in the array representing a corresponding subpixel in the input.
Training involves iteratively optimizing a loss function that minimizes an error 4806 (e.g., a sigmoid error) between the binary map 1720 and the ground truth binary map 1404, and updating parameters of the binary classification model 4600 based on the error 4806.
In one implementation, the loss function is a custom weighted binary cross-entropy loss, and the error 4806 is minimized for each subpixel between the predicted scores (e.g., sigmoid scores) of corresponding subpixels in the binary map 1720 and the ground truth binary map 1404, as shown in FIG. 48, and the labeled scores (e.g., sigmoid scores) of corresponding subpixels in the binary map 1720 and the ground truth binary map 1404, as shown in FIG. 48.
The custom-weighted loss function gives more weight to the COM subpixel each time it is misclassified, multiplied by the corresponding reward (or penalty) weight specified in the reward (or penalty) matrix. Further details regarding the custom-weighted loss function can be found in the Appendix entitled "Custom-Weighted Loss Function".
Training 4800 includes hundreds, thousands, and/or millions of forward propagations 4808 and backward propagations 4810, including parallelogram techniques such as batching. Training data 1504 includes a series of upsampled and downsized image sets as input image data 1702. Training data 1504 is annotated with ground truth labels by annotator 2806. Training 2800 can be operated on by trainer 1510 using a stochastic gradient update algorithm such as Adam.
FIG. 49 shows another implementation of the input image data 1702 provided to a binary classification model 4600 and the corresponding class labels 4904 used to train the binary classification model 4600.
In another embodiment shown, the input image data 1702 includes a series of upsampled and downsized image sets 4902. The class labels 4904 include two classes: (1) "no cluster center" and (2) "cluster center" are distinguished using different output values: (1) light green units/subpixels 4906 represent subpixels predicted by the binary classification model 4600 that do not contain cluster centers, and (2) dark green subpixels 4908 represent units/subpixels predicted by the binary classification model 4600 as containing cluster centers.
(inference)
Figure 50 is an implementation of template generation by a binary classification model 4600 during inference 5000 in which a binary map 1720 is generated by the binary classification model 4600 as an inference output during inference 5000. An example of a binary map 1720 includes binary classification scores 5010 per unit that together represent the binary map 1720. In a softmax application, the binary map 1720 has a first array 5002a of classification scores per unit for non-center classes and a second array 5002b of classification scores per unit for cluster center classes.
Inference 5000 includes hundreds, thousands, and/or millions of forward propagations 5004, including parallelogram techniques such as batching. Inference 5000 is performed on inference data 2908, which includes a series of upsampled and downsized image sets as input image data 1702. Inference 5000 can be operated on by tester 2906.
In some implementations, the binary map 1720 is subjected to the post-processing techniques described above, such as thresholding, peak detection, and/or watershed division, to generate cluster metadata.
(Peak detection)
Figure 51 illustrates one embodiment of subjecting a binary map 1720 to peak detection to identify cluster centers. As described above, the binary map 1720 is a unit array that classifies each subpixel based on a predicted classification score, with each unit in the array representing a corresponding subpixel in the input. The classification score can be a softmax score or a sigmoid score.
In a softmax application, the binary map 1720 includes two arrays: (1) a first array 5002a of classification scores per unit of the non-center classes, and (2) a second array 5002b of classification scores per unit of the cluster center classes, where each unit in both arrays represents a corresponding subpixel in the input.
To determine which sub-pixels in the input contain cluster centers and which do not, the peak locator 1806 applies peak detection on the units in the binary map 1720. Peak detection identifies units that have a classification score (e.g., softmax/sigmoid score) above a pre-set threshold. The identified units are inferred as cluster centers and their corresponding sub-pixels in the input are determined to contain cluster centers and are stored as cluster center sub-pixels in the sub-pixel classification data store 5102. Further details regarding the peak locator 1806 can be found in the Appendix entitled "Peak Detection".
The remaining units in the input and their corresponding subpixels do not contain the cluster center and are stored as non-central subpixels in the subpixel classification data store 5102.
In some implementations, before applying peak detection, units with classification scores below a certain background threshold (e.g., 0.3) are set to zero. In some implementations, such units and their corresponding subpixels in the input are inferred to represent the background surrounding the cluster and are stored as background subpixels in the subpixel classification data store 5102. In other implementations, such units can be considered noise and ignored.
(Example of model output)
Figure 52a shows an example binary map generated by the binary classification model 4600 on the left. Figure 52a also shows an example ground truth binary map on the right that the binary classification model 4600 approximates during training. The binary map has multiple sub-pixels and classifies each sub-pixel as either a cluster center or a non-center. Similarly, the ground truth binary map has multiple sub-pixels and classifies each sub-pixel as either a cluster center or a non-center.
(Experimental results and discussion)
Figure 52b shows the performance of the binary classification model 4600 using recalibration and refinement statistics. By applying these statistics, the binary classification model 4600 performs the RTA base caller.
(Network Structure)
53 is a table showing an example structure of a binary classification model 4600, along with details of the layers, the dimensionality of the layer outputs, the model parameter magnitudes, and the interconnections between layers of the binary classification model 4600. Similar details are disclosed in an appendix entitled "Binary_Classification_Model_Example_Architecture."
3. Three-way (three-class) classification model
FIG. 54 illustrates one implementation of a ternary classification model 5400. In another embodiment shown, the ternary classification model 5400 is a deep fully convolutional segmentation neural network that processes input image data 1702 through an encoder sub-network and a corresponding decoder sub-network. The encoder sub-network includes a hierarchy of encoders. The decoder sub-network includes a hierarchy of decoders that map the low-resolution encoder feature map to the full input resolution ternary map 1718. In another embodiment, the ternary classification model 5400 is a U-Net network with skip connections between the decoders and encoders. Further details regarding segmentation networks can be found in the appendix entitled "Segmentation Networks."
(Tripartite Map)
The final output layer of the ternary classification model 5400 is a unit-wise classification layer that generates a classification label for each unit in the output array. In some implementations, the unit-wise partitioning layer is a subpixel-wise classification layer that generates a softmax classification score distribution for each subpixel in the ternary map 1718 across three classes, i.e., background class, cluster center class, and cluster/intra-cluster class, and the classification label for a given subpixel is determined from the corresponding softmax classification score distribution.
The ternary map 1718 represents each subpixel based on its predicted classification score. The ternary map 1718 also stores the predicted classification scores in a unit array, with each unit in the array representing a corresponding subpixel in the input.
(Training)
FIG. 55 illustrates one implementation of training 5500 a ternary classification model 5400 using a backpropagation-based gradient update technique that modifies parameters of the ternary classification model 5400 until the ternary map 1718 of the ternary classification model 5400 progressively approaches or matches the training ground truth ternary map 1304.
In the illustrated implementation, the final output layer of the ternary classification model 5400 is a softmax-based per-subpixel classification layer. In another implementation, the softmax ternary map generator 1402 assigns each ground truth triplet either (i) a background value triplet (e.g., [1, 0, 0]), (ii) a cluster center value triplet (e.g., [0, 1, 0]), or (iii) a cluster/cluster interior value triplet (e.g., [0, 0, 1]).
Background subpixels are assigned background value triplets, center of mass (COM) subpixels are assigned cluster center value triplets, and cluster/interior-cluster subpixels are assigned cluster/interior-cluster value triplets.
In the background value triplet [1,0,0], the first value [1] represents the background class label, the second value [0] represents the cluster center label, and the third value [0] represents the cluster/inner-cluster class label.
In the cluster center triplet [0,1,0], the first value [0] represents the background class label, the second value [1] represents the cluster center label, and the third value [0] represents the cluster/intra-cluster class label.
In a cluster/cluster inner value triplet [0,0,1], the first value [0] represents the background class label, the second value [0] represents the cluster center label, and the third value [1] represents the cluster/cluster inner class label.
The ground truth ternary map 1304 represents each subpixel based on an assigned value triplet. The ground truth ternary map 1304 also stores the assigned triplets in a unit array, with each unit in the array representing a corresponding subpixel in the input.
Training involves iteratively optimizing a loss function that minimizes an error 5506 (e.g., a softmax error) between the ternary map 1718 and the ground truth ternary map 1304, and updating parameters of the ternary classification model 5400 based on the error 5506.
In one implementation, the loss function is a custom weighted categorization cross-entropy loss, and the error 5506 is minimized for each subpixel between the predicted classification score (e.g., softmax score) and the labeled class score (e.g., softmax score), as shown in FIG. 54, and between the labeled class scores (e.g., softmax scores) of the corresponding subpixels in the ternary map 1718 and the ground truth ternary map 1304.
The custom-weighted loss function gives more weight to the COM subpixel each time it is misclassified, multiplied by the corresponding reward (or penalty) weight specified in the reward (or penalty) matrix. Further details regarding the custom-weighted loss function can be found in the Appendix entitled "Custom-Weighted Loss Function".
Training 5500 includes hundreds, thousands, and/or millions of forward propagations 5508 and backward propagations 5510, including parallelogram techniques such as batching. Training data 1504 includes a series of upsampled and downsized image sets as input image data 1702. Training data 1504 is annotated with ground truth labels by annotator 2806. Training 5500 can be operated on by trainer 1510 using a stochastic gradient update algorithm such as Adam.
FIG. 56 illustrates one embodiment of input image data 1702 provided to a ternary classification model 5400 and the corresponding class labels used to train the ternary classification model 5400.
In another embodiment shown, the input image data 1702 includes a series of upsampled and downsized image sets 5602. The class labels 5604 include three classes: (1) a "background class", (2) a "cluster center class", and (3) a "cluster interior class", which are distinguished using different output values. For example, some of these different output values can be visually represented as follows: (1) gray units/subpixels 5606 represent subpixels predicted by the ternary classification model 5400 to be background, (2) dark green units/subpixels 5608 represent subpixels predicted by the ternary classification model 5400 to contain cluster centers, and (3) light green subpixels 5610 represent subpixels predicted by the ternary classification model 5400 to contain the interiors of clusters.
(Network Structure)
57 is a table showing an example structure of a ternary classification model 5400, along with details of the layers, the dimensionality of the layer outputs, the magnitude of the model parameters, and the interconnections between the layers of the ternary classification model 5400. Similar details are disclosed in an appendix entitled "Ternary_Classification_Model_Example_Architecture."
(inference)
FIG. 58 is an implementation of template generation by a ternary classification model 5400 during inference 5800 in which a ternary map 1718 is generated by the ternary classification model 5400 as an inference output during inference 5800. An example of a ternary map 1718 is disclosed in the appendix entitled "Ternary_Classification_Model_Ouput". The appendix includes binary classification scores 5810 per unit that together represent the ternary map 1718. In a softmax application, the appendix has a first array 5802a of classification scores per unit for background classes, a second array 5802b of classification scores per unit for cluster center classes, and a third array 5802c of classification scores per unit for cluster/intra-cluster classes.
Inference 5800 includes hundreds, thousands, and/or millions of forward propagations 5804, including parallelogram techniques such as batching. Inference 5800 is performed on inference data 2908, which includes a series of upsampled and downsized image sets as input image data 1702. Inference 5000 can be operated on by tester 2906.
In some implementations, the ternary map 1718 is generated by the ternary classification model 5400 using post-processing techniques described above, such as thresholding, peak detection, and/or watershed division.
FIG. 59 graphically illustrates the ternary map 1718 generated by the ternary classification model 5400 with the ternary softmax classification score distributions of three corresponding classes, namely, the background class 5906, the cluster center class 5902, and the cluster/cluster inner class 5904, respectively.
Figure 60 shows a unit array generated by a ternary classification model 5400 along with output values for each unit. As shown, each unit has three output values for three corresponding classes, namely background class 5906, cluster center class 5902, and cluster/cluster inner class 5904. For each classification (column wise), each unit is assigned the class with the highest output value, as indicated by the class in parentheses for each unit. In some implementations, output values 6002, 6004, and 6006 are analyzed for each of the respective classes 5906, 5902, and 5904 (row wise).
(Peak detection and watershed division)
Figure 61 illustrates one embodiment of subjecting the ternary map 1718 to post-processing to identify cluster centers, cluster backgrounds, and cluster interiors. As described above, the ternary map 1718 is an array of units that classifies each subpixel based on a predicted classification score, with each unit in the array representing a corresponding subpixel in the input. The classification score may be a softmax score.
In the softmax application, the ternary map 1718 includes three arrays: (1) a first array 5802a of per-unit classification scores for the background class, (2) a second array 5802b of per-unit classification scores for the cluster center classes, and (3) a third array 5802c of per-unit classification scores for the cluster inner classes. In all three arrays, each unit represents a corresponding sub-pixel in the input.
To determine which sub-pixels in the input contain cluster centers that contain the interior of a cluster and include background, the peak locator 1806 applies peak detection to the softmax values in the ternary map 1718 of the cluster center class 5802b. Peak detection identifies units that have a classification score (e.g., a softmax score) above a pre-set threshold. The identified units are inferred as cluster centers and their corresponding sub-pixels in the input are determined to contain cluster centers and are stored as cluster center sub-pixels in the sub-pixel classification and segmentation data store 6102. Further details regarding the peak locator 1806 can be found in the Appendix entitled "Peak Detection".
In some implementations, before applying peak detection, units with classification scores below a certain noise threshold (e.g., 0.3) are set to zero. Such units can be considered as noise and can be ignored.
Also, those subpixels having a classification score of the background class 5802a above a certain background threshold (e.g., 0.5 or greater) and their corresponding subpixels in the input are inferred to indicate background surrounding the cluster and are stored as background subpixels in the subpixel classification and segmentation data store 6102.
A watershed segmentation algorithm operated by watershed segment 3102 is then used to determine the shapes of the clusters. In some implementations, the background units/subpixels are used as a mask by the watershed segmentation algorithm. The classification scores of the units/subpixels inferred as cluster centers and cluster interiors are summed to generate a so-called "cluster label". The cluster centers are used as watershed markers for separation by intensity valleys by the watershed segmentation algorithm.
In one embodiment, the negatively polarized cluster labels are provided as an input image to a watershed divider 3102, which performs the segmentation and generates cluster shapes as discontinuous regions of adjacent cluster interior subpixels separated by background subpixels. Additionally, each discontinuous region includes a corresponding cluster center subpixel. In some embodiments, the corresponding cluster center subpixel is the center of the region to which it belongs. In other embodiments, the center of mass (COM) of the discontinuous region is calculated based on the underlying position coordinates and stored as the new center of the cluster.
The output of the watershed segmenter 3102 is stored in the subpixel classification and segmentation data store 6102. Further details regarding the watershed segmentation algorithm and other segmentation algorithms can be found in the Appendix entitled "Watershed Segmentation."
Example outputs of the peak locator 1806 and the watershed divider 3102 are shown in FIGS.
(Example of model output)
FIG. 62a shows an example prediction of the ternary classification model 5400. FIG. 62a shows four maps, each with a unit arrangement. The first map 6202 (far left) shows the output value of each unit of the cluster center class 5802b. The second map 6204 shows the output value of each unit of the cluster/cluster inner class 5802c. The third map 6206 (far right) shows the output value of each unit of the background class 5802a. The fourth map 6208 (bottom) is a binary mask of the ground truth ternary map 6008, which assigns each unit the class label with the highest output value.
Figure 62b shows another example prediction of the ternary classification model 5400. Figure 62b shows four maps, each with a unit arrangement. The first map 6212 (bottom) shows the output value of each unit of the cluster/cluster inner class. The second map 6214 shows the output value of each unit of the cluster center class. The third map 6216 (rightmost) shows the output value of each unit of the background class. The fourth map (top) 6210 is a ground truth ternary map that assigns each unit the class label with the highest output value.
Figure 62c shows yet another example prediction of the ternary classification model 5400. Figure 64 shows four maps, each with a unit arrangement. The first map 6220 (bottom) shows the output value of each unit of the cluster/cluster inner class. The second map 6222 shows the output value of each unit of the cluster center class. The third map 6224 (rightmost) shows the output value of each unit of the background class. The fourth map 6218 (top) is a ground truth ternary map that assigns each unit the class label with the highest output value.
Figure 63 shows one embodiment of deriving cluster centers and cluster shapes from the output of the ternary classification model 5400 of Figure 62a by subjecting the output to post-processing. The post-processing (e.g., peak locations, watershed divisions) produces cluster shape data and other metadata that are identified in a cluster map 6310.
(Experimental results and discussion)
FIG. 64 compares the performance of the binary classification model 4600, the regression model 2600, and the RTA base caller. Performance is evaluated using various sequencing metrics. One metric is the total number of clusters detected ("#clusters"), which can be measured by the number of unique cluster centers detected. Another metric is the number of detected clusters that pass the chacity filter ("%PF" (pass filter)). During cycles 1-25 of the sequencing run, the chacity filter removes the least reliable clusters from the image extraction results. A cluster "passes the filter" if no more than one base call has a chacity value of less than 0.6 in the first 25 cycles. The chacity is defined as the ratio of the brightest base intensity divided by the sum of the brightest test and the second brightest base intensity. This metric goes beyond the amount of clusters detected and also their quality, i.e., how many of the detected clusters can be used for accurate base calling and downstream secondary and ternary analyses, such as variant calling and variant pathogenicity annotation.
Other metrics measuring how good a detected cluster is for downstream analysis include the number of aligned reads generated from the detected cluster ("% aligned"), the number of duplicate reads generated from the detected cluster ("% Duplicate"), the number of reads generated from the detected cluster that mismatch the reference sequence for all reads aligned to the reference sequence ("Mismatched"), the number of reads generated from the detected cluster that are ignored for alignment ("% Soft Clipped") because their portions do not match the reference sequence sufficiently on either side, the number of bases called for the detected cluster that have a quality score of 30 and are above ("% Q30 Bases"), the number of paired reads generated from the detected cluster that are aligned to the inboard reads within a reasonable distance ("Total Correct Read Pairs"), and the number of unique or overlapping correct read pairs generated from the detected cluster ("Non-overlapping Correct Read Pairs").
As shown in FIG. 64, both the binary classification model 4600 and the regression model 2600 perform well with the RTA base caller in generating templates for the majority of metrics.
FIG. 65 compares the performance of the ternary classification model 5400 with that of the RTA-based caller under three conditions, five sequencing metrics, and two run densities.
In the first scenario, called "RTA", the cluster centers are detected by the RTA base caller, intensity extraction from the clusters is performed by the RTA base caller, and the clusters are also base called using the RTA base caller. In the second scenario, called "RTA IE", the cluster centers are detected by the ternary classification model 5400, but intensity extraction from the clusters is performed by the RTA base caller, and the clusters are also base called using the RTA base caller. In the third scenario, called "Self IE", the cluster centers are detected by the ternary classification model 5400, and intensity extraction from the clusters is performed using the cluster shape based intensity extraction technique disclosed herein (note that the cluster shape information is generated by the ternary classification model 5400). However, the clusters are base called using the RTA base caller.
Performance is compared between the ternary classification model 5400 and the RTA base caller along five metrics: (1) total number of clusters detected ("#Clusters"), (2) number of detected clusters passing the chasity filter ("#PF"), (3) number of unique or overlapping suitable read pairs generated from detected clusters ("#Unoverlapping Suitable Read Pairs"), (4) sequence reads generated from detected clusters and reference sequences after alignment ("Mismatch Rate"), and (5) percentage of mismatches between detected clusters with a quality score of 30 ("%Q30").
We compare the performance between the ternary classification model 5400 and the RTA base caller under three conditions and compare five metrics for two types of sequencing runs: (1) a normal run with 20 pM library concentration, and (2) a high-density run with 30 pM library concentration.
As shown in FIG. 65, a 3-way classification model 5400 implements the RTA base caller for all metrics.
FIG. 66 shows that under the same three conditions, five metrics, and two execution densities, regression model 2600 outperforms the RTA base caller for all metrics.
FIG. 67 focuses on the final layer 6702 of the neural network-based template generator 1512.
Figure 68 visualizes what the final layer 6702 of the neural network-based template generator 1512 has learned as a result of backpropagation-based gradient update training. The illustrated embodiment visualizes 24 out of 32 convolution filters of the final layer 6702 overlaid on the ground truth cluster shapes. As shown in Figure 68, the final layer 6702 has learned cluster metadata including the spatial distribution of clusters such as cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and cluster boundaries.
Figure 69 overlays the cluster center predictions of binary classification model 4600 (in blue) on the RTA-based colored ones (in pink). Predictions are made on sequencing image data from an Illumina NextSeq sequencer.
Figure 70 overlays cluster center predictions made by RTA base calling (in pink) on a visualization of the trained convolutional filters of the final layer of a binary classification model 4600. These convolutional filters are learned as a result of sequencing image data from an Illumina NextSeq sequencer.
71 shows one embodiment of the training data used to train the neural network based template generator 1512. In this alternative embodiment, the training data is obtained from a high density flow cell that generates data using Storm probe images. In another embodiment, the training data is obtained from a high density flow cell that generates data with fewer bridge amplification cycles.
FIG. 72 is an example of using beads for image registration based on cluster center predictions in a neural network based template generator 1512.
73 shows one embodiment of cluster statistics for clusters identified by the neural network based template generator 1512. The cluster statistics include cluster size based on the number of contributing subpixels and GC content.
FIG. 74 shows how the ability of the neural network-based template generator 1512 to distinguish between adjacent clusters improves as the number of initial sequencing cycles in which the input image data 1702 is used is increased from 5 to 7. For 5 sequencing cycles, a single cluster is identified by a single discontinuous region of contiguous sub-pixels. For 7 sequencing cycles, the single cluster is split into two adjacent clusters, each with its own discontinuous region of adjacent sub-pixels.
Figure 75 shows the difference in base calling performance when RTA-based color uses ground truth mass (COM) positions as cluster centers as opposed to when non-COM positions are used as cluster centers.
FIG. 76 shows the performance of the neural network-based template generator 1512 on additional detected clusters.
FIG. 77 shows different data sets used to train the neural network-based template generator 1512.
(Sequencing System)
78A and 78B show one embodiment of a sequencing system 7800A. The sequencing system 7800A includes a configurable processor 7846. The configurable processor 7846 implements the base calling techniques disclosed herein. A sequencing system is also referred to as a "sequencer."
The sequencing system 7800A can obtain any information or data related to at least one of a biological material or a chemical. In some embodiments, the sequencing system 7800A is a workstation, which can be similar to a benchtop device or a desktop computer. For example, most (or all) of the systems and components for carrying out the desired reactions can be in a common housing 7802.
In certain embodiments, the sequencing system 7800A is a nucleic acid sequencing system configured for a variety of applications, including, but not limited to, de novo sequencing, whole genome or targeted genomic region resequencing, and metagenomics. The sequencer may also be used for DNA or RNA analysis. In some embodiments, the sequencing system 7800A may also be configured to generate reaction sites within a biosensor. For example, the sequencing system 7800A may be configured to receive a sample and generate surface-bound clusters of clonovirus amplified nucleic acid from the sample. Each cluster may constitute or be part of a reaction site within a biosensor.
The exemplary sequencing system 7800A may include a system receptacle or interface 7810 configured to interact with a biosensor 7812 to effect a desired reaction within the biosensor 7812. In the discussion that follows with respect to FIG. 78A , the biosensor 7812 is loaded into the system receptacle 7810. However, it is understood that a cartridge including the biosensor 7812 may be inserted into the system receptacle 7810, and in some conditions, the cartridge may be temporarily or permanently removed. As discussed above, the cartridge may include, among other things, fluid control and fluid storage components.
In certain embodiments, the sequencing system 7800A is configured to perform multiple parallel reactions within the biosensor 7812. The biosensor 7812 includes one or more reaction sites where a desired reaction can occur. The reaction sites may be immobilized, for example, on a solid surface of the biosensor or on beads (or other movable substrates) located within corresponding reaction chambers of the biosensor. The reaction sites may include, for example, clusters of clonovirus amplified nucleic acids. The biosensor 7812 may include a solid-state imager (e.g., a CCD or CMOS imager) and a flow cell attached thereto. The flow cell may include one or more flow paths that receive solutions from the sequencing system 7800A and direct the solutions toward the reaction sites. Optionally, the biosensor 7812 may be configured to engage a thermal element for transferring thermal energy into and out of the flow paths.
The sequencing system 7800A may include various components, assemblies, and systems (or subsystems) that interact with each other to perform a given method or assay protocol for biological or chemical analysis. For example, the sequencing system 7800A includes a system controller 7806 that may communicate with the various components, assemblies, and subsystems of the sequencing system 7800A and also includes a biosensor 7812. For example, in addition to the system vessel 7810, the sequencing system 7800A may also include a fluid control system 7808 that controls the flow of fluids in the fluidic network and biosensor 7812 of the sequencing system 7800A, a fluid reservoir system 7814 that holds any fluids (e.g., gases or liquids) that may be used by the bioassay system, a temperature control system 7804 that may regulate the temperature of the fluids in the fluidic network, the fluid reservoir system 7814, and/or the biosensor 7812, and an illumination system 7816 configured to illuminate the biosensor 7812. As described above, when a cartridge having a biosensor 7812 is loaded into the system receptacle 7810, the cartridge may also include fluid control and fluid storage components.
The sequencing system 7800A may also include a user interface 7818 for interacting with a user. For example, the user interface 7818 may include a display 7820 for displaying or requesting information from a user, and a user input device 7822 for receiving user input. In some implementations, the display 7820 and the user input device 7822 are the same device. For example, the user interface 7818 may include a touch-sensitive display configured to detect the presence of an individual touch and to identify the location of the touch on the display. However, other user input devices 7822, such as a mouse, touchpad, keyboard, keypad, handheld scanner, voice recognition system, motion recognition system, etc. may also be used. As described in more detail below, the sequencing system 7800A may communicate with various components, including a biosensor 7812 (e.g., in the form of a cartridge), to perform the desired reaction. The sequencing system 7800A may also be configured to analyze data obtained from the biosensor to provide the desired information to the user.
The system controller 7806 comprises a microcontroller, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a coarse grained reconfigurable architecture (CGRAs), a logic circuit, and any other circuit or processor capable of performing the functions described herein. The above examples are merely exemplary and are therefore not intended to limit the definition and/or meaning of the term system controller. In an exemplary embodiment, the system controller 7806 executes a set of instructions stored in one or more storage elements, memories, or modules for at least one of acquiring and analyzing detection data. The detection data can include multiple sequences of pixel signals, whereby sequences of pixel signals from each of millions of sensors (or pixels) can be detected over many base call cycles. The storage elements can be in the form of information sources or physical memory elements in the sequencing system 7800A.
The set of instructions may include various commands that instruct the sequencing system 7800A or biosensor 7812 to perform certain operations, such as the methods and processes of various embodiments described herein. The set of instructions may be in the form of a software program that may form a part of a tangible non-transitory computer readable medium or medium. As used herein, the terms "software" and "firmware" are interchangeable and include any computer program stored in memory that is executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely exemplary and therefore not limiting to the types of memory that can be used to store a computer program.
The software may be in various forms, such as system software or application software. Furthermore, the software may be in the form of a collection of separate programs, or a program module or a portion of a program module within a larger program. The software may also include modular programming in the form of object-oriented programming. After acquiring the detection data, the detection data may be automatically processed by the sequencing system 7800A processed in response to user input, or may be processed in response to a request made by another processing machine (e.g., a remote request via a communication link). In another embodiment shown, the system controller 7806 includes an analysis module 7844. In another embodiment, the system controller 7806 does not include the analysis module 7844, but instead has access to the analysis module 7844 (e.g., the analysis module 7844 may be separately hosted on the cloud).
The system controller 7806 may be connected to the biosensor 7812 and other components of the sequencing system 7800A via a communication link. The system controller 7806 may also be communicatively connected to an off-site system or server. The communication link may be a wire, a cord, or wireless. The system controller 7806 may receive user input or commands from a user interface 7818 and a user input device 7822.
The fluid control system 7808 includes a fluid network and is configured to direct the flow of one or more fluids through the fluid network. The fluid network may be in fluid communication with a biosensor 7812 and a fluid storage system 7814. For example, fluid may be selected from the fluid storage system 7814 and directed to the biosensor 7812 in a controlled manner, or fluid may be drawn from the biosensor 7812 and directed, for example, to a waste reservoir in the fluid storage system 7814. Although not shown, the fluid control system 7808 may include a flow sensor that detects the flow rate or pressure of the fluid in the fluid network. The sensor may be in communication with the system controller 7806.
The temperature control system 7804 is configured to regulate the temperature of fluids in different regions of the fluid network, the fluid reservoir system 7814, and/or the biosensor 7812. For example, the temperature control system 7804 may include a thermal cycler that interacts with the biosensor 7812 and controls the temperature of the fluid flowing along a reaction site within the biosensor 7812. The temperature control system 7804 may also regulate the temperature of solid elements or components of the sequencing system 7800A or the biosensor 7812. Although not shown, the temperature control system 7804 may include sensors for detecting the temperature of the fluids or other components. The sensors may be in communication with the system controller 7806.
The fluid storage system 7814 is in fluid communication with the biosensor 7812 and may store various reaction components or reactants used to carry out a desired reaction. The fluid storage system 7814 may also store fluids for washing or rinsing the fluidic network and the biosensor 7812 and for diluting reactants. For example, the fluid storage system 7814 may include various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, and the like. Additionally, the fluid storage system 7814 may also include a waste reservoir for receiving waste from the biosensor 7812. In embodiments that include a cartridge, the cartridge may include one or more of a fluid storage system, a fluid control system, or a temperature control system. Thus, one or more of the components described herein with respect to these systems may be contained within the cartridge housing. For example, the cartridge may have various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, waste, and the like. Thus, one or more of the fluid reservoir system, fluid control system, or temperature control system may be removably engaged with the bioassay system via a cartridge or other biosensor.
The illumination system 7816 may include a light source (e.g., one or more LEDs) and multiple optical components for illuminating the biosensor. Examples of light sources may include lasers, arc lamps, LEDs, or laser diodes. The optical components may be, for example, reflectors, polarizers, beam splitters, collimators, lenses, filters, wedges, prisms, mirrors, detectors, and the like. In embodiments using an illumination system, the illumination system 7816 may be configured to direct excitation light to the reaction sites. As an example, a fluorophore may be excited by a wavelength of green light, and thus the wavelength of the excitation light may be about 532 nm. In one embodiment, the illumination system 7816 is configured to generate illumination parallel to a surface normal of the surface of the biosensor 7812. In another embodiment, the illumination system 7816 is configured to generate illumination that is off-angled to the surface normal of the surface of the biosensor 7812. In yet another embodiment, the illumination system 7816 is configured to generate illumination having multiple angles, including some parallel illumination and some off-angle illumination.
The system receptacle or interface 7810 is configured to engage the biosensor 7812 in at least one of mechanical, electrical, and fluidic ways. The system receptacle 7810 can hold the biosensor 7812 in a desired orientation to facilitate fluid flow through the biosensor 7812. The system receptacle 7810 can also include electrical contacts configured to engage the biosensor 7812 such that the sequencing system 7800A can communicate with and/or provide power to the biosensor 7812. Additionally, the system receptacle 7810 can include a fluid port (e.g., a nozzle) configured to engage the biosensor 7812. In some embodiments, the biosensor 7812 is removably coupled to the system receptacle 7810 both electrically and fluidically.
In addition, the sequencing system 7800A may communicate remotely with other systems or networks, or with other bioassay systems 7800A. Detection data obtained by the bioassay system 7800A may be stored in a remote database.
FIG. 78B is a block diagram of a system controller 7806 that can be used in the system of FIG. 78A. In one implementation, the system controller 7806 includes one or more processors or modules that can communicate with each other. Each of the processors or modules may include algorithms (e.g., instructions stored on a tangible and/or non-transitory computer-readable storage medium) or sub-algorithms for performing a particular process. The system controller 7806 is conceptually illustrated as a collection of modules, but may be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 7806 may be implemented using an off-the-shelf PC with a single processor or multiple processors, with functional operations distributed among the processors. As a further option, the modules described below may be implemented using a hybrid configuration in which certain modular functions are implemented using dedicated hardware, while the remaining modular functions are implemented using off-the-shelf PCs, etc. The modules may also be implemented as software modules within a processing unit.
In operation, the communication port 7850 may transmit information (e.g., commands) to the biosensor 7812 (FIG. 78A) and/or subsystems 7808, 7814, 7804 (FIG. 78A). In an embodiment, the communication port 7850 may output multiple sequences of pixel signals. The communication link 7834 may receive user input from the user interface 7818 (FIG. 78A) and transmit data or information to the user interface 7818. Data from the biosensor 7812 or subsystems 7808, 7814, 7804 may be processed in real-time by the system controller 7806 during a bioassay session. Additionally or alternatively, data may be temporarily stored in system memory during a bioassay session and processed slower than real-time or offline operation.
As shown in FIG. 78B, the system controller 7806 may include multiple modules 7826-7848 in communication with a main control module 7824 along with a central processing unit (CPU) 7852. The main control module 7824 may communicate with a user interface 7818 (FIG. 78A). Although the modules 7826-7848 are shown in direct communication with the main control module 7824, the modules 7826-7848 may also communicate directly with each other, with the user interface 7818, and with the biosensor 7812. The modules 7826-7848 may also communicate with the main control module 7824 via other modules.
The plurality of modules 7826-7848 includes system modules 7828-7832, 7826 in communication with subsystems 7808, 7814, 7804, and 7816, respectively. The fluid control module 7828 may communicate with the fluid control system 7808 to control valves and flow sensors of the fluid network to control the flow of one or more fluids through the fluid network. The fluid storage module 7830 may notify a user when fluid is low or when a waste reservoir is at or near capacity. The fluid storage module 7830 may also communicate with a temperature control module 7832 so that fluid may be stored at a desired temperature. The illumination module 7826 may communicate with the illumination system 7816 to illuminate the reaction sites at specified times during a protocol, such as after a desired reaction (e.g., a binding event) has occurred. In some embodiments, the illumination module 7826 may communicate with the illumination system 7816 to illuminate the reaction sites at a specified angle.
The plurality of modules 7826-7848 may also include an instrument module 7836 that communicates with the biosensor 7812 and an identification module 7838 that determines identification information associated with the biosensor 7812. The instrument module 7836 may, for example, communicate with the system receptacle 7810 to verify that the biosensor has established electrical and fluidic connection with the sequencing system 7800A. The identification module 7838 may receive a signal that identifies the biosensor 7812. The identification module 7838 may use the identification information of the biosensor 7812 to provide other information to a user. For example, the identification module 7838 may determine and then display a lot number, date of manufacture, or a recommended protocol to be run on the biosensor 7812.
The plurality of modules 7826-7848 also includes an analysis module 7844 (also referred to as a signal processing module or signal processor) that receives and analyzes signal data (e.g., image data) from the biosensor 7812. The analysis module 7844 includes memory (e.g., RAM or flash) for storing the detection/image data. The detection data can include multiple sequences of pixel signals, such that a sequence of pixel signals from each of millions of sensors (or pixels) can be detected over many base call cycles. The signal data can be stored for subsequent analysis or sent to the user interface 7818 to display desired information to a user. In some embodiments, the signal data can be processed by a solid-state imager (e.g., a CMOS image sensor) before the analysis module 7844 receives the signal data.
The analysis module 7844 is configured to acquire image data from the photodetector at each of the plurality of sequencing cycles. The image data is derived from the luminescence signals detected by the photodetector and processes the image data for each of the plurality of sequencing cycles via the neural network based template generator 1512 and/or the neural network based base caller 1514 to generate base calls for at least some of the analytes at each of the plurality of sequencing cycles. The photodetector may be part of one or more overhead cameras (e.g., a CCD camera in an Illumina GAIIx that takes images of the clusters on the biosensor 7812 from above) or may be part of the biosensor 7812 itself (e.g., a CMOS image sensor in an Illumina iSeq that is below the clusters on the biosensor 7812 and takes images of the clusters from the bottom).
The output of the photodetector is a sequence image showing the intensity emissions of each cluster and their surrounding background. The sequence images show the intensity emissions generated as a result of incorporating nucleotides into a sequence during sequencing. The intensity emissions are from the associated analytes and their surrounding background. The sequence images are stored in memory 7848.
Protocol modules 7840 and 7842 communicate with main control module 7824 to control the operation of subsystems 7808, 7814, and 7804 in carrying out a predetermined assay protocol. Protocol modules 7840 and 7842 may include instruction sets for instructing sequencing system 7800A to perform specific operations according to a predetermined protocol. As shown, the protocol module may be a sequence synthesis (SBS) module 7840 configured to issue various commands to carry out a sequence-by-sequence synthesis process. In SBS, the extension of a nucleic acid primer along a nucleic acid template is monitored to determine the sequence of nucleotides in the template. The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase enzyme) or ligation (e.g., catalyzed by a ligase enzyme). In certain polymer-based SBS embodiments, fluorescently labeled nucleotides are added to the primer (thereby extending the primer) in a template-dependent manner such that detection of the order and type of nucleotides added to the primer can be used to determine the sequence of the template. For example, to initiate a first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc. can be delivered into/through a flow cell housing an array of nucleic acid templates. The nucleic acid templates may be located at corresponding reaction sites. These reaction sites can be detected where primer extension allows the incorporated labeled nucleotides to be detected through an imaging event. During the imaging event, the illumination system 7816 can provide excitation light to the reaction sites. Optionally, the nucleotides can further include a reversible termination feature that terminates further primer extension once the nucleotide is added to the primer. For example, a nucleotide analog with a reversible terminator moiety can be added to the primer such that no further extension can occur until a deblocking agent is delivered to remove the moiety. Thus, in another embodiment using a reversible termination, a command can be given to deliver a deblocking reagent to the flow cell (either before or after detection). One or more commands can be given to effect wash(s) between the various delivery steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary sequencing techniques are described, for example, in Bentley et al., Nature 456:53-59 (20078), WO 04/0178497, U.S. Patent No. 7,057,026, WO 91/066778, U.S. Patent No. 07/123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,2781, and U.S. Patent No. 20078/01470780782 (each of which is incorporated herein by reference).
In the nucleotide delivery step of the SBS cycle, any one of a single type of nucleotide can be delivered at a time, or multiple different nucleotide types (e.g., A, C, T, and G) can be delivered. In nucleotide delivery configurations where only a single type of nucleotide is present at a time, different nucleotides do not need to have separate labels because they can be distinguished based on the temporal separation inherent to the individualized delivery. Thus, the sequencing method or device can use a single color detection. For example, the excitation source only needs to provide excitation of a single wavelength or a single wavelength range. In nucleotide delivery configurations where delivery results in multiple different nucleotides being present in the flow cell at a given time, the sites incorporating different nucleotide types can be distinguished based on the different fluorescent labels attached to each nucleotide type in the mixture. For example, four different nucleotides can be used, each with one of four different fluorophores. In one embodiment, the four different fluorophores can be distinguished using excitation in four different regions of the spectrum. For example, four different excitation radiation sources can be used. Alternatively, less than four different excitation sources can be used, but optical filtering of the excitation radiation from a single source can be used to generate a range of different excitation radiation in the flow cell.
In some embodiments, less than four different colors can be detected in a mixture with four different nucleotides. For example, pairs of nucleotides can be detected at the same wavelength, but can be distinguished based on the difference in intensity for one member of the pair, or based on a change to one member of the pair (e.g., via chemical modification, photochemical modification, or physical modification) that causes a distinct signal to appear or disappear compared to the signal detected for the other member of the pair. Exemplary devices and methods for distinguishing four different nucleotides using detection of less than four colors are described, for example, in U.S. Patent No. 61/5378,294 and U.S. Patent No. 61/619,78778, which are incorporated herein by reference in their entirety. U.S. Patent Application No. 13/624,200, filed September 21, 2012, is incorporated by reference in its entirety.
The multiple protocol modules may also include a sample preparation (or generation) module 7842 configured to issue commands to the fluidic control system 7808 and the temperature control system 7804 to amplify the product in the biosensor 7812. For example, the biosensor 7812 may be engaged to a sequencing system 7800A. The amplification module 7842 may issue instructions to the fluidic control system 7808 to deliver the necessary amplification components to a reaction chamber in the biosensor 7812. In other embodiments, the reaction site may already contain some components for amplification, such as template DNA and/or primers. After delivering the amplification components to the reaction chamber, the amplification module 7842 may instruct the temperature control system 7804 to cycle through different temperature steps according to a known amplification protocol. In some embodiments, the amplification and/or incorporation of nucleotides is performed isothermally.
The SBS module 7840 can issue commands to perform a bridge PCR in which a cluster of clonal amplicons is formed over a localized region within the channel of the flow cell. After generating the amplicons via bridge PCR, the amplicons may be "linearized" to create single-stranded template DNA, and sstDNA and sequencing primers may be hybridized to universal sequences flanking the region of interest. For example, a reversible terminator-based sequencing by synthesis method may be used as described above or as follows.
Each base calling or sequencing cycle can extend the sstDNA by a single base, which can be achieved, for example, by using a modified DNA polymerase and a mixture of four types of nucleotides. The different types of nucleotides can have unique fluorescent labels, and each nucleotide can further have a reversible terminator that allows only a single base incorporation to occur in each cycle. After addition of a single base to the sstDNA, excitation light can enter the reaction site and the fluorescent emission can be detected. After detection, the fluorescent label and the terminator can be chemically cleaved from the sstDNA. Another similar base calling or sequencing cycle can be as follows. In such a sequencing protocol, the SBS module 7840 can instruct the fluid control system 7808 to direct the flow of reagents and enzyme solutions through the biosensor 7812. Exemplary reversible terminator-based SBS methods that can be utilized with the devices and methods described herein are described in U.S. Patent Application Publication No. 2007/0166705, U.S. Patent Application Publication No. 2006/017878901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006/0240439, U.S. Patent Application Publication No. 2006/027814714709, WO 05/0657814, U.S. Patent Application Publication No. 2005/014700900, WO 06/078B199, and WO 07/01470251, each of which is incorporated by reference in its entirety. Exemplary reagents for reversible terminator-based SBS are described in U.S. Pat. No. 7,541,444, U.S. Pat. No. 7,057,026, U.S. Pat. No. 7,414,14716, U.S. Pat. No. 7,427,673, U.S. Pat. No. 7,566,537, U.S. Pat. No. 7,592,435, and WO 07/1478353678, each of which is incorporated herein by reference in its entirety.
In some embodiments, the amplification and SBS modules may operate in a single assay protocol, eg, template nucleic acid is amplified and subsequently sequenced within the same cartridge.
The sequencing system 7800A may also allow the user to reconfigure the assay protocol. For example, the determination system 7800A may provide the user with an option through the user interface 7818 to modify the determined protocol. For example, if it is determined that the biosensor 7812 is to be used for amplification, the sequencing system 7800A may request the temperature of the annealing cycle. Additionally, the sequencing system 7800A may issue a warning to the user if the user provides user input that is not generally accepted for the selected assay protocol.
In one embodiment, the biosensor 7812 includes a million sensors (or pixels), each of which generates a sequence of multiple pixel signals over successive base call cycles. The analysis module 7844 detects the multiple sequences of pixel signals and attributes them to corresponding sensors (or pixels) according to the row-wise and/or column-wise positions of the sensors on the array of sensors.
FIG. 79 is a simplified block diagram of a system for analysis of sensor data from a sequencing system 7800A, such as base calling sensor output. In the example of FIG. 79, the system includes a configurable processor 7846. The configurable processor 7846 can execute a base caller (e.g., neural network based template generator 1512 and/or neural network based base caller 1514) in coordination with a run-time program executed by a central processing unit (CPU) 7852 (i.e., a host processor). The sequencing system 7800A includes a biosensor 7812 and a flow cell. The flow cell can include one or more tiles in which clusters of genetic material are exposed to a series of analyte flows that are used to trigger reactions within the clusters to identify bases in the genetic material. A sensor senses reactions for each cycle of sequencing in each tile of the flow cell to provide tile data. Genetic sequencing is a data-intensive operation that converts base call sensor data into a sequence of base calls for each group of genetic material sensed during the base calling operation.
The system of this example includes a CPU 7852 that executes a run-time program for coordinating base calling operations, a memory 7848B that stores sequences of arrays of tile data, base call reads generated by the base calling operations, and other information used in the base calling operations. In this figure, the system also includes a memory 7848A that stores configuration files (or files), such as FPGA bit files, and model parameters of a neural network used to configure and reconfigure the configurable processor 7846. The sequencing system 7800A can include programs for configuring the configurable processor, and in some embodiments, can include a reconfigurable processor that runs a neural network.
The sequencing system 7800A is coupled to the configurable processor 7846 by a bus 7902. The bus 7902 can be implemented using a high throughput technology, such as a bus technology compatible with the PCIe standard (Peripheral Component Interconnect Express) currently maintained and developed by the PCI Special Interest Group (PCI-SIG). Also, in this example, the memory 7848A is coupled to the configurable processor 7846 by a bus 7906. The memory 7848A can be an on-board memory located on a circuit board having the configurable processor 7846. The memory 7848A is used for fast access by the configurable processor 7846 of working data used in base calling operations. The bus 7906 can also be implemented using a high throughput technology, such as a bus technology compatible with the PCIe standard.
Configurable processors, including field programmable gate arrays FPGAs, coarse-grained configurable reconfigurable arrays CGRAs, and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than can be achieved using a general-purpose processor executing a computer program. Configuring a configurable processor involves compiling a functional description to generate a configuration file, sometimes referred to as a bitstream or bitfile, and distributing the configuration file to configurable elements on the processor. The configuration file configures the circuitry to set data flow patterns, including the use of distributed memory and other on-chip memory resources, lookup table contents, the operation of configurable logic blocks, and configurable execution units such as configurable interconnects and other elements of the configurable array. A configuration file is reconfigurable if it can be changed in the field by changing a loaded configuration file. For example, the configuration file may be stored in a volatile SRAM element, in a non-volatile read-write memory element, or distributed among an array of configurable elements on a configurable or reconfigurable processor. A variety of commercially available configurable processors are suitable for use in base calling operations as described herein. Examples include Google's Tensor Processing Unit (TPU), GX4 Rackmount Series, GX9 Rackmount Series, NVIDIA DGX-1, Microsoft's Stratix V FPGA, Graphcore's Intelligent Processor Unit (IPU), Qualcomm's Zeroth Platform (Snapdragon processors), NVIDIA Volta, NVIDIA's Drive PX, NVIDIA's JETSON TX1/TX2 MODULE, Intel's Nirvana, Movidius VPU, Fujitsu DPI, Arm DynamicIQ, IBM TrueNorth, Lambda GPU Server with Testa V100s, Xilinx Alveo U200, Xilinx Alveo U250, Xilinx Alveo U280, Intel/Altera Stratix GX2800, Intel/Altera Stratix GX2800, and Intel Stratix GX10M. In some embodiments, the host CPU may be implemented on the same integrated circuit as the configurable processor.
The embodiments described herein use a configurable processor 7846 to implement the neural network based template generator 1512 and/or the neural network based basis caller 1514. The configuration file for the configurable processor 7846 may be implemented by specifying the logic functions to be performed using a high level description language HDL or a register transfer level RTL language specification. This specification may be compiled using resources designed by a selected configurable processor to generate the configuration file. The same or similar specifications may be compiled for the purpose of generating a design for an application specific integrated circuit that may not be a configurable processor.
Thus, alternatives to the configurable processor 7846 in all embodiments described herein include a configured processor including an application specific ASIC or dedicated integrated circuit or set of integrated circuits, or is a system-on-chip SOC device, or a graphics processing unit (GPU) processor or a coarse-grained reconfigurable architecture (CGRA) processor configured to perform neural network based base call operations as described herein.
In general, the configurable and configured processors described herein that are configured to perform the execution of neural networks are referred to herein as neural network processors.
The configurable processor 7846, in this example, is configured using a program executed by the CPU 7852 or by a configuration file loaded by another source to configure an array of configurable elements 7916 (e.g., configuration logic blocks (CLBs), such as look-up tables (LUTs), flip-flops, arithmetic processing units (PMUs), and computational memory units (CMUs), configurable I/O blocks, programmable interconnects) to perform base calling functions. In this example, the configuration includes data flow logic 7908 coupled to buses 7902 and 7906, which performs the function of distributing data and control parameters among the elements used in the base calling operations.
The configurable processor 7846 is also configured with base calling execution logic 7908 to execute the neural network based template generator 1512 and/or the neural network based base caller 1514. The logic 7908 includes multi-cycle execution clusters (e.g., 7914), which in this example include execution cluster 1 through execution cluster X. The number of multi-cycle execution clusters can be selected according to tradeoffs with the desired throughput of operation and available resources on the configurable processor 7846.
The multi-cycle execution clusters are coupled to the data flow logic 7908 by data paths 7910 implemented using configurable interconnect and memory resources on the configurable processor 7846. The multi-cycle execution clusters are also coupled to the data flow logic 7908 by control paths 7912 implemented using configurable interconnect and memory resources, such as the configurable processor 7846. Provisions include providing control signals indicative of available execution clusters, providing input units to the available execution clusters for execution of the neural network based template generator 1512 and/or the neural network based base caller 1514, providing trained parameters for the neural network based template generator 1512 and/or the neural network based base caller 1514, providing output patches of base call classification data, and other control data used in the execution of the neural network based template generator 1512 and/or the neural network based base caller 1514.
The configurable processor 7846 is configured to execute an execution of the neural network based template generator 1512 and/or the neural network based base caller 1514 using the trained parameters to generate classification data for detection cycles of the base calling operation. Execute the execution of the neural network based template generator 1512 and/or the neural network based base caller 1514 to generate classification data for subject detection cycles of the base calling operation. The execution of the neural network based template generator 1512 and/or the neural network based base caller 1514 operates in a sequence including a number N of arrays of tile data from each detection cycle of the N detection cycles, the N detection cycles providing sensor data for different base calling operations for one base position per operation in the time sequence in the embodiment described herein. Optionally, some of the N detection cycles can be out of the sequence as necessary according to the particular neural network model being executed. The number N can be any number greater than 1. In some embodiments described herein, the N detection cycles represent a set of detection cycles for at least one detection cycle preceding the subject detection cycle and at least one detection cycle following the subject cycle. Embodiments described herein include where the number N is an integer equal to or greater than 5.
The data flow logic 7908 is configured to use an input unit for a given run that includes tile data for an N array of spatially aligned patches to move the tile data and at least some of the trained parameters of the model parameters from the memory 7848A to the configurable processor 7846 for execution of the neural network based template generator 1512 and/or the neural network based base caller 1514. The input unit may be moved by a direct memory access operation in a single DMA operation or in smaller units that move during available time slots in coordination with the execution of the deployed neural network.
The tile data of the sensing cycle described herein can include an array of sensor data having one or more features. For example, the sensor data can include two images that are analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The tile data can also include metadata about the images and the sensor. For example, in an embodiment of a base calling operation, the tile data can include information about the alignment of the images with the clusters, such as distance from center information indicating the distance of each pixel in the array of sensor data from the center of the group of genetic material on the tile.
During execution of the neural network-based template generator 1512 and/or the neural network-based base caller 1514, the tile data may also include data that is generated during execution of the neural network-based template generator 1512 and/or the neural network-based base caller 1514, referred to as intermediate data that is not recalculated but can be recalculated during execution of the neural network-based template generator 1512 and/or the neural network-based base caller 1514. For example, during execution of the neural network-based template generator 1512 and/or the neural network-based base caller 1514, the data flow logic 7908 may write the intermediate data to the memory 7848A in place of the sensor data for a given patch of the array of tile data. Such embodiments are described in more detail below.
As shown, a system for analysis of base calling sensor output is described that includes a memory (e.g., 7848A) accessible by a runtime program that stores tile data including sensor data for tiles from a sensing cycle of a base calling operation. The system also includes a neural network processor, such as a configurable processor 7846, having access to the memory. The neural network processor is configured to perform an execution of the neural network using the trained parameters to generate classification data for the sensing cycle. As described herein, the execution of the neural network operates on a sequence of N arrays of tile data from each sensing cycle of the N sensing cycles that comprise the subject cycle to generate classification data for the subject cycle. Data flow logic 908 is provided to move the tile data and trained parameters from the memory to the neural network processor for execution of the neural network using an input unit including data for the spatially aligned patches of the N arrays from each sensing cycle of the N sensing cycles.
Also described is a system in which a neural network processor has access to a memory and includes a plurality of execution clusters, the execution clusters being configured to execute a neural network. Data flow logic 7908 has access to the memory and executes a cluster in the plurality of execution clusters to provide an input unit of tile data to an available execution cluster in the plurality of execution clusters, the input unit including an input unit including a number N of spatially aligned patches of the array of tile data from a respective sensing cycle, and a subject sensing cycle, causing the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patches of the subject sensing cycle, where N is greater than 1.
FIG. 80 is a simplified diagram showing aspects of a base calling operation, including the functionality of a run-time program executed by a host processor. In this diagram, the output of an image sensor from a flow cell is provided on line 8000 to an image processing thread 8001, which can perform processes on the image, such as aligning and positioning individual tiles in an array of sensor data, and resampling the image, which can be used by a process to calculate a tile cluster mask for each tile in the flow cell, which can be used by a process to identify pixels in the array of sensor data that correspond to clusters of genetic material on the corresponding tile of the flow cell. The output of the image processing thread 8001 is provided on line 8002 to dispatch logic 8010 in the CPU, which is forwarded to a data cache 8004 (e.g., SSD storage) on high speed bus 8003 or on high speed bus 8005, according to the state of the base calling operation, to neural network processor hardware 8020, such as the configurable processor 7846 of FIG. 79. The processed and transformed image can be stored on the data cache 8004 to detect previously used cycles. The hardware 8020 returns the classification data output by the neural network to the dispatch logic 8080, which passes the information to a data cache 8004 or on line 8011 to thread 8002, which can use the classification data to perform base calling and quality score calculations and place the data in a standard format for base called reads. The output of thread 8002, which performs base calling and quality score calculations, is provided on line 8012 to thread 8003, which aggregates the base called reads, performs other operations such as data compression, and writes the resulting base calling output to a specified destination for consumption by the customer.
In some embodiments, the host may include a thread (not shown) that performs final processing of the output of the hardware 8020 supporting the neural network. For example, the hardware 8020 may provide an output of classification data from a final layer of a multi-cluster neural network. The host processor may perform output activation functions, such as a softmax function, over the classification data to populate the data used by the base calling and quality score thread 8002. The host processor may also perform input operations (not shown), such as batch normalization of the tile data before input to the hardware 8020.
FIG. 81 is a simplified diagram of a configurable processor 7846 configuration such as that of FIG. 79. In FIG. 81, the configurable processor 7846 includes an FPGA with multiple high-speed PCIe interfaces. The FPGA is configured with a wrapper 8100 including data flow logic 7908 as described with reference to FIG. 79. The wrapper 8100 manages interfacing and coordinating with a run-time program in the CPU via a CPU communication link 8109 and manages communication with an on-board DRAM 8102 (e.g., memory 7848A) via a DRAM communication link 8110. The data flow logic 7908 in the wrapper 8100 provides patch data obtained by traversing an array of tile data on the on-board DRAM 8102 to the clusters 8101 for a number N of cycles, and obtains and delivers process data 8115 from the clusters 8101 to the on-board DRAM 8102. The wrapper 8100 also manages the transfer of data between the on-board DRAM 8102 and the host memory for both the input array of tile data and the output patch of classification data. The wrapper forwards the patch data on line 8113 to the assigned cluster 8101. The wrapper provides trained parameters such as weights and biases on line 8112 to the cluster 8101 obtained from on-board DRAM 8102. The wrapper provides configuration and control data on line 8111 to the cluster 8101 provided from, or generated in response to, a runtime program on the host via CPU communication link 8109. The cluster can also provide status signals on line 8116 to the wrapper 8100 that are used in conjunction with control signals from the host to provide spatially aligned patch data and to run a multi-cycle neural network on the patch data using the resources of the cluster 8101.
As described above, there may be multiple clusters on a single configurable processor managed by a wrapper 8100 configured to run on corresponding ones of the multiple patches of tile data. Each cluster may be configured to provide classification data for base calls in a subject detection cycle using the tile data of multiple sensing cycles as described herein.
In an example system, model data including kernel data such as filter weights and biases can be sent from the host CPU to the configurable processor, so that the model can be updated as a function of cycle number. The base calling operation can include, in a representative example, on the order of hundreds of sensing cycles. The base calling operation can include paired end reads in some embodiments. For example, the model trained parameters may be updated every 20 cycles (or other number of cycles) or according to an update pattern implemented in the particular system and neural network model. In some embodiments, where a sequence for a given string in a genetic cluster on a tile includes paired end reads including a first portion extending from (or up) a first end of the string and a second portion extending up (or down) from a second end of the string, the trained parameters can be updated at the transition from the first portion to the second portion.
In some implementations, image data of multiple cycles of sensor data for a tile may be sent from the CPU to the wrapper 8100. The wrapper 8100 may optionally perform some pre-processing and conversion of the sensor data and write the information to the on-board DRAM 8102. The input tile data for each sensing cycle may include an array of sensor data including 4000x3000 pixels/tile or more per tile, with two features representing the colors of the two images of the tile, and including one or two bytes per pixel. In an embodiment where the number N is three sensing cycles used in each implementation of the multi-cycle neural network, the array of tile data for each implementation of the multi-cycle neural network may consume several hundred megabytes per number. In some embodiments of the system, the tile data also includes an array of DFC data stored once per tile, or other types of metadata about the sensor data and the tile.
In operation, if a multi-cycle cluster is available, the wrapper assigns the patch to the cluster. The wrapper fetches the next patch of tile data for the cross section of the tile and sends it to the assigned cluster along with the appropriate control and configuration information. The cluster can be configured with enough memory on the configurable processor to have enough memory to hold the patch of data, including the patch, from multiple cycles in some systems being processed in place, and in various embodiments are processed using ping-pong buffer techniques or raster scan techniques.
When the assigned cluster completes its operation of the neural network of the current patch and generates an output patch, it signals the wrapper. The wrapper either reads the output patch from the assigned cluster or the assigned cluster pushes the data to the wrapper. The wrapper then assembles the output patch for the processed tile in DRAM 8102. Once the processing of the entire tile is complete and the output patch of data is transferred to DRAM, the wrapper sends the processed output array back to the host/CPU in a specific format. In some embodiments, the on-board DRAM 8102 is managed by memory management logic in the wrapper 8100. The runtime program can control the sequencing operations to complete the analysis of the array of all tile data for every cycle executed in a continuous flow to provide real-time analysis.
(Technical Improvements and Terminology)
Base calling involves incorporating or attaching a fluorescently labeled tag with the analyte. The analyte may be a nucleotide or oligonucleotide, and the tag may be a specific nucleotide type (A, C, T, or G). Excitation light is directed at the tagged analyte, and the tag emits a detectable fluorescent signal or intensity emission. The intensity emission indicates the photons emitted by the excitation tag chemically bound to the analyte.
Throughout this application, including the claims, "when used, images, image data, or image regions showing intensity emissions of analytes and their surrounding background, they refer to the intensity emissions of tags attached to the analytes. Those skilled in the art will understand that the intensity emissions of an attached tag represent or correspond to the intensity emissions of the analyte to which the tag is attached, and are therefore used interchangeably. Similarly, a characteristic of an analyte refers to a characteristic of the tag attached to the analyte, or the intensity emissions from the attached tag. For example, the center of the analyte refers to the center of the intensity emissions emitted by the tag attached to the analyte. In another example, the background around the analyte refers to the background around the intensity emissions emitted by the tag attached to the analyte.
All literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and web pages, are expressly incorporated by reference in their entirety. In the event that one or more of the incorporated literature and similar materials differs from or conflicts with this application, including but not limited to defined terms, term usage, techniques described, etc., this application controls.
The disclosed technology uses neural networks to improve the quality and quantity of nucleic acid sequence information that can be obtained from a nucleic acid template or its complement, e.g., a nucleic acid sample, such as a DNA or RNA polynucleotide or other nucleic acid sample. Thus, certain implementations of the disclosed technology provide higher throughput polynucleotide sequencing, e.g., higher rates of collection of DNA or RNA sequence data, greater efficiency in collecting sequence data, and/or lower costs of obtaining such sequence data, as compared to previously available methods.
The disclosed technology uses neural networks to identify centers of solid-phase nucleic acid clusters and analyze optical signals generated during sequencing of such clusters to unambiguously distinguish between adjacent, neighboring, or overlapping clusters and assign sequencing signals to single, discrete source clusters. These and related embodiments thus enable the recovery of meaningful information, such as sequence data, from regions of high-density cluster arrays where useful information may not have been previously obtained from such regions due to confounding effects of overlapping or closely spaced neighboring clusters, including the effects of overlapping signals (e.g., as used in nucleic acid sequencing).
As described in more detail below, in certain embodiments, a composition is provided that comprises a solid support immobilized with one or more nucleic acid clusters as provided herein.Each cluster comprises multiple immobilized nucleic acids of the same sequence, and has a distinguishable center with a detectable central label as provided herein, and the distinguishable center is distinguishable from the nucleic acids immobilized in the surrounding area within the cluster.Also described herein are methods for making and using such clusters with distinguishable centers.
Embodiments of the present disclosure will find use in many contexts where benefits derive from the ability to identify, determine, annotate, record, or otherwise assign the location of a substantially central location within a cluster, such as in high throughput nucleic acid sequencing, development of image analysis algorithms for assigning optical or other signals to individual source clusters, and other applications where recognition of the centers of immobilized nucleic acid clusters is desirable and beneficial.
In certain embodiments, the present invention contemplates methods related to high-throughput nucleic acid analysis, such as nucleic acid sequencing (e.g., "sequencing"). Exemplary high-throughput nucleic acid analysis includes, but is not limited to, de novo sequencing, resequencing, whole genome sequencing, gene expression analysis, gene expression monitoring, epigenetics analysis, genome methylation analysis, allele-specific primer extension (APSE), genetic diversity profiling, whole genome polymorphism discovery and analysis, single nucleotide polymorphism analysis, hybridization-based sequencing, and the like. Those skilled in the art will appreciate that a variety of different nucleic acids can be analyzed using the methods and compositions of the present invention.
Although the implementations of the present invention are described in the context of nucleic acid sequencing, they are applicable in any field where image data acquired at different times, spatial locations, or other temporal or physical aspects are analyzed. For example, the methods and systems described herein are useful in the fields of molecular and cell biology, where image data from microarrays, biological specimens, cells, organisms, etc. are acquired and analyzed at different times or perspectives. Images can be obtained using any number of techniques known in the art, including, but not limited to, fluorescent microscopy, optical microscopy, confocal microscopy, optical imaging, magnetic resonance imaging, tomographic scanning, etc. As another example, the methods and systems described herein can be applied when image data acquired by surveillance, aerial, or satellite imaging techniques, etc. are acquired and analyzed at different times or perspectives. The methods and systems are particularly useful for analyzing images acquired in a field of view, where the observed specimens remain in the same location relative to each other in the field of view. However, specimens may have different properties in separate images, e.g., specimens may appear different in separate images of the field of view. For example, the analyte may appear to be different in color for a given analyte detected in different images, may indicate a change in the intensity of the signal detected for a given analyte in different images, or even the appearance of a signal for a given analyte in one image and the disappearance of the analyte signal in another image.
The examples described herein may be used in various biological or chemical processes and systems for academic or commercial analysis. More specifically, the examples described herein may be used in various processes and systems in which it is desirable to detect an event, characteristic, quality, or property indicative of a specified reaction. For example, the examples described herein include optical detection devices, biosensors, and components thereof, as well as bioassay systems that operate with biosensors. In some embodiments, the devices, biosensors, and systems may include a flow cell and one or more optical sensors coupled (removably or fixedly) together in a substantially monolithic structure.
The devices, biosensors, and bioassay systems may be configured to perform multiple designated reactions that may be detected individually or collectively. The devices, biosensors, and bioassay systems may be configured to perform multiple cycles in which multiple designated reactions occur in parallel. For example, the devices, biosensors, and bioassay systems may be used to array high-density arrays of DNA features through repeated cycles of enzymatic manipulation and light or image detection/capture. Thus, the devices, biosensors, and bioassay systems (e.g., via one or more cartridges) may include one or more microfluidic channels that deliver reagents or other reaction components into the reaction solution, biosensors, and bioassay systems. In some examples, the reaction solution may be substantially acidic, such as comprising a pH of about 5 or less, or about 4 or less, or about 3 or less. In some other examples, the reaction solution may be substantially alkaline/basic, such as comprising a pH of about 8 or more, or about 9 or more, or about 10 or more. As used herein, the term "acidic" and grammatical variants thereof refer to a pH value less than about 7, and the terms "basic," "alkaline," and grammatical variants thereof refer to a pH value greater than about 7.
In some embodiments, the reaction sites are provided or spaced in a predetermined manner, such as a uniform or repeating pattern. In some other embodiments, the reaction sites are randomly distributed. Each of the reaction sites can be associated with one or more light guides and one or more light sensors that detect light from the associated reaction site. In some embodiments, the reaction sites are located within a reaction recess or chamber that can at least partially compartmentalize a designated reaction.
As used herein, a "designated reaction" includes a change in at least one of the chemical, electrical, physical, or optical properties (or qualities) of a chemical or biological substance of interest, such as an analyte of interest. In certain examples, the designated reaction is a positive binding event, such as the incorporation of a fluorescently labeled biomolecule with a fluorescently labeled biomolecule of interest. More generally, the designated reaction may be a chemical conversion, chemical change, or chemical interaction. The designated reaction may also be a change in an electrical property. In certain examples, the designated reaction includes the incorporation of an analyte with a fluorescently labeled molecule. The analyte may be an oligonucleotide, and the fluorescently labeled molecule may be a nucleotide. The designated reaction may be detected when excitation light is directed to an oligonucleotide with a labeled nucleotide, and the fluorophore emits a detectable fluorescent signal. In alternative examples, the detected fluorescence is the result of chemiluminescence or bioluminescence. A specified reaction can also, for example, increase Fluorescence (or Forster) Resonance Energy Transfer (FRET) by bringing a donor fluorophore into close proximity with an acceptor fluorophore, decrease FRET by separating the donor and acceptor fluorophores, increase fluorescence by separating a quencher from a fluorophore, or decrease fluorescence by colocalizing a quencher and a fluorophore.
As used herein, a "reaction solution," "reaction component," or "reactant" includes any material that may be used to obtain at least one specified reaction. For example, potential reaction components include, for example, reagents, enzymes, samples, other biomolecules, and buffers. A reaction component may be delivered to a reaction site in solution and/or immobilized at a reaction site. A reaction component may directly or indirectly interact with another material, such as an analyte of interest immobilized at a reaction site. As noted above, a reaction solution may be substantially acidic (i.e., includes relatively high acidity) (e.g., includes a pH of about 5 or less, a pH of about 4 or less), or a pH of about 3 or less, or substantially alkaline/basic (i.e., includes relatively high alkaline/basicity) (e.g., includes a pH of about 8 or more, a pH of about 9 or more, or a pH of about 10 or more).
As used herein, the term "reaction site" is a localized area where at least one designated reaction can occur. A reaction site may include a support surface of a reaction structure or substrate on which a substance can be immobilized. For example, a reaction site may include a surface of a reaction structure (which may be disposed within a channel of a flow cell) having reaction components thereon, such as colonies of nucleic acids thereon. In some such examples, the nucleic acids in the colonies have the same sequence, e.g., are clonal copies of a single-stranded or double-stranded template. However, in some examples, a reaction site may contain only a single nucleic acid molecule, e.g., in single-stranded or double-stranded form.
The reaction sites may be randomly distributed along the reaction structure or may be arranged in a predetermined manner (e.g., in parallel in a matrix such as a microarray). The reaction sites may also include reaction chambers or recesses that at least partially define a spatial region or volume configured to compartmentalize a designated reaction. As used herein, the term "reaction chamber" or "reaction recess" includes a defined spatial region of a support structure (often in fluid communication with a flow path). The reaction recess may be at least partially isolated from the surrounding environment or spatial region. For example, the reaction recesses may be separated from each other by a shared wall, such as a detection surface. As a more specific example, the reaction recess may be a nanocell that includes a recess, well, groove, cavity, or depression defined by an inner surface of the detection surface, and may have an opening or aperture (i.e., an open side) so that the nanocell can be in fluid communication with the flow path.
In some embodiments, the reaction recesses of the reaction structure are sized and shaped relative to a solid (including a semi-solid) such that the solid can be fully or partially inserted therein. For example, the reaction recesses may be sized and shaped to accommodate a capture bead. The capture bead may have clonovirus amplified DNA or other material thereon. Alternatively, the reaction recesses may be sized and shaped to receive an approximate number of beads or solid substrates. As another example, the reaction recesses may be filled with a porous gel or material configured to control diffusion or filter fluids or solutions that may flow into the reaction recesses.
In some examples, a light sensor (e.g., a photodiode) is associated with a corresponding reaction site. The light sensor associated with a reaction site is configured to detect light emission from the associated reaction site via at least one light guide when a designated reaction occurs at the associated reaction site. In some cases, multiple light sensors (e.g., several pixels of a light detection or camera device) may be associated with a single reaction site. In other cases, a single light sensor (e.g., a single pixel) may be associated with a single reaction site or with a group of reaction sites. The light sensor, reaction site, and other features of the biosensor may be configured such that at least a portion of the light is directly detected by the light sensor without being reflected.
As used herein, "biological or chemical" includes biomolecules, subject samples, subject analytes, and other chemical compounds. Biological or chemical substances may be used to detect, identify, or analyze other chemical compounds, or act as intermediaries to study or analyze other chemical compounds. In certain examples, biological or chemical substances include biomolecules. As used herein, "biomolecules" include at least one of biopolymers, nucleotides, nucleic acids, polynucleotides, oligonucleotides, proteins, enzymes, polypeptides, antibodies, antigens, ligands, receptors, polysaccharides, carbohydrates, polyphosphates, cells, tissues, organisms, or fragments thereof, or any other biologically active chemical compounds, such as analogs or mimetics of the aforementioned species. In further examples, the biological or chemical substances or biomolecules detect the products of another reaction, such as an enzyme or reagent, e.g., the enzyme or reagent used to detect pyrophosphate in a pyrosequencing reaction. Enzymes and reagents useful for pyrophosphate detection are described, for example, in U.S. Patent Publication No. 2005/0244870(A1), which is incorporated by reference in its entirety.
The biomolecules, samples, and biological materials or chemicals may be naturally occurring or synthetic and may be suspended in a solution or mixture within the reaction wells or regions. The biomolecules, samples, and biological materials or chemicals may also be bound to a solid phase or gel material. The biomolecules, samples, and biological materials or chemicals may also include pharmaceutical compositions. In some cases, the biomolecules, samples, and biological materials or chemicals of interest may be referred to as targets, probes, or analytes.
As used herein, a "biosensor" includes a device that includes a reaction structure having a plurality of reaction sites configured to detect a designated reaction occurring at or near the reaction site. The biosensor may include a solid-state photodetector or "imaging" device (e.g., a CCD or CMOS photodetector device) and, optionally, a flow cell attached thereto. The flow cell may include at least one flow path in fluid communication with the reaction site. As one particular example, the biosensor is configured to fluidly and electrically couple to a biological assay system. The bioassay system may deliver reaction solutions to the reaction sites according to a predetermined protocol (e.g., sequence number synthesis) and perform a plurality of imaging events. For example, the bioassay system may flow the reaction solutions along the reaction sites. At least one of the reaction solutions may include four types of nucleotides with the same or different fluorescent labels. The nucleotides may bind to corresponding oligonucleotides, etc., in the reaction sites. The bioassay system may then illuminate the reaction sites using an excitation light source (e.g., a solid-state light source such as a light emitting diode (LED)). The excitation light may have a predetermined wavelength or wavelengths including a range of wavelengths. Fluorescent labels excited by incident excitation light can provide an emission signal (e.g., of a different wavelength or wavelengths of light than the excitation light, and potentially different from each other) that can be detected by a photosensor.
As used herein, the term "immobilized" when used in reference to a biomolecule or biological substance or chemical includes substantially attaching the biomolecule or biological substance or chemical to a surface, such as the detection surface of an optical detection device or a reaction structure. For example, the biomolecule or biological substance or chemical may be immobilized to the surface of a reaction structure using adsorption techniques including non-covalent bonding (e.g., electrostatic forces, van der Waals, and hydrophobic interfacial dehydration), as well as covalent bonding techniques in which a functional group or linker facilitates binding of the biomolecule to the surface. Immobilizing the biomolecule or biological substance or chemical to a surface may be based on the properties of the surface, the liquid medium carrying the biomolecule or biological substance or chemical, and the properties of the biomolecule or biological substance or chemical itself. In some cases, the surface may be functionalized (e.g., chemically or physically modified) to facilitate immobilization of the biomolecule (or biological substance or chemical) to the surface.
In some examples, the nucleic acid can be immobilized on a reaction structure, such as a surface of the reaction well. In certain examples, the devices, biosensors, bioassay systems and methods described herein may include the use of naturally occurring nucleotides and enzymes configured to interact with the naturally occurring nucleotides. Naturally occurring nucleotides include, for example, ribonucleotides or deoxyribonucleotides. Naturally occurring nucleotides can be in monophosphate, diphosphate, or triphosphate form and can have a base selected from adenine (A), thymine (T), uracil (U), guanine (G), or cytosine (C). However, it will be understood that non-naturally occurring nucleotides, modified nucleotides, or analogs of the above nucleotides can be used.
As described above, biomolecules or biological substances or chemicals may be immobilized at reaction sites within the reaction recesses of the reaction structure. Such biomolecules or biological substances may be physically held or immobilized within the reaction recesses by interference fitting, adhesion, covalent bonding, or entrapment. Examples of articles or solids that may be placed within the reaction recesses include polymer beads, pellets, agarose gels, powders, quantum dots, or other solids that may be compressed and/or held within the reaction chamber. In certain embodiments, the reaction recesses may be coated or filled with a hydrogel layer that may be covalently bonded to DNA oligonucleotides. In certain examples, nucleic acid superstructures such as DNA balls may be placed within or in the reaction recesses, for example, by attaching them to the inner surface of the reaction recesses or by being suspended in a liquid within the reaction recesses. DNA balls or other nucleic acid superstructures may be implemented and then placed within or in the reaction recesses. Alternatively, DNA balls may be synthesized in situ in the reaction recesses. The substances immobilized within the reaction recesses may be in a solid, liquid, or gas state.
As used herein, the term "analyte" is intended to mean a point or region of a pattern that can be distinguished . An individual analyte can include one or more molecules of a particular type. For example, an analyte can include a single target nucleic acid molecule with a particular sequence, or an analyte can include several nucleic acid molecules with the same sequence (and/or its complementary sequence). Different molecules that are different analytes of a pattern can be differentiated from each other according to the location of the analyte within the pattern. Exemplary analytes include wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, pads of gel material on a substrate, or channels in a substrate.
Any of a variety of target analytes to be detected, characterized, or identified can be used in the devices, systems, or methods described herein. Exemplary analytes include, but are not limited to, nucleic acids (e.g., DNA, RNA, or analogs thereof), proteins, polysaccharides, cells, antibodies, epitopes, receptors, ligands, enzymes (e.g., kinases, phosphatases, or polymerases), small molecule drug candidates, cells, viruses, organisms, and the like.
The terms "analyte", "nucleic acid", "nucleic acid molecule", and "polynucleotide" are used interchangeably herein. In various embodiments, a nucleic acid may be used as a template (e.g., a nucleic acid template, or a nucleic acid complement complementary to a nucleic acid template) as provided herein for certain types of nucleic acid analysis, including, but not limited to, nucleic acid amplification, nucleic acid expression analysis, and/or nucleic acid sequencing, or suitable combinations thereof. Nucleic acids in certain implementations include, for example, linear polymers of deoxyribonucleotides in 3'-5' phosphodiester, or deoxyribonucleic acid (DNA), such as single-stranded and double-stranded DNA, genomic DNA, copy DNA or complementary DNA (cDNA), recombinant DNA, or any form of synthetic or modified DNA. In other embodiments, the nucleic acid may be, for example, a linear polymer of ribonucleotides in 3'-5' phosphodiester or other linkages such as ribonucleic acid (RNA), e.g., single-stranded and double-stranded RNA, messenger (mRNA), copy RNA or complementary RNA (cRNA), or spliced mRNA, ribosomal RNA, small nuclear RNA (snoRNA), microRNA (miRNA), small interfering RNA (sRNA), piRNA (piRNA), or any form of synthetic or modified RNA. The nucleic acids used in the compositions and methods of the invention may vary in length and may be intact or full-length molecules or fragments, or smaller portions of larger nucleic acid molecules. In certain embodiments, the nucleic acid may bear one or more detectable labels, as described elsewhere herein.
The terms "specimen", "cluster", "nucleic acid cluster", "nucleic acid colony", and "DNA cluster" are used interchangeably and refer to multiple copies of a nucleic acid template and/or its complement attached to a solid support. Typically, in certain preferred embodiments, a nucleic acid cluster comprises multiple copies of a template nucleic acid and/or its complement attached to a solid support via their 5' ends. The copies of the nucleic acid strands that make up a nucleic acid cluster may be in single-stranded or double-stranded form. The copies of the nucleic acid template present within a cluster may have nucleotides at corresponding positions that differ from each other due to, for example, the presence of a label moiety. The corresponding positions may also include analog structures with different chemical structures but similar Watson-Crick base pairing properties, such as in the case of uracil and thymine.
Colonies of nucleic acids may also be referred to as "nucleic acid clusters." Nucleic acid colonies may optionally be generated by cluster amplification or bridge amplification techniques, as described in more detail elsewhere herein. Multiple repeats of a target sequence may be present in a single nucleic acid molecule, such as a disruptor generated using a rolling circle amplification procedure.
The nucleic acid clusters of the present invention can have different shapes, sizes, and densities depending on the conditions used. For example, the clusters can have a substantially circular, multi-sided, donut-shaped, or ring-shaped shape. The diameter of the nucleic acid clusters can be designed to be about 0.2 μm to about 6 μm, about 0.3 μm to about 4 μm, about 0.4 μm to about 3 μm, about 0.5 μm to about 2 μm, about 0.75 μm to about 1.5 μm, or any intervening diameter. In certain embodiments, the diameter of the nucleic acid clusters is about 0.5 μm, about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The diameter of the nucleic acid clusters can be influenced by a number of parameters, including, but not limited to, the number of amplification cycles performed in the production of the clusters, the length of the nucleic acid template, or the density of the primers attached to the surface on which the clusters are formed. The density of the nucleic acid clusters is typically less than 0.1/mm<sup>2</sup>, 1/mm<sup>2</sup>, 10/mm<sup>2</sup>, 100/mm<sup>2</sup>, 1,000/mm<sup>2</sup>, 10,000/mm<sup>2</sup>~100,000/mm<sup>2</sup>The present invention, in part, can be designed to produce higher density nucleic acid clusters, e.g., 100,000/mm<sup>2</sup>~1,000,000/mm<sup>2</sup>, and 1,000,000/mm<sup>2</sup>~10,000,000/mm<sup>2</sup>The following are further intended:
As used herein, an "analyte" is a specimen or region of interest within a field of view. When used in connection with a microarray device or other molecular analysis device, an analyte refers to a region occupied by similar or identical molecules. For example, an analyte can be an amplified oligonucleotide, or any other group of polynucleotides or polypeptides having the same or similar sequence. In other embodiments, an analyte can be any element or group of elements that occupies a physical area on a sample. For example, an analyte can be a parcel of land, a body of water, etc. When analytes are imaged, each analyte has some area. Thus, in many embodiments, an analyte is not simply a pixel.
The distance between the analytes can be described in any number of ways. In some embodiments, the distance between the analytes can be described from the center of one analyte to the center of another analyte. In other embodiments, the distance can be described from the edge of one analyte to the edge of another analyte, or between the outermost identifiable points of each analyte. The edge of the analyte can be described as a theoretical or actual physical boundary on the chip, or some point within the boundary of the analyte. In other embodiments, the distance can be described with respect to a fixed point on the sample, or an image of the sample.
Generally, some embodiments are described herein with respect to the analysis method. It will be understood that a system for performing the method in an automated or semi-automated manner is also provided. Thus, the present disclosure provides a neural network-based template generation and base calling system, the system can include a processor, a storage device, and a program for image analysis, the program includes instructions for performing one or more of the methods described herein. Thus, the methods described herein can be performed, for example, on a computer having components described herein or known in the art.
The methods and systems described herein are useful for analyzing any of a variety of objects. Particularly useful objects are solid supports or solid surfaces with analytes attached. The methods and systems described herein provide advantages when used with objects having repeating patterns of analytes in the xy plane. One example is a microarray having a collection of cells, viruses, nucleic acids, proteins, antibodies, carbohydrates, small molecules (such as drug candidates), biologically active molecules, or other analytes of interest.
There has been an increase in the number of applications of arrays with analytes having biological molecules such as nucleic acids and polypeptides. Such microarrays typically contain deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) probes. These are specific for nucleotide sequences present in humans and other organisms. In certain applications, for example, individual DNA or RNA probes can be attached to individual analytes of the array. Test samples, such as those from known humans or organisms, can be exposed to the array such that the target nucleic acids (e.g., gene fragments, mRNA, or amplicons) hybridize to complementary probes at each analyte in the array. The probes can be labeled by target-specific processes (e.g., due to a label present on the target nucleic acid or due to an enzyme label of the probe or target present in hybridized form in the analyte). They can then be examined by scanning a specific frequency of light over the analytes to identify which target nucleic acids are present in the sample.
Biological microarrays can be used for gene sequencing and similar applications. In general, gene sequencing involves determining the order of nucleotides in a length of target nucleic acid, such as a fragment of DNA or RNA. A relatively short sequence is typically sequenced in each analyte, and the resulting sequence information may be used in a variety of bioinformatics methods to reliably determine the sequence of many broad lengths of genetic material from which the fragments are derived. Automated computer-based algorithms of signature fragments have been developed and have been used more recently in genome mapping, identification of genes and their functions, and the like. Microarrays are particularly useful for characterizing genome content, since there are many variants, which is an alternative to performing many experiments on individual probes and targets. Microarrays are an ideal format for carrying out such studies in a practical manner.
Any of a variety of analyte arrays (also called "microarrays") known in the art can be used in the methods or systems described herein. A typical array includes analytes, each with an individual probe or population of probes. In the latter case, the population of probes in each analyte is typically homogenous, with a single type of probe. For example, in the case of nucleic acid sequences, each analyte can have multiple nucleic acid molecules, each with a common sequence. However, in some embodiments, the population in each analyte of the array can be heterogeneous. Similarly, protein sequences can have analytes with a single protein or population of proteins, typically, but not necessarily, with the same amino acid sequence. The probes can be attached to the surface of the array, for example, by covalently binding the probes to the surface, or through non-covalent interaction(s) between the probes and the surface. In some embodiments, the probes, such as nucleic acid molecules, can be attached to the surface via a gel layer, for example, as described in U.S. Patent Application Serial No. 13/784,368 and U.S. Patent Application Serial No. 2011/0059865(A1), each of which is incorporated herein by reference.
Exemplary arrays include, but are not limited to, BeadChip arrays available from Illumina, Inc. (San Diego, Calif.) or others such as those described below in which probes are attached to beads present on a surface (e.g., beads in wells on a surface). U.S. Pat. Nos. 6,266,459, 6,355,431, 6,770,441, 6,859,570, or 7,622,294, or PCT Publication WO 00/63437, each of which is incorporated herein by reference. Further examples of commercially available microarrays that can be used include, for example, VLSIPS (Very Large Scale Immobilized Polymer Examples of suitable spotted microarrays include Affymetrix® GeneChip® microarrays or other microarrays synthesized according to a technology sometimes referred to as the 'Synthesis' technology. Spotted microarrays can also be used in the methods or systems according to some embodiments of the present disclosure. An exemplary spotted microarray is the CodeLink Array available from Amersham Biosciences. Another useful microarray is one manufactured using inkjet printing techniques, such as the SurePrint Technology available from Agilent Technologies.
Other useful sequences include those used in nucleic acid sequencing applications. For example, sequences with amplicons of genomic fragments (often called clusters) are particularly useful, such as those described in Bentley et al., Nature 456:53-59 (2008), WO 04/018497, WO 91/06678, WO 07/123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281, or U.S. Patent No. 7,057,026, or U.S. Patent Application Publication No. 2008/0108082, each of which is incorporated herein by reference. Another type of sequence useful for nucleic acid sequencing is the sequence of particles generated from emulsion PCR technology. Examples are described in Dressman et al., Proc. Natl. Acad. Sci. USA, 1999, 10, 1411-1423, 1999, 10, 1412-1423 ... 100:8817-8822 (2003), WO 05/010145, U.S. Patent Application Publication No. 2005/0130173, or U.S. Patent Application Publication No. 2005/0064460, each of which is incorporated herein by reference in its entirety.
Arrays used for nucleic acid sequencing often have random spatial patterns of nucleic acid analytes. For example, the HiSeq or MiSeq sequencing platforms available from Illumina Inc (San Diego, Calif.) utilize flow cells in which nucleic acid sequences are formed by random seeding followed by bridge amplification. However, patterned arrays can also be used for nucleic acid sequencing or other analytical applications. Examples of patterned arrays, their use and methods of use are described in U.S. Patent Application Nos. 13/787,396, 13/783,043, 13/784,368, U.S. Patent Application Publication Nos. 2013/0116153(A1), and 2012/0316086(A1), each of which is incorporated herein by reference. Analytes in such patterned arrays can be used to capture single nucleic acid template molecules for subsequent formation of homogenous colonies, for example, via bridge amplification. Such patterned arrays are particularly useful for nucleic acid sequencing applications.
The size of the analytes on an array (or other object used in the methods or systems herein) can be selected to suit a particular application. For example, in some embodiments, the analytes of the array can have a size that accommodates only a single nucleic acid molecule. A surface with multiple analytes in this size range is useful for constructing an array of molecules for detection with single molecule resolution. Analytes in this size range are also useful for use in arrays with analytes each comprising a colony of nucleic acid molecules. Thus, the analytes of the array each have a size of about 1 mm.<sup>2</sup>Below, about 500 μm<sup>2</sup>Below, about 100 μm<sup>2</sup>Below, about 10 μm<sup>2</sup>Below, about 1 μm<sup>2</sup>Below, about 500 nm<sup>2</sup>Less than or equal to about 100 nm<sup>2</sup>Below, about 10 nm<sup>2</sup>Below, about 5 nm<sup>2</sup>Less than or equal to 1 nm<sup>2</sup>Alternatively or additionally, the specimens in the array can have an area of about 1 mm<sup>2</sup>Above 500μm<sup>2</sup>More than 100μm<sup>2</sup>Above 10μm<sup>2</sup>Above 1μm<sup>2</sup>Above 500 nm<sup>2</sup>More than 100 nm<sup>2</sup>More than 10 nm<sup>2</sup>More than 5 nm<sup>2</sup>More than or about 1 nm<sup>2</sup>And that's it. In practice, the analyte can have a size within a range between upper and lower limits selected from those exemplified above. Although some size ranges of surface analytes have been exemplified with respect to nucleic acids and nucleic acid scales, it will be understood that analytes in these size ranges can be used in applications that do not involve nucleic acids. It will be further understood that the size of the analyte need not necessarily be limited to the scale used for nucleic acid applications.
In embodiments involving objects with multiple analytes, such as arrays of analytes, the analytes can be distinct, separated by a space between each other. Arrays useful in the present invention can have analytes separated by an edge-to-edge distance of at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm or less. Alternatively or additionally, arrays can have analytes separated by an edge-to-edge distance of at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can apply to the average edge-to-edge spacing and edge-to-edge spacing of the analytes, as well as the minimum or maximum spacing.
In some embodiments, the analytes of the array need not be distinct, instead, adjacent analytes can abut each other. Whether the analytes are distinct or not, the size of the analytes and/or the pitch of the analytes can vary so that the array can have a desired density. For example, the average analyte pitch in a regular pattern can be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm or less. Alternatively or additionally, the average analyte pitch in a regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can also apply to the maximum or minimum pitch of the regular pattern. For example, the maximum analyte pitch in the regular pattern can be 100 μm or less, 50 μm or less, 10 μm or less, 5 μm or less, 1 μm or less, 0.5 μm or less, and/or the minimum analyte pitch in the regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more.
The density of analytes in an array can also be understood in terms of the number of analytes present per unit area. For example, the average density of analytes on an array is at least about 1×10<sup>3</sup>Samples/mm<sup>2</sup>、1×10<sup>4</sup>Samples/mm<sup>2</sup>、1×10<sup>5</sup>Samples/mm<sup>2</sup>、1×10<sup>6</sup>Samples/mm<sup>2</sup>、1×10<sup>6</sup>Samples/mm<sup>2</sup>、1×10<sup>7</sup>Samples/mm<sup>2</sup>、1×10<sup>8</sup>Samples/mm<sup>2</sup>, or 1 × 10<sup>9</sup>Samples/mm<sup>2</sup>Alternatively, or in addition, the average density of analytes on the array can be at most about 1×10<sup>9</sup>Samples/mm<sup>2</sup>、1×10<sup>8</sup>Samples/mm<sup>2</sup>、1×10<sup>7</sup>Samples/mm<sup>2</sup>、1×10<sup>6</sup>Samples/mm<sup>2</sup>、1×10<sup>5</sup>Samples/mm<sup>2</sup>、1×10<sup>4</sup>Samples/mm<sup>2</sup>, or 1 × 10<sup>3</sup>Samples/mm<sup>2</sup>It can be the following:
The above ranges can apply, for example, to all or a portion of a regular pattern that includes all or a portion of an array of analytes.
The analytes in the pattern can have any of a variety of shapes. For example, when viewed in a two-dimensional plane, such as on the surface of an array, the analytes may appear rounded, circular, elliptical, rectangular, square, symmetrical, asymmetrical, triangular, polygonal, etc. The analytes can be arranged in a regular repeating pattern, including, for example, a hexagonal or rectilinear pattern. The pattern can be selected to achieve a desired level of packing. For example, circular analytes are optimally packed in a hexagonal arrangement. Of course, other packaging configurations can also be used for circular analytes, and vice versa.
A pattern can be characterized in terms of the number of analytes present in a subset that forms the smallest geometric unit of the pattern. A subset can include, for example, at least about 2, 3, 4, 5, 6, 10, or more analytes. Depending on the size and density of the analytes, a geometric unit can be as small as 1 mm<sup>2</sup>, 500μm<sup>2</sup>, 100μm<sup>2</sup>, 50μm<sup>2</sup>, 10μm<sup>2</sup>, 1 μm<sup>2</sup>, 500 nm<sup>2</sup>, 100 nm<sup>2</sup>, 50 nm<sup>2</sup>, 10 nm<sup>2</sup>Alternatively or additionally, the geometric unit may be 10 nm<sup>2</sup>, 50 nm<sup>2</sup>, 100 nm<sup>2</sup>, 500 nm<sup>2</sup>, 1 μm<sup>2</sup>, 10μm<sup>2</sup>, 50μm<sup>2</sup>, 100μm<sup>2</sup>, 500μm<sup>2</sup>, 1mm<sup>2</sup>The properties of the analytes in the geometric units, such as shape, size, pitch, etc., can be selected from those described herein more generally with respect to the analytes in the array or pattern.
An array with a regular pattern of analytes may be ordered with respect to the relative location of the analytes, but random with respect to one or more other characteristics of each analyte. For example, in the case of nucleic acid sequences, the nucleic acid analytes may be regular with respect to their relative location, but random with respect to the knowledge of the sequence for the nucleic acid species present in any particular analyte. As a more specific example, a nucleic acid sequence formed by seeding a repeating pattern of analytes with a template nucleic acid and amplifying the template in each analyte to form copies of the template in the analyte (e.g., via cluster amplification or bridge amplification) will have a regular pattern of nucleic acid analytes, but will be random with respect to the distribution of the sequence of the nucleic acid across the array. Thus, detection of the presence of nucleic acid material on an array can result in a repeating pattern of analytes, whereas sequence-specific detection can result in a non-repeating distribution of signal across the array.
It will be understood that the descriptions of patterns, orders, randomness, etc. herein relate not only to analytes on an object, such as analytes on an array, but also to analytes in an image. Thus, the patterns, orders, randomness, etc. can exist in any of a variety of formats used to store, manipulate, or communicate image data, including, but not limited to, computer readable media or computer components such as graphical user interfaces or other output devices.
As used herein, the term "image" is intended to mean a representation of all or a part of an object. The representation may be an optically detected reproduction. For example, an image may be obtained from fluorescence, luminescence, scattering, or absorption signals. The part of the object present in the image may be the surface or other xy plane of the object. Typically, an image is a two-dimensional representation, but in some cases, the information in the image may be derived from three or more dimensions. An image need not include optically detected signals. Non-optical signals may be present instead. An image may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.
As used herein, "image" refers to a reproduction or representation of at least a portion of a sample or other object. In some embodiments, the reproduction is an optical reproduction, for example, produced by a camera or other optical detector. The reproduction can be a non-optical reproduction, for example, a representation of an electrical signal obtained from an array of nanopore analytes, or a representation of an electrical signal obtained from an ion-sensitive CMOS detector. In certain embodiments, non-optical reproductions can be excluded from the methods or devices described herein. The image can have a resolution that can distinguish analytes present at any of a variety of intervals, including, for example, those that are less than 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm apart.
As used herein, "acquisition," "capture," and like terms refer to any part of the process of acquiring an image file. In some implementations, data acquisition can include generating an image of the specimen, looking for a signal in the specimen, directing a detection device to look for or generate an image of the signal, providing instructions for further analysis or transformation of the image file, and instructions for any number of transformations or manipulations of the image file.
As used herein, the term "template" refers to a representation of the location or relationship between signals or analytes. Thus, in some embodiments, the template is a physical grid with a representation of signals corresponding to analytes in a sample. In some embodiments, the template can be a chart, table, text file, or other computer file that indicates locations corresponding to analytes. In the embodiments presented herein, the template is generated to track the location of analytes across a set of images of the sample captured at different reference points. For example, the template can be x, y coordinates or a series of values that describe the direction and/or distance of one analyte relative to another analyte.
As used herein, the term "specimen" can refer to an object or area of an object from which an image is captured. For example, in an embodiment where an image is taken from the surface of soil, a parcel of land can be a specimen. In other embodiments where biomolecule analysis is performed within a flow cell, the flow cell can be divided into any number of subdivisions, each of which can be a specimen. For example, the flow cell can be divided into various channels or lanes, and each lane can be further divided into 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 140, 160, 180, 200, 400, 600, 800, 1000 or more separate areas to be imaged. One example of a flow cell has eight lanes, with each lane divided into 120 specimens or tiles. In another embodiment, a sample can be made in multiple tiles, or even the entire flow cell. Thus, the image of each specimen can represent a larger area of the surface being imaged.
It will be understood that references to ranges and sequential lists of numbers herein include not only the numbers recited but all real numbers between the recited numbers.
As used herein, a "reference point" refers to any temporal or physical distinction between images. In another preferred embodiment, the reference point is a time point. In a more preferred embodiment, the reference point is a time point or cycle during the sequencing reaction. However, the term "reference point" can include other aspects that distinguish or separate images, such as angle, rotation, time, or other aspects that can distinguish or separate the images.
As used herein, "subset of images" refers to a group of images in a set. For example, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or any number of images selected from a set of images. In certain other embodiments, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60 or less or any number of images selected from a set of images. In another preferred embodiment, the images are obtained from one or more sequencing cycles, with four images correlating to each cycle. Thus, for example, a subset may be a group of 16 images acquired over four cycles.
Base refers to a nucleotide base or nucleotide, (adenine), C (cytosine), T (thymine), or G (guanine). This application uses "base(s)" and "nucleotide(s)" interchangeably.
The term "chromosome" refers to a genetic carrier of the present invention of a living cell, derived from a chromatin strand containing DNA and protein components (especially histones). In this specification, the conventional internationally recognized individual human genome chromosome numbering system is used herein.
The term "site" refers to a unique location (e.g., chromosome ID, chromosomal location and orientation) on a reference genome. In some embodiments, a site may be a residue, a sequence tag, or the location of a segment on a sequence. The term "locus" may be used to refer to a specific location of a nucleic acid sequence or polymorphism on a reference chromosome.
The term "sample" as used herein typically refers to a sample derived from a biological fluid, cell, tissue, organ, or organism containing the nucleic acid to be sequenced and/or phased, or a sample derived from a mixture of nucleic acids containing at least one nucleic acid sequence to be sequenced and/or phased. Such samples include, but are not limited to, sputum/oral fluid, amniotic fluid, blood, blood fractions, fine needle biopsy samples (e.g., surgical biopsy, needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, tissue explants, organ cultures, and any other tissue or cell preparations thereof, or fractions or derivatives thereof. Samples are often taken from human subjects (e.g., patients), but samples can be taken from any organism that has chromosomes, including, but not limited to, dogs, cats, horses, goats, sheep, cows, pigs, etc. Samples can be used directly as obtained from a biological source, or after pretreatment to modify the characteristics of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, and the like.
The term "sequence" includes or refers to a chain of nucleotides linked together. The nucleotides can be based on DNA or RNA. It is understood that one sequence may include multiple subsequences. For example, a single sequence (e.g., a PCR amplicon) may have 350 nucleotides. A sample read may include multiple subsequences within these 350 nucleotides. For example, a sample read may include a first and a second flanking subsequence, e.g., having 20-50 nucleotides. The first and second flanking subsequences may be located on either side of a repeat segment with a corresponding subsequence (e.g., 40-100 nucleotides). Each of the flanking subsequences may include (or may include a portion of) a primer subsequence (e.g., 10-30 nucleotides). For ease of reading, the term "subsequence" is referred to as "sequence", but it is understood that the two sequences need not be separate from each other on a common strand. To distinguish the various sequences described herein, the sequences may be given different labels (e.g., target sequence, primer sequence, flanking sequence, reference sequence, etc.). Other terms, such as "allele," may be given different labels to distinguish similar entities. Applications use "read(s)" and "sequence read(s)" interchangeably.
The term "paired end sequencing" refers to a sequencing method that sequences both ends of a target fragment. Paired end sequencing can facilitate the detection of genomic rearrangements and repeated segments, as well as the detection of gene fusions and novel transcripts. Methods of paired end sequencing are described in International Publication No. WO 07010252, International Application No. PCTGB2007/003798, and US Patent Publication No. 2009/0088327, each of which is incorporated herein by reference. In one embodiment, a series of operations may be performed as follows: (a) generating clusters of nucleic acids; (b) linearizing the nucleic acids; (c) hybridizing a first sequencing primer and performing repeated cycles of extension, scanning, and deblocking; (d) "flipping" the target nucleic acid on the flow cell surface by synthesizing a complementary copy; (e) linearizing the resynthesized strand; and (f) hybridizing a second sequencing primer and performing repeated cycles of extension, scanning, and deblocking. The inversion operation can deliver the reagents described above for a single cycle of bridge amplification.
The term "reference genome" or "reference sequence" refers to a specific known genomic sequence, either partial or complete, of any organism that can be used to reference an identified sequence from a subject. For example, the reference genomes used for human subjects, as well as many other organisms, are available from the National Center for Biotechnology Information at Found at ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus, expressed in a nucleic acid sequence. A genome includes both genes and non-coding sequences of DNA. A reference sequence may be larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105 times larger, or at least about 106 times larger, or at least about 107 times larger. In one example, the reference genome sequence is of the full-length human genome. In another example, the reference genome sequence is limited to a particular human chromosome, such as chromosome 13. In some embodiments, the reference chromosome is a chromosome sequence from the human genome version hg19. Such sequences may be referred to as chromosomal reference sequences, although the term reference genome is intended to encompass such sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, sub-chromosomal regions (strands, etc.) of any species, and the like. In various embodiments, the reference genome is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a specific individual. In other embodiments, "genome" also covers so-called "graph genomes" that use a specific storage format and representation of genome sequences. In one embodiment, a graph genome stores data in a linear file. In another embodiment, a graph genome refers to a representation in which alternative sequences (e.g., different copies of a chromosome with small differences) are stored as different paths in a graph. Further information regarding the implementation of graph genomes can be found at https://www.biorxiv.org/content/biorxiv/early/2018/03/20/194530.full.pdf, the contents of which are incorporated herein by reference in their entirety.
The term "read" refers to a collection of sequence data describing a fragment of a nucleotide sample or reference. The term "read" can refer to a sample read and/or a reference read. Typically, but not necessarily, a read represents a short sequence of consecutive base pairs in a sample or reference. A read may be symbolically represented by the base pair sequence (ATCG) of the sample or reference fragment. A read may be stored in a memory device and appropriately processed to determine whether the read matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing instrument or indirectly from stored sequence information about the sample. In some cases, the DNA sequence is of sufficient length (e.g., at least about 25 bp) that it can be used to identify a larger sequence or region that can be aligned and specifically assigned to, for example, a chromosome or genomic region or gene.
Next generation sequencing methods include, for example, sequencing by synthesis technology (Illumina), pyrosequencing (454), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences), and sequencing by ligation (SOLiD sequencing). Depending on the sequencing method, the length of each read can vary from about 30 bp to over 10,000 bp. For example, DNA sequencing using a SOLiD sequencer generates nucleic acid reads of about 50 bp. In another example, Ion Torrent Sequencing generates nucleic acid reads of up to 400 bp, and 454 pyrosequencing generates nucleic acid reads of about 700 bp. In yet another example, single molecule real-time sequencing can generate reads of 10,000 bp to 15,000 bp. Thus, in certain embodiments, the nucleic acid sequence reads have a length of 30-100 bp, 50-200 bp, or 50-400 bp.
The term "sample read", "sample sequence" or "sample fragment" refers to sequence data relating to a genomic sequence of interest from a sample. For example, a sample read includes sequence data from a PCR amplicon having forward and reverse primer sequences. The sequence data can be obtained from any selected sequence procedure. A sample read can be, for example, a sequence-by-synthesis (SBS) reaction, a sequencing and ligation reaction, or any other suitable sequencing method in which it is desired to determine the length and/or identity of repetitive elements. A sample read can be a consensus (e.g., average or weighted) sequence derived from multiple sample reads. In certain embodiments, providing a reference sequence includes identifying a locus of interest based on primer sequences of a PCR amplicon.
The term "raw fragment" refers to sequence data of a portion of a genomic sequence of interest that at least partially overlaps a designated or secondary location of interest in a sample read or sample fragment. Non-limiting examples of product fragments include double stitched fragments, simple stitched fragments, and simple non-stitched fragments. The term "raw" is used to indicate that the raw fragment contains sequence data that has some relationship to the sequence data in the sample read, regardless of whether the raw fragment shows supporting variants that correspond to and authenticate or confirm potential variants in the sample read. The term "raw fragment" does not indicate that the fragment necessarily contains supporting variants that validate the variant call in the sample read. For example, when a sample read is determined by a variant calling application to exhibit a first variant, the variant calling application can determine that one or more raw fragments lack a corresponding type of "supporting" variant that would otherwise be expected to occur in light of the variant in the sample read.
The terms "mapped," "aligned," "aligning," or "aligning" refer to the process of comparing a read or tag to a reference sequence, thereby determining whether the reference sequence contains the read sequence. If the reference sequence is read, the read may be mapped to the reference sequence, or in certain alternative embodiments, may be mapped to a specific location within the reference sequence. In some cases, the alignment simply tells whether the read is a member of a particular reference sequence (i.e., whether the read is present or absent in the reference sequence). For example, the alignment of a read to a reference sequence for human chromosome 13 tells whether the read is present in the reference sequence for chromosome 13. Tools that provide this information are sometimes called set membership testers. In some cases, the alignment also indicates the location within the reference sequence where the read or tag map is. For example, if the reference sequence is the entire human genome sequence, the alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and/or site of chromosome 13.
The term "indel" refers to the insertion and/or deletion of bases in the DNA of an organism. Microindels refer to indels that result in a net change of 1-50 nucleotides. Unless the length of the indel is a multiple of three, a frameshift mutation occurs in coding regions of the genome. Indels can be contrasted with point mutations. An indel insertion deletes a nucleotide from the sequence, whereas a point mutation is a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels can also be contrasted with tandem base mutations (TBM), which can be defined as a substitution at adjacent nucleotides (mostly two adjacent nucleotides are replaced, although substitutions at three adjacent nucleotides have been observed).
The term "variant" refers to a nucleic acid sequence that differs from a nucleic acid reference. Exemplary nucleic acid sequence variants include, but are not limited to, single nucleotide polymorphisms (SNPs), short deletion and insertion polymorphisms (Indels), copy number variations (CNVs), microsatellite markers, or short tandem repeats and structural variants. Somatic variant calling is an effort to identify variants that are present at low frequency in a DNA sample. Somatic variant calling is of interest in the context of cancer treatment. Cancer is caused by the accumulation of mutations in DNA. DNA samples from tumors are generally heterogeneous, containing some normal cells, early stages of cancer progression (with fewer mutations), and some late-stage cells (with more mutations). Due to this heterogeneity, somatic mutations often appear at low frequency when sequencing tumors (e.g., from FFPE samples). For example, an SNV may be found in only 10% of the reads that cover a given base. A variant that is classified as somatic or germline by a variant classifier is also referred to herein as a "variant under test."
The term "noise" refers to erroneous variant calls that arise from one or more errors in the sequencing process and/or variant calling application.
The term "variant frequency" refers to the relative frequency of an allele (variant of a gene) at a particular locus in a population, expressed as a fraction or percentage. For example, the fraction or percentage may be the percentage of all chromosomes in a population that carry that allele. As an example, a sample variant frequency refers to the relative frequency of an allele/variant at a particular locus/position along a genomic sequence of interest across a "population" corresponding to the number of reads and/or samples obtained for the genomic sequence of interest from an individual. As another example, a baseline variant frequency refers to the relative frequency of an allele/variant at a particular locus/position along one or more baseline genomic sequences, where the relative frequency of an allele/variant at a particular locus/position along one or more baseline genomic sequences is obtained for one or more baseline genomic sequences.
The term "variant allele frequency (VAF)" refers to the proportion of sequenced reads that carry a variant divided by the overall coverage at the target position. VAF is a measure of the proportion of sequenced reads that carry a variant.
The terms "position," "designated position," and "locus" refer to the position or coordinates of one or more nucleotides within a nucleotide sequence. The terms "position," "designated position," and "locus" also refer to the position or coordinates of one or more base pairs in a sequence of nucleotides.
The term "haplotype" refers to a combination of alleles at adjacent sites on a chromosome that are inherited together. A haplotype, if present, may be one locus, several loci, or an entire chromosome, depending on the number of recombination events that have occurred between a given pair of loci.
The term "threshold" herein refers to a number or values used as a cutoff to characterize a sample, a nucleic acid, or a portion thereof (e.g., a readout). The threshold may vary based on empirical analysis. The threshold may be compared to a measured or calculated value to determine whether the source giving rise to such a value should be classified in a particular way. The threshold may be identified empirically or analytically. The choice of threshold depends on the confidence with which the user desires the classification to be made. The threshold may be selected for a particular purpose (e.g., for a balance of sensitivity and selectivity). As used herein, the term "threshold" refers to a point at which the course of an analysis may change and/or an action may be triggered. A threshold does not have to be a predetermined number. Instead, the threshold may be, for example, a function based on multiple factors. The threshold may be adaptive to the situation. Furthermore, the threshold may indicate an upper limit, a lower limit, or a range between limits.
In some embodiments, an index or score based on the sequencing data may be compared to a threshold. As used herein, the term "metric" or "score" may include a value or result determined from the sequencing data, or may include a function based on a value or result determined from the sequencing data. As with a threshold, an index or score may be adaptive to the context. For example, an index or score may be a normalized value. As an example of a score or metric, one or more embodiments may use a count score in analyzing the data. The count score may be based on the number of sample reads. The sample reads may have been through one or more filtering stages such that the sample reads have at least one common characteristic or quality. For example, each of the sample reads used to determine the count score may be aligned with a reference sequence or assigned as a potential allele. The number of sample reads that have a common characteristic may be counted to determine the read count. The count score may be based on the read count. In some embodiments, the count score may be a value equal to the read count. In other examples, the count score may be based on the read count and other information. For example, the counting score may be based on the read counts of a particular allele of a locus and the total number of reads of the locus. In some embodiments, the counting score may be based on the read counts of a locus and previously obtained data. In some embodiments, the counting score may be a normalized score between predetermined values. The counting score may also be a function of read counts from other loci of the sample, or read counts from other samples run simultaneously with the sample of interest. For example, the counting score may be a function of the read counts of a particular allele and read counts of other loci in the sample, and/or read counts from other samples. As an example, the read counts from other loci and/or read counts from other samples may be used to normalize the counting score for a particular allele.
The term "coverage" or "fragment coverage" refers to a count or other measure of multiple sample reads to the same fragment of a sequence. The read count may represent a count of the number of reads that cover the corresponding fragment. Alternatively, coverage may be determined by multiplying the read count by a specified factor based on historical knowledge, knowledge of the sample, knowledge of the locus, etc.
The term "read depth" (conventionally a number followed by "x") refers to the number of sequenced reads with overlapping alignment at the target position. It is often expressed as an average or percentage above a cutoff for a set of intervals (such as exons, genes, or panels). For example, a clinical report may state that the panel average coverage is 1,105x with 98% of the targeted bases covered >100x.
The term "base call quality score" or "Q score" refers to a PHRED-scaled probability ranging from 0-50 that is inversely proportional to the probability that a single sequenced base is correct. For example, a T base call with a Q of 20 is considered 99.99% likely to be correct. Any base call with a Q<20 should be considered low quality, and any variant identified where a significant proportion of sequenced reads supporting the variant is low should be considered a potential false positive.
The term "variant read" or "variant read number" refers to the number of sequenced reads that support the presence of a variant.
In terms of "stringency" (or DNA strands), the genetic message in DNA can be represented as the letters A, G, C, and T, e.g., 5'-AGGACA-3'. Often, sequences are written in the orientation shown here, i.e., 5' end to the left and 3' end to the right. Although DNA can occur as a single-stranded molecule (like certain viruses), we usually find DNA as a double-stranded unit. It has a double helix structure with two antiparallel strands. In this case, the word "antiparallel" means that the two strands run parallel, but of opposite polarity. Double-stranded DNA is held together by base pairing, and pairing is always maintained such that adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). This pairing is called complementarity, and one DNA strand is said to be the complement of the other. Thus, double-stranded DNA can be represented as two strings in this way: 5'-AGGACA-3' and 3'-TCCTGT-5'. Note that the two strands have opposite polarity. Thus, the strandedness of the two DNA strands can be referred to as the reference strand and its complement, the forward and reverse strands, the top and bottom strands, the sense and antisense strands, or the Watson and Crick strands.
Read alignment (also called read mapping) is the process by which sequences in a genome are referred to as originating. Once aligned, the "mapping quality" or "mapping quality score (MAPQ)" of a given read quantifies the probability that its location on the genome is correct. Mapping quality is encoded on a phase scale, where P is the probability that the alignment is incorrect. The probability is calculated as follows: P=10<sup>(-MAQ/10)</sup>Where, MAPQ is the mapping quality. For example, a mapping quality of 40=10 for a power of -4 means that there is a 0.01% chance that the reads are inaccurately aligned. Thus, mapping quality is associated with several alignment factors such as the base quality of the reads, the complexity of the reference genome, and the pallid-end information. First, if the base quality of the reads is low, it means that the observed sequence is likely to be erroneous and therefore its alignment is incorrect. Second, mapping ability refers to the complexity of the genome. Repetitive regions are more difficult to map and the reads contained in these regions usually have a lower mapping quality. In this context, MAPQ reflects the fact that the reads are not uniquely aligned and their actual origin cannot be determined. Third, for pallid-end sequencing data, concordant pairs are more likely to be aligned. The higher the mapping quality, the better the alignment. Reads aligned with good mapping quality usually mean that the reads sequence well and aligned with only minor mismatches within the high mappable regions. The MAPQ value can be used as a quality control for the alignment results: the percentage of aligned reads with a MAPQ higher than 20 is usually for downstream analysis.
As used herein, "signal" refers to a detectable event, such as, for example, a light emission, preferably a light emission, in an image. Thus, in another preferred embodiment, the signal can represent any detectable light emission (i.e., a "spot") captured in an image. Thus, as used herein, "signal" can refer to both the actual emission from the analyte of the specimen, and to a spurious emission that does not correlate with the actual analyte. Thus, the signal may result from noise and can be subsequently discarded as not representative of the actual analyte of the test strip.
As used herein, the term "cluster" refers to a group of signals. In certain embodiments, the signals are from different specimens. In another preferred embodiment, the signal cluster is a group of signals that cluster together. In a more preferred embodiment, the signal cluster represents the physical area covered by one amplification oligonucleotide. Each signal cluster should ideally be observed as several signals (one per template cycle, possibly more due to crosstalk). Thus, overlapping signals are detected, where two (or more) signals are included in the template from the same signal cluster.
As used herein, terms such as "minimum," "maximum," "minimize," "maximize," and grammatical variations thereof may include values that are not absolute maximums or minimums. In some implementations, the values include near maximums and minimums. In other examples, the values may include local maximums and/or local minimums. In some implementations, the values may include only absolute maximums or minimums.
As used herein, "crosstalk" refers to the detection of a signal in one image that is also detected in a separate image. In another preferred embodiment, crosstalk may occur when an emitted signal is detected in two separate detection channels. For example, if an emitted signal occurs in one color, the emission spectrum of that signal may overlap with another emitted signal in another color. In a preferred embodiment, the fluorescent molecules used to indicate the presence of nucleotide bases A, C, G, and T are detected in separate channels. However, because the emission spectra of A and C overlap, a portion of the C color signal may be detected during detection using a color channel. Thus, crosstalk between the A and C signals allows a signal from one color image to appear in the other color image. In some embodiments, there is G and T crosstalk. In some embodiments, the amount of crosstalk between channels is asymmetric. It will be understood that the amount of crosstalk between channels can be controlled by, among other things, the selection of signal molecules with appropriate emission spectra, as well as the selection of the size and wavelength range of the detection channel.
As used herein, "register," "registration," "registration," and similar terms refer to any process for correlating signals in an image or dataset with signals in an image or dataset from another time or perspective. For example, registration can be used to align signals from a set of images to form a template. In another example, registration can be used to register signals from other images to a template. One signal may be directly or indirectly registered to another signal. For example, a signal from image "S" may be directly registered to image "G." As another example, a signal from image "N" may be directly registered to image "G," or a signal from image "N" may be registered to image "S" that was previously registered to image "G." Thus, the signal from image "N" is indirectly registered to image "G."
As used herein, the term "fiducial" is intended to mean a distinguishable reference point in or on an object. The fiducial point may be, for example, a mark, a second object, a shape, an edge, an area, an irregularity, a channel, a pit, a post, etc. The fiducial point may be present in an image of the object or in another data set derived from detecting the object. The fiducial point may be specified by an x and/or y coordinate in the plane of the object. Alternatively or additionally, the fiducial point may be specified by a z coordinate orthogonal to the xy plane, for example defined by the relative position of the object and the detector. One or more coordinates for the fiducial point may be specified with respect to one or more other analytes of the object, or an image or other data set derived from the object.
As used herein, the term "optical signal" is intended to include, for example, fluorescent, luminescent, scattering, or absorption signals. Optical signals may be detected in the ultraviolet (UV) range (about 200-390 nm), visible (VIS) range (about 391-770 nm), infrared (IR) range (about 0.771-25 micrometers), or other ranges of the electromagnetic spectrum. Optical signals may be detected in a manner that excludes all or part of one or more of these ranges.
As used herein, the term "signal level" is intended to mean an amount or quantity of detected energy or encoded information having a desired or predetermined characteristic. For example, optical signals can be quantified by one or more of intensity, wavelength, energy, frequency, power, brightness, etc. Other signals can be quantified according to characteristics such as voltage, current, electric field strength, magnetic field strength, frequency, power, temperature, etc. The absence of a signal is understood to be a signal level of zero, or a signal level that is not significantly distinguished from noise.
As used herein, the term "simulate" is intended to mean to create a representation or model of a physical or behavioral thing that predicts a physical or behavioral characteristic. The representation or model may often be distinguishable from the thing or behavior. For example, the representation or model may be distinguishable with respect to one or more characteristics, such as color, texture, size, or strength of signal detected from all or a portion of the shape. In certain implementations, the representation or model may be idealized, exaggerated, muted, or incomplete compared to something or behavior. Thus, in some implementations, the representation of the model may be, for example, a representation that represents with respect to at least one of the above characteristics. The representation or model may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.
As used herein, the term "specific signal" is intended to mean a detected energy or encoded information that is selectively observed over other energies or information, such as background energy or information. For example, a specific signal can be an optical signal detected at a particular intensity, wavelength, or color, an electrical signal detected at a particular frequency, power, or field strength, or other signal known in the art of spectroscopic and analytical detection.
As used herein, the term "swing" is intended to mean a rectangular portion of an object. A swing may be an elongated strip that is scanned by relative movement between the object and a detector in a direction parallel to the longest dimension of the strip. Generally, the width of the rectangular portion or strip is constant along its entire length. The multiple swages of the object may be parallel to one another. The multiple swages of the object may overlap one another, be adjacent to one another, or be separated from one another by interstitial regions.
As used herein, the term "variance" is intended to mean the expected difference and the observed difference, or the difference between two or more observations. For example, variance can be the discrepancy between an expected value and a measured value. Statistical functions such as standard deviation, squared standard deviation, coefficient of variation, etc. can be used to express variance.
As used herein, the term "xy coordinates" is intended to mean information that specifies a location, size, shape, and/or orientation within an xy plane. The information may be, for example, numerical coordinates in a Cartesian system. The coordinates may be provided relative to one or both of the x and y axes, or may be provided relative to another location within the xy plane. For example, the coordinates of an analyte in an object may specify the location of the analyte relative to a fiducial or the location of another analyte in the object.
As used herein, the term "xy-plane" is intended to mean a two-dimensional region defined by linear axes x and y. When used with reference to a detector and an object observed by the detector, it may be further specified as being orthogonal to the observation direction between the detector and the object being detected.
As used herein, the term "z-coordinate" is intended to mean information that specifies the location of a point, line, or region along an axis perpendicular to the xy plane. In certain other embodiments, the z-axis is perpendicular to the region of the object observed by the detector. For example, the direction of the focus of an optical system may be specified along the z-axis.
In some embodiments, the acquired signal data is transformed using an affine transformation. In some such embodiments, template generation utilizes the fact that affine transformations between color channels are consistent between runs. Because of this consistency, a set of default offsets can be used in determining the coordinates of analytes in a specimen. For example, a default offset file can include relative transformations (shifts, scales, skews) for different channels relative to one channel, such as the A channel. However, in other embodiments, offsets between color channels drift within and/or between runs, making offset-driven template generation difficult. In such examples, the methods and systems provided herein can utilize offset template generation, which is further described below.
In some aspects of the above embodiments, the system may include a flow cell. In some aspects, the flow cell includes lanes or other configurations of tiles, at least some of the tiles including one or more analyte groups. In some aspects, the analytes include a plurality of molecules, such as nucleic acids. In certain aspects, the flow cell is configured to extend a primer that hybridizes to a nucleic acid in the analyte to deliver a labeled nucleotide base to a sequence of the nucleic acid, thereby generating a signal corresponding to the analyte including the nucleic acid. In preferred embodiments, the nucleic acids in the analyte are identical or substantially identical to each other.
In some of the image analysis systems described herein, each image in the set of images includes a color signal, with different colors corresponding to different nucleotide bases. In some embodiments, each image in the set of images includes a signal having a single color selected from at least four different colors. In some embodiments, each image in the set of images includes a signal having a single color selected from four different colors. In some of the systems described herein, a nucleic acid can be sequenced by providing four different labeled nucleotide bases to an array of molecules to generate four different images, each image includes a signal having a single color, and the signal color is different for each of the four different images, thereby generating a cycle of four color images corresponding to the four possible nucleotides present at a particular position in the nucleic acid. In certain embodiments, the system includes a flow cell configured to deliver additional labeled nucleotide bases to the array of molecules, thereby generating a cycle of multiple color images.
In preferred embodiment forms, the methods provided herein may include determining whether a processor is actively acquiring data or whether the processor is in a low activity state. Acquiring and storing a large number of high quality images typically requires a large amount of storage capacity. Furthermore, once acquired and stored, analysis of image data may be resource intensive and may impede processing power for other functions, such as acquisition and storage of additional image data. Thus, as used herein, the term "low activity state" refers to the processing power of a processor at a given time. In some implementations, a low activity state occurs when a processor is not acquiring and/or storing data. In some implementations, a low activity state occurs when some data acquisition and/or storage is taking place, but additional processing power remains such that image analysis can occur simultaneously without interfering with other functions.
As used herein, "identifying a conflict" refers to identifying a situation in which multiple processes are competing for a resource. In some such implementations, one process is given priority over another process. In some implementations, the conflict may relate to the need to give priority to the allocation of time, processing power, storage power, or any other resource that is given priority. Thus, in some implementations, when processing time or capacity is distributed between two processes, such as either analyzing a data set and acquiring and/or storing a data set, a discrepancy between the two processes exists that can be resolved by giving priority to one of the processes.
Also provided herein is a system for performing image analysis. The system may include a processor, a storage capacity, and a program for image analysis, the program including instructions for processing a first data set for storage and a second data set for analysis, the processing including acquiring and/or storing the first data set on the storage device, and analyzing the second data set when the processor is not acquiring the first data set. In certain aspects, the program includes instructions for identifying at least one instance of a conflict between acquiring and/or storing the first data set and analyzing the second data set, and priority is given to acquiring and/or storing the image data such that acquiring and/or storing the first data set is given priority. In certain aspects, the first data set includes an image file acquired from an optical imaging device. In certain aspects, the system further includes an optical imaging device. In some aspects, the optical imaging device includes a light source and a detection device.
As used herein, the term "program" refers to instructions or commands for performing a task or process. The term "program" may be used interchangeably with the term "module." In certain implementations, a program may be a compilation of various instructions executed under the same command set. In other implementations, a program may refer to a separate batch or file.
The following are some of the surprising effects of utilizing the methods and systems for performing image analysis described herein. In some sequencing implementations, an important measure of the usefulness of a sequencing system is its overall efficiency. For example, the amount of mappable data generated per day, as well as the total cost of installing and running the instrument, are important aspects of an economical sequencing solution. To reduce the time to generate mappable data and increase the efficiency of the system, real-time base calling can be enabled on the instrument computer and can be performed in parallel with sequencing chemistry and imaging. This allows data processing and analysis to be completed before sequencing chemistry finishing. In addition, it can reduce the storage required for intermediate data and limit the amount of data that needs to be moved across the network.
While the sequencing output is increasing, the data per run transferred from the system provided herein to the network and secondary analysis processing hardware is substantially reduced. By converting the data on the instrument computer (acquisition computer), the network load is dramatically reduced. Without these on-instrument, off-network data reduction techniques, the image output of the DNA sequencing instrument FRET would cripple most networks.
The widespread adoption of high-throughput DNA sequencing instruments has been driven in part by their ease of use, their support for a range of applications, and their suitability for virtually any lab environment. The highly efficient algorithms presented herein allow for the addition of significant analytical capabilities to simple workstations capable of controlling sequencing instruments. This reduction in computational hardware requirements has several practical advantages that will become even more important as sequencing output levels continue to increase. For example, by performing image analysis and base calling with simple towers, heat generation, laboratory footprint, and power consumption are kept to a minimum. In contrast, other commercial sequencing technologies have recently ramped up their computing infrastructure, with up to five times the processing power for primary analysis, before beginning to increase heat output and power consumption. Thus, in some embodiments, the computational efficiency of the methods and systems provided herein allows for increased sequencing throughput while minimizing server hardware.
Thus, in some embodiments, the methods and/or systems presented herein function as a state machine, keeping track of each sample's individual state, and upon detecting that a sample is ready to progress to the next state, taking appropriate action to advance the sample to that state. A more detailed example of how a state machine monitors the file system to determine when a sample is ready to progress to the next state in accordance with a preferred embodiment is provided in Example 1 below.
In preferred embodiments, the methods and systems provided herein are multi-threaded and can work with a configurable number of threads. Thus, for example, in the context of nucleic acid sequencing, the methods and systems provided herein can work in the background during live sequencing runs for real-time analysis, or can be run using existing image datasets for offline analysis. In certain preferred embodiments, the methods and systems handle multi-threading by giving each thread its own subset of the analytes it is involved in. This minimizes the possibility of thread retention.
The disclosed method may include using a detection device to acquire a target image of an object, the image including a repeating pattern of analytes on the object. Detection devices capable of high resolution imaging of a surface are particularly useful. In certain embodiments, the detection device will have sufficient resolution to distinguish analytes at the densities, pitches, and/or analyte sizes described herein. Detection devices capable of acquiring images or image data from a surface are particularly useful. Exemplary detectors are those configured to acquire area images while maintaining a static relationship between the object and the detector. Scanning devices may also be used. For example, devices that acquire continuous area images (e.g., referred to as "step and shot" detectors) may be used. Also useful are devices that continuously scan points or lines on the surface of an object to accumulate data to build an image of the surface. Point scanning detectors may be configured to scan points (i.e., small detection areas) on the surface of an object via a raster motion in the xy plane of the surface. Line scanning detectors may be configured to scan a line along the y dimension of the surface of the object, with the longest dimension of the line occurring along the x dimension. It will be appreciated that scanning detection may be accomplished by moving the detection device, the object, or both. For example, detection devices that are particularly useful in nucleic acid sequencing applications are described in U.S. Patent Application Publication No. 2012/0270305(A1), U.S. Patent Application Publication No. 2013/0023422(A1), and U.S. Patent Application Publication No. 2013/0260372(A1), and U.S. Patent Nos. 5,528,050, 5,719,391, 8,158,926, and 8,241,573, each of which is incorporated herein by reference.
The embodiments disclosed herein may be implemented as a method, apparatus, system, or article of manufacture using programming or engineering techniques to generate software, firmware, hardware, or any combination thereof. As used herein, the term "article of manufacture" refers to code or logic embodied in hardware or computer readable media, such as optical storage devices, as well as volatile or non-volatile memory devices. Such hardware may include, but is not limited to, field programmable gate arrays (FPGAs), coarse grain reconfigurable structures (CGRAs), application specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microprocessors, or other similar processing devices. In certain embodiments, the information or algorithms described herein reside in a non-transitory storage medium.
In certain embodiments, the computer-implemented methods described herein can occur in real-time while multiple images of an object are being acquired. Such real-time analysis is particularly useful for nucleic acid sequencing applications where nucleic acid sequences are subjected to repeated cycles of fluidic and detection steps. While analysis of sequencing data can often be beneficial to perform the methods described herein in real-time or in the background, it can be beneficial to perform the methods described herein while other data acquisition or analysis algorithms are in process. Examples of real-time analysis methods that can be used in the present methods are those commercially available from Illumina, Inc. (San Diego, Calif.) and/or used in the MiSeq and HiSeq sequencing instruments described in U.S. Patent Application Publication No. 2012/0020537(A1), which is incorporated herein by reference.
An exemplary data analysis system, formed by one or more programmed computers, having programming stored on one or more machine-readable media with code executed to perform one or more steps of the methods described herein. In one embodiment, for example, the system includes an interface designed to enable networking of the system to one or more detection systems (e.g., optical imaging systems) configured to acquire data from a target object. The interface can receive and condition the data, if appropriate. In certain embodiments, the detection system outputs image data representing, for example, individual image elements or pixels that together form an image of an array or other object. The processor processes the received detection data according to one or more routines defined by the processing code. The processing code may be stored in various types of memory circuits.
According to currently contemplated embodiments, the processing code executed on the detection data includes data analysis routines designed to analyze the detection data to determine the location of individual analytes visible or encoded in the data, as well as locations where no analyte is detected (i.e., where no analyte is present or no significant signal is detected from an existing analyte) and metadata. In certain embodiments, analyte locations in the array typically appear brighter than non-analyte locations due to the presence of a fluorescent dye attached to the imaged analyte. It will be understood that analytes need not appear brighter than their surrounding areas if, for example, no target of a probe in the analyte is present in the array being detected. The color in which the individual analytes appear can be a function of the dye used, as well as the wavelength of light used by the imaging system for imaging purposes. Analytes without bound targets or without a specific label can be identified according to other characteristics, such as their expected location in the microarray.
Once the data analysis routine has located the individual analytes in the data, a value assignment may be performed. Generally, the value assignment assigns a digital value to each analyte based on the characteristics of the data represented by the detector element (e.g., pixel) at the corresponding location. That is, for example, when the imaging data is processed, the value assignment routine may be designed to recognize that a particular color or wavelength of light has been detected at a particular location. In a typical DNA imaging application, for example, the four common nucleotides are represented by four separate and distinguishable colors. Each color may then be assigned a value corresponding to that nucleotide.
As used herein, the terms "module," "system," or "system controller" may include hardware and/or software systems and circuits that operate to perform one or more functions. For example, a module, system, or system controller may include a computer processor, controller, or other log-based device that performs operations based on instructions stored on a tangible and non-transitory computer-readable storage medium, such as a computer memory. Alternatively, a module, system, or system controller may include a hardwired device that performs operations based on hardwired logic and circuitry. The modules, systems, or system controllers shown in the accompanying drawings may represent hardware and circuits that operate based on software or hardwired instructions, software that instructs hardware to operate, or a combination thereof. A module, system, or system controller may include or represent hardware circuits or circuits that include and/or are connected with one or more processors, such as one or more computer microprocessors.
As used herein, the terms "software" and "firmware" are used interchangeably and include any computer program stored in memory that is executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely examples and are not intended to be limiting of the types of memory that can be used to store computer programs.
In the field of molecular biology, one of the processes for nucleic acid sequencing in use is sequence synthesis. This technology can be applied to highly parallel sequencing projects. For example, by using an automated platform, it is possible to perform millions of sequencing reactions simultaneously. Therefore, one embodiment of the present invention relates to an apparatus and method for acquiring, storing, and analyzing image data generated during nucleic acid sequencing.
The enormous gain in the amount of data that can be acquired and stored makes streamlined image analysis methods even more beneficial. For example, the image analysis methods described herein allow both designers and end users to make efficient use of existing computer hardware. Thus, methods and systems are presented herein that reduce the computational complexity of processing data in terms of rapidly increasing data output. For example, in the field of DNA sequencing, yields have expanded 15-fold in recent times and can reach hundreds of gigases in a single run of a DNA sequencing device. Large-scale genome-scale experiments are out of reach for most researchers if the computational infrastructure requirements increase proportionately. Thus, the generation of more raw sequence data increases the need for secondary analysis and data storage, making optimization of data transport and storage highly beneficial. Some embodiments of the methods and systems presented herein can reduce the time, hardware, networking, and laboratory infrastructure requirements required to generate usable sequence data.
The present disclosure describes various methods and systems for performing the methods. Some examples of the methods are described as a series of steps. However, it should be understood that the implementations are not limited to the specific steps and/or order of steps described herein. Steps may be omitted, steps may be modified, and/or other steps may be added. Furthermore, the steps described herein may be combined, steps may be performed simultaneously, steps may be divided into multiple sub-steps, steps may be performed in a different order, or steps (or a series of steps) may be re-performed iteratively. In addition, although different methods are described herein, it should be understood that in other implementations, different methods (or steps of different methods) may be combined.
In some embodiments, a processing unit, processor, module, or computing system that is "configured" to perform a task or operation may be understood to be specifically structured to perform the task or operation (e.g., having one or more programs or instructions that are adapted or intended to perform a task or operation, and/or having an arrangement of processing circuitry that is adapted or intended to perform a task or operation). For clarity and avoidance of doubt, a general purpose computer (which may be "configured to perform a task or operation" when suitably programmed) is not configured as being "configured" to perform a task or operation unless specifically programmed or structurally modified to perform the task or operation).
Moreover, the operations of the methods described herein may be sufficiently complex such that the operations cannot be performed by the average person or by a person skilled in the art in a commercially reasonable period of time. For example, the methods may rely on relatively complex calculations such that such a person cannot complete the methods in a commercially reasonable period of time.
Throughout this application, various publications, patents, or patent applications are referenced. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art to which this invention pertains.
The term "comprising," as used herein, is intended to be open ended, encompassing not only the recited elements, but any additional elements as well.
As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set. Exceptions may occur where express disclosure or context clearly dictates otherwise.
Although the invention has been described with reference to the above embodiments, it will be understood that various modifications can be made without departing from the invention.
The modules of the present application may be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some may be implemented on different processors or computers, or may be spread among many different processors or computers. In addition, it will be understood that some of the modules may be operated in parallel or in a different order than that shown in the figures, without affecting the functionality achieved. Also, as used herein, the term "module" may include "sub-modules," which may be considered herein to constitute a module. The blocks of the figures designated as modules may also be considered as flow chart steps in a method.
As used herein, "identification" of an item of information does not necessarily require a direct specification of that item of information. Information may be "identified" within a field by simply referencing the actual information through one or more layers in one direction, or by identifying one or more items of different information that are sufficient to determine the actual item of information. Additionally, the term "designate" is used herein to mean the same thing as "identify."
As used herein, a given signal, event, or value is "dependent on a pre-decessor signal, event, or value of a pre-decessor signal, an event, or value that is affected by the given signal, event, or value. If an intervening processing element, step, or time period is present, then the given signal, event, or value may "exist" depending on a "pre-decessor signal, event, or value." If an intervening processing element or step combines two or more signals, events, or values, then the signal output of the processing element or step is considered to be dependent on "each of the signal, event, or value inputs." If a given signal, event, or value is the same as a pre-decessor signal, event, or value, then this is simply considered to mean that the given signal, event, or value is "dependent" or "depends" on or "depends on" a "pre-decessor signal, event, or value" or "depends on" a "base decessor signal, event, or value." The "responsiveness" of a given signal, event, or value to another signal, event, or value is defined similarly.
As used herein, "concurrently" or "parallel" does not require exact simultaneity. It is sufficient if the assessment of one of the individuals begins before the assessment of another of the individuals is completed.
(Computer Systems)
FIG. 82 is a computer system 8200 that may be used by sequencing system 800A to implement the techniques disclosed herein. Computer system 8200 includes at least one central processing unit (CPU) 8272 that communicates with a number of peripheral devices via a bus subsystem 8255. These peripheral devices may include, for example, a storage subsystem 8210 including memory devices and file storage subsystem 8236, user interface input devices 8238, user interface output devices 8276, and a network interface subsystem 8274. The input and output devices enable user interaction with computer system 8200. Network interface subsystem 8274 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
In one embodiment, the system controller 7806 is communicatively linked to a storage subsystem 8210 and a user interface input device 8238 .
The user interface input devices 8238 can include keyboards, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, scanners, touch screens integrated into displays, audio input devices such as voice recognition systems and microphones, and other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 8200.
The user interface output devices 8276 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual displays such as an audio output device. In general, use of the term "output device(s)" is intended to include all possible types of devices and methods for outputting information from the computer system 8200 to a user or to another machine or computer system.
The storage subsystem 8210 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the deep learning processor 8278.
The deep learning processor 8278 may be a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and/or coarse grained reconfigurable architectures (CGRAs). The deep learning processor 8278 may be hosted by a deep learning cloud platform such as Google Cloud Platform, Xilinx, and Cirrascale. Examples of deep learning processors 8278 include Google's Tensor Processing Unit (TPU), GX4 Rackmount Series, GX82 Rackmount Series, NVIDIA DGX-1, Microsoft's Stratix V FPGA, Graphcore's Intelligent Processor Unit (IPU), Snapdragon processors, NVIDIA's Volta, NVIDIA's Drive PX, NVIDIA's JETSON TX1/TX2 MODULE, Intel's Nirvana, Movidius VPU, Fujitsu DPI, Arm DynamicIQ, IBM TrueNorth, Lambda GPU servers were used along with Testa V100s and others.
The memory subsystem 8222 used in the storage subsystem 8210 may include a number of memories including a main random access memory (RAM) 8232 for storing instructions and data during program execution, and a read only memory (ROM) 8234 in which fixed instructions are stored. The file storage subsystem 8236 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, drives, optical drives, or removable media cartridges. Modules implementing the functionality of a particular embodiment may be stored by the file storage subsystem 8236 in the storage subsystem 8210 or in other machines accessible by the processor.
The bus subsystem 8255 provides a mechanism for allowing the various components and subsystems of the computer system 8200 to communicate with each other as intended. Although the bus subsystem 8255 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.
The computer system 8200 itself can be of a variety of types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the varying nature of computers and networks, the description of computer system 8200 shown in Figure 82 is intended only as a specific example for purposes of illustrating a preferred embodiment of the invention. Many other configurations of computer system 8200 can have more or fewer components than the computer system shown in Figure 82.
(Specific Improvements)
We have described various embodiments of neural network-based template generation and neural network-based base calling. One or more features of the embodiments can be combined with the base embodiments. Non-mutually exclusive embodiments are taught as combinable. One or more features of the embodiments can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of the enumeration of repeating these options should not be construed as limiting the combinations taught in the preceding sections. These descriptions are incorporated herein by reference in each of the following implementations.
(Subpixel base call)
We disclose a computer-implemented method for determining metadata for analytes on tiles of a flow cell. The method includes accessing a series of image sets generated during a sequencing run, each image set being set in a series generated during a respective sequencing cycle of the sequencing run, each image in the series having a plurality of sub-pixels. The method includes obtaining base calls from a base color that classify each of the sub-pixels as one of four bases (A, C, T, and G), thereby generating a base call sequence for each of the sub-pixels over a plurality of sequencing cycles of the sequencing operation. The method includes generating an analyte map that identifies the analytes as discontinuous regions of contiguous sub-pixels that share substantially matching base call sequences. The method includes determining a spatial distribution of the analytes, including their shapes and sizes, based on the discontinuous regions, and storing the analyte map in a memory for use as ground truth for training a classifier.
The methods described in this and other sections of the disclosed technology may include one or more of the following features and/or characteristics described in connection with the additional methods disclosed. For the sake of brevity, combinations of features disclosed in this application are not individually recited and are not repeated for each base set of features. The reader will understand how features specified in these embodiments can be readily combined with the base feature sets specified in other embodiments.
In one embodiment, the method includes identifying subpixels in the analyte map that do not belong to any of the non-joint regions. In one embodiment, the method includes classifying each of the subpixels as one of five bases (A, C, T, G, and N) from a base caller. In one embodiment, the analyte map identifies analyte boundaries between two consecutive subpixels where the base call sequences do not substantially match.
In one embodiment, the method includes identifying an origin subpixel in the preliminary center coordinates of the analyte determined by the base color, and a breadth-first search to substantially match the base call sequence by starting from the origin subpixel and continuing with consecutive non-origin subpixels. In one embodiment, the method includes determining analyte center coordinates on an analyte-by-analyte basis, calculating a center of mass of a discontinuous region of the analyte map as an average of the coordinates of each consecutive subpixel forming the discontinuous region, and storing the hyperlocation center coordinates of the analyte on the analyte on an analyte-by-analyte basis to use as ground truth for training the classifier.
In one embodiment, the method includes identifying centers of mass subpixels in discrete regions of an analyte map on an analyte by analyte basis, upsampling the analyte map using interpolation, and storing the upsampled analyte map in memory for use as ground truth for training a classifier, and assigning, for each analyte, a value to each contiguous subpixel in the discrete region based on an attenuation coefficient proportional to the distance of the adjacent subpixel from the center of mass subpixel in the discrete region to which it belongs in the analyte by analyte upsampled analyte map. In one embodiment, the value is an intensity value normalized between zero and one. In one embodiment, the method includes assigning the same predetermined value to all subpixels identified as background in the upsampled analyte map. In one embodiment, the predetermined value is a zero intensity value.
In one embodiment, the method includes generating an attenuation map from an upsampled analyte map representing contiguous subpixels in the separated regions and subpixels identified as background based on their assigned values, and storing the attenuation map in memory for use as ground truth for training the classifier. In one embodiment, each subpixel in the attenuation map has a normalized value between zero and one. In one embodiment, the method includes classifying contiguous subpixels in the upsampled analyte map on an analyte by analyte basis as analyte interior subpixels belonging to the same analyte and analyte center subpixels as analyte boundary subpixels and subpixels identified as background as background subpixels, storing the classification in memory for use as ground truth for training the classifier.
In one embodiment, the method includes storing analyte-by-analyte coordinates of analyte interior subpixels, analyte center subpixels, boundary subpixels, and background subpixels on an analyte-by-analyte basis, downscaling the coordinates by a factor used to upsample the analyte map, and storing the downscaled coordinates in a memory for use as ground truth for training a classifier.
In one embodiment, the method includes labeling analyte center subpixels in binary ground truth data generated from the upsampled analyte map using color coding to label analyte center subpixels belonging to an analyte center class, and storing the binary ground truth data in memory for use as ground truth for training the classifier. In one embodiment, the method includes labeling background subpixels belonging to a background class in ternary ground truth data generated from the upsampled analyte map using color coding to label background subpixels belonging to a background class, and storing the ternary ground truth data in memory for use as ground truth for training the classifier.
In one embodiment, the method includes generating an analyte map for a number of tiles of a flow cell, storing the analyte map in a memory, and determining a spatial distribution of analytes within the tiles based on the analyte map including their shapes and sizes, and classifying, on an analyte by analyte basis, analyte interior subpixels as belonging to the same analyte, analyte center subpixels, boundary subpixels, and background subpixels in the upsampled analyte map of the analytes, and storing the classification in memory for use as ground truth for training a classifier, storing coordinates of the analyte interior subpixels, analyte center subpixels, and boundary subpixels on an analyte by analyte basis, downscaling coordinates by a factor used to upsample the analyte map, and storing the downscaled coordinates in memory for use as ground truth for training the classifier.
In one embodiment, base call sequences substantially match when a predetermined portion of the base calls match position by position in the order. In one embodiment, the base color generates the base call sequence by interpolating the intensities of the subpixels, including at least one of nearest neighbor intensity extraction, Gaussian intensity extraction, intensity extraction based on average 2x2 subpixel area, intensity extraction based on brightest test of 2x2 subpixel area, intensity extraction based on average 3x3 subpixel area, bilinear intensity extraction, bicubic intensity extraction, and/or intensity extraction based on weighted area coverage. In one embodiment, the subpixels are identified to the base caller based on their integer or non-integer coordinates.
In one embodiment, the method includes requiring that at least some of the discrete regions have a predetermined minimum number of subpixels. In one embodiment, the flow cell has at least one patterned surface having an array of wells that occupy analytes. In such an embodiment, the method determines which of the wells are substantially occupied by at least one analyte with one of the wells being minimally occupied based on the determined shapes and sizes of the analytes, and which of the wells are co-occupied by multiple analytes.
In one embodiment, the flow cell has at least one non-patterned surface, and the analytes are non-uniformly scattered across the non-patterned surface. In one embodiment, the density of the analytes is about 100,000 analytes/mm<sup>2</sup>~Approx. 1,000,000 specimens/mm<sup>2</sup>In one embodiment, the density of the specimens is in the range of about 1,000,000 specimens/mm<sup>2</sup>~Approx. 10,000,000 specimens/mm<sup>2</sup>In one embodiment, the subpixel is a quarter subpixel. In another embodiment, the subpixel is a half subpixel. In one embodiment, the preliminary center coordinates of the analyte determined by the base caller are defined in a template image of the tile, and the pixel resolution, image coordinate system, and measurement scale of the image coordinate system are the same as the template image and the image. In one embodiment, each image set has four images. In another embodiment, each image set has two images. In yet another embodiment, each image set has one image. In one embodiment, the sequencing operation utilizes four channel chemistry. In another embodiment, the sequencing operation utilizes two channel chemistry. In yet another embodiment, the sequencing run utilizes one channel chemistry.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
We disclose a computer-implemented method for determining metadata for analytes on tiles of a flow cell. The method includes accessing a set of images of the tile captured during a sequencing run and preliminary center coordinates of the analytes determined by base colors. The method includes obtaining, for each image set, one of four base origin subpixels that include a base center coordinate, and one of four base origin subpixels that include a predetermined neighborhood of contiguous subpixels that are contiguous to each of the origin subpixels, thereby generating a base call sequence for each of the origin subpixels and each of the predetermined neighborhood of contiguous subpixels. The method includes generating an analyte map that is contiguous to at least a portion of a corresponding one of the origin subpixels and shares a substantially matching base call sequence of one of the four bases with at least a portion of the corresponding one of the origin subpixels. The method includes storing the analyte map in a memory and determining a shape and size of the analyte based on the discontinuous regions in the analyte map.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the predetermined neighborhood of contiguous subpixels is an m x n subpixel patch centered on a pixel that includes the origin subpixel, where the subpixel patch is 3 x 3 pixels. In one embodiment, the predetermined neighborhood of contiguous subpixels is a neighborhood of n connected subpixels centered on a pixel that includes the origin subpixel. In one embodiment, the method includes identifying subpixels in the analyte map that do not fall into any of the discontinuous regions as background.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Training data generation)
We disclose a computer-implemented method of generating training data for neural network-based template generation and base calling. The method includes accessing a number of images of a flow cell captured over multiple cycles of a sequencing run, the flow cell having a number of tiles, and in the number of images, each tile having a sequence of image sets generated over multiple cycles, each image in the sequence of image sets showing intensity radiation of analytes and their surrounding background on a particular one of the tiles at a particular cycle. The method includes constructing a training set having a number of training examples, each training example corresponding to a particular one of the tiles and including image data from at least some of the image sets in the sequence of image sets for the particular one of the tiles. The method includes generating at least one ground truth data representation for each of the training examples, the ground truth data representation identifying at least one of the spatial distribution of analytes and their surrounding background on the particular one of the tiles whose intensity radiation is depicted by the image data, and including at least one of the analyte shape, analyte size, and/or analyte boundary, and/or analyte center.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the image data includes images of each of at least a portion of the image sets in the sequence of the image sets of a particular one of the tiles, the images having a resolution of 1800x1800. In one embodiment, the image data includes at least one image patch from each of the images, the image patch covering a portion of the particular one of the tiles and having a resolution of 20x20. In one embodiment, the image data includes an upsampled representation of the image patch, the upsampled representation having a resolution of 80x80. In one embodiment, the ground truth data representation has an upsampled resolution of 80x80.
In one embodiment, the multiple training examples correspond to the same particular one of the tiles, and each includes as image data a different image patch from each image in the sequence of image sets of the same particular image set of the tiles, where at least a portion of the different image patches overlap one another. In one embodiment, the ground truth data representation identifies the analytes as discontinuous regions of adjacent sub-pixels, and the centers of the analytes as centers of mass sub-pixels within each one of the discontinuous regions, and the analytes as their surrounding background. In one embodiment, the ground truth data representation uses color coding to identify each sub-pixel as either analyte-centered or non-centered. In one embodiment, the ground truth data representation uses color coding to identify each sub-pixel as either analyte-inside, analyte-centered, or surrounding background.
In one embodiment, the method includes storing in memory training examples in a training set and associated ground truth data representations as training data for neural network based template generation and base calling. In one embodiment, the method includes generating training data for various flow cells, sequencing instruments, sequencing protocols, sequencing chemistries, sequencing reagents, and analyte densities.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Metadata and call generation based)
In one embodiment, the method includes accessing sequence images of the analyte generated by a sequencer, generating training data from the sequence determination images, and using the training data to train a neural network to generate metadata about the analyte. Each of the features described in the specific embodiment section for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the set of base features specified in other embodiments. Other embodiments of the method described in this section can include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-mentioned methods. Yet another embodiment of the method described in this section can include a system including a memory and one or more processors operable to execute instructions stored in the memory, and can perform any of the above-mentioned methods.
In one embodiment, the method includes accessing sequence images of the specimen generated by a sequencer, generating training data from the sequence images, and using the training data to train a neural network to call the specimen. Each of the features described in the specific embodiment section for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the set of base features specified in other embodiments. Other embodiments of the method described in this section can include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-mentioned methods. Yet another embodiment of the method described in this section can include a system including a memory and one or more processors operable to execute instructions stored in the memory, and can perform any of the above-mentioned methods.
(Regression model)
The inventors have disclosed a computer-implemented method for identifying analytes on a tile of a flow cell and associated analyte metadata. The method includes processing input image data from a sequence of image sets through a neural network to generate alternative representations of the input image data. Each image in the sequence of image sets covers a tile and shows the intensity emission of analytes on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. The method includes processing the alternative representations through an output layer and generating an output of the analytes whose intensity emission is represented by the input image data as discontinuous regions of adjacent sub-pixels, the centers of the analytes as central sub-pixels at the center of mass of corresponding ones of the discontinuous regions, and the centers of the analytes as background sub-pixels that do not belong to any of the discontinuous regions.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, adjacent sub-pixels in corresponding ones of the discontinuous regions have intensity values weighted according to the distance of the adjacent sub-pixels from a central sub-pixel in the discontinuous region to which they belong. In one embodiment, the central sub-pixel has the highest intensity value in the corresponding ones of the discontinuous regions. In one embodiment, all background sub-pixels have the same lowest intensity value in the output. In one embodiment, the output layer normalizes the intensity values between zero and one.
In one embodiment, the method includes applying a peak locator to the output to find peak intensities in the output, determining location coordinates of centers of analytes based on the peak intensities, downscaling the location coordinates by an upsampling factor used to create the input image data, and storing the downscaled location coordinates in memory for use on an analyte-calling basis. In one embodiment, the method includes classifying adjacent subpixels as analyte-internal subpixels belonging to the same analyte, and storing the analyte-by-analyte-based classifications of analyte-internal subpixels and the downscaled location coordinates for use on an analyte-calling basis. In one embodiment, the method includes determining, on an analyte-by-analyte basis, distances of corresponding analyte-internal subpixels of centers of analytes, and storing the analyte-by-analyte-based distances in memory for use on an analyte-calling basis.
In one embodiment, the method includes extracting intensities from analyte interior subpixels in corresponding ones of the discontinuous regions including using at least one of nearest neighbor intensity extraction, Gaussian intensity extraction, intensity extraction based on average 2x2 subpixel area, intensity extraction based on brightest test of 2x2 subpixel areas, intensity extraction based on average 3x3 subpixel area, bilinear intensity extraction, quadratic intensity extraction, and/or intensity extraction based on weighted area coverage, and/or intensity extraction.
In one embodiment, the method includes determining a spatial distribution of the analytes including at least one of analyte shape, analyte size, and/or analyte boundary based on the discrete regions, and storing associated analyte metadata in an analyte-by-analyte-based memory for use on a analyte-by-analyte basis.
In one embodiment, the input image data includes images in a sequence of images set, the images having a resolution of 3000x3000. In one embodiment, the input image data includes at least one image patch from each of the images in the sequence of images set, the image patch covering a portion of a tile and having a resolution of 20x20. In one embodiment, the input image data includes upsampled representations of image patches from each of the images in the sequence of images set, the upsampled representation having a resolution of 80x80. In one embodiment, the output has an upsampled resolution of 80x80.
In one embodiment, the neural network is a deep fully convolutional segmentation neural network having an encoder sub-network and a corresponding decoder network, where the encoder sub-network includes a hierarchy of encoders and the decoder sub-network includes a hierarchy of decoders that map the low resolution encoder feature maps to the full input resolution feature maps. In one embodiment, the density of the specimens is about 100,000 specimens/mm<sup>2</sup>~Approx. 1,000,000 specimens/mm<sup>2</sup>In another embodiment, the density of the specimens is in the range of about 1,000,000 specimens/mm<sup>2</sup>~Approx. 10,000,000 specimens/mm<sup>2</sup>The range is.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Training regression models)
The inventors have disclosed a computer-implemented method for training a neural network to identify analytes and associated analyte metadata. The method includes obtaining training data for training the neural network. The training data includes a plurality of training examples and corresponding ground truth data to be generated by the neural network by processing the training examples. Each training example includes image data from a sequence of image sets. Each image in the sequence of image sets covers a tile of a flow cell and shows intensity emission of analytes on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. Each ground truth data identifies an analyte represented by the image data of the corresponding training example as a discontinuous region of adjacent sub-pixels, a center of the analyte as a central sub-pixel at the center of mass of each one of the discontinuous regions, and their surrounding background. The method includes training a neural network to generate outputs of the training examples, including iteratively optimizing a loss function that minimizes an error between the output and ground truth data, and updating parameters of the neural network based on the error; training a neural network to generate outputs of the training examples, and updating parameters of the neural network based on the error.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the method includes storing updated parameters of the neural network in memory upon error convergence after the last iteration to apply to further neural network based template generation and base calling. In one embodiment, in the ground truth data, adjacent sub-pixels in corresponding ones of the discontinuous regions have intensity values weighted according to the distance of the adjacent sub-pixels from a central sub-pixel in the joint region to which the adjacent sub-pixels belong. In one embodiment, in the ground truth data, the central sub-pixel has the highest intensity value in each discontinuous region. In one embodiment, in the ground truth data, all background sub-pixels have the same lowest intensity value in the output. In one embodiment, in the ground truth data, the intensity values are normalized between zero and one.
In one embodiment, the loss function is a mean squared error, which is minimized on a sub-pixel basis between normalized intensity values of corresponding sub-pixels in the output and the ground truth and the ground truth. In one embodiment, the ground truth data identifies a spatial distribution of the analytes including at least one of analyte shape, analyte size, and/or analyte boundary as part of the associated analyte metadata. In one embodiment, the image data includes images in a sequence of image sets, the images having a resolution of 1800x1800. In one embodiment, the image data includes at least one image patch from each of the images in the sequence of image sets, the image patch covering a portion of a tile and having a resolution of 20x20. In one embodiment, the image data includes an upsampled representation of an image patch from each of the images in the sequence of image sets, the upsampled representation of the image patch having a resolution of 80x80.
In one embodiment, in the training data, the multiple training examples are each represented as a different image patch of image data from each image in the sequence of the image set of the same tile, and at least a portion of the different image patches overlap each other. In one embodiment, the ground truth data has an upsampled resolution of 80x80. In one embodiment, the training data includes training examples of multiple tiles of a flow cell. In one embodiment, the training data includes training examples of various flow cells, sequencing installations, sequencing protocols, sequencing chemistries, sequencing reagents, and analyte densities. In one embodiment, the neural network is a deep fully convolutional segmentation neural network having an encoder sub-network and a corresponding decoder network, the encoder sub-network including a hierarchy of encoders, and the decoder sub-network including a hierarchy of decoders that map a low resolution encoder feature map to a full input resolution feature map for sub-pixel-by-subpixel classification by a final classification layer.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Neural network based template generator)
We disclose a computer-implemented method for determining metadata about an analyte on a flow cell, the method including accessing image data depicting an intensity emission of the analyte, processing the image data through one or more layers of a neural network, generating an alternative representation of the image data, and processing the alternative representation through an output layer to generate an output that identifies at least one of a shape and a size of a center of the analyte and/or specimen.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the image data further indicates intensity radiation of the background surrounding the analytes. In such an embodiment, the method includes an output identifying a spatial distribution of the analytes on the flow cell, including surrounding background and boundaries between the analytes. In one embodiment, the method includes determining a center location coordinate of the analyte on the flow cell based on the output. In one embodiment, the neural network is a convolutional neural network. In one embodiment, the neural network is a recurrent neural network. In one embodiment, the neural network is a deep full convolutional segmentation neural network having an encoder sub-network and a corresponding decoder network, followed by an output layer, where the encoder sub-network includes a hierarchy of encoders and the decoder sub-network includes a hierarchy of decoders that map the low resolution encoder feature map to the full input resolution feature map.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Binary Classification Model)
The inventors have disclosed a computer-implemented method for identifying analytes on a tile of a flow cell and associated analyte metadata. The method includes processing input image data from a sequence of image sets through a neural network and generating an alternative representation of the image data. In one embodiment, each image in the sequence of image sets covers a tile and shows the intensity emission of the analytes on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. The method includes processing the alternative representations through a classification layer and generating an output that identifies a center of the analyte whose intensity emission is represented by the input image data. The output has a plurality of sub-pixels, each sub-pixel in the plurality of sub-pixels is classified as either an analyte center or a non-center.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the classification layer assigns each subpixel in the output a first likelihood score that is analyte-centered and a second likelihood score that is non-centered. In one embodiment, the first and second likelihood scores are determined based on a softmax function and are exponentially normalized between zero and one. In one embodiment, the first and second likelihood scores are determined based on a sigmoid function and are normalized between zero and one. In one embodiment, each subpixel in the output is classified as either analyte-centered or non-centered based on whether one of the first and second likelihood scores is higher than the other. In one embodiment, each subpixel in the output is classified as either analyte-centered or non-centered based on whether the first and second likelihood scores are above a predetermined threshold likelihood score. In one embodiment, the output identifies a center of mass of a corresponding one of the analytes. In one embodiment, all subpixels in the output that are classified as analyte-centered are assigned the same first predetermined value and all subpixels that are classified as non-centered are assigned the same second predetermined value. In one embodiment, the first and second predetermined values are intensity values.In one embodiment, the first and second predetermined values are continuous values.
In one embodiment, the method includes determining location coordinates of subpixels classified as analyte centers, downscaling the location coordinates by an upsampling factor used to prepare the input image data, and storing the downscaled location coordinates in memory for use on an analyte calling basis. In one embodiment, the input image data includes images in a sequence of an image set, the images having a resolution of 3000x3000. In one embodiment, the input image data includes at least one image patch from each of the images in the sequence of an image set, the image patch covering a portion of a tile and having a resolution of 20x20. In one embodiment, the input image data includes upsampled representations of image patches from each of the images in the sequence of an image set, the upsampled representation having a resolution of 80x80. In one embodiment, the output has an upsampled resolution of 80x80.
In one embodiment, the neural network is a deep fully convolutional segmentation neural network having an encoder sub-network and a corresponding decoder network followed by a classification layer, where the encoder sub-network includes a hierarchy of encoders and the decoder sub-network includes a hierarchy of decoders that map the low resolution encoder feature maps to full input resolution feature maps for sub-pixel classification by the classification layer. In one embodiment, the density of the specimens is about 100,000 specimens/mm<sup>2</sup>~Approx. 1,000,000 specimens/mm<sup>2</sup>In another embodiment, the density of the specimens is in the range of about 1,000,000 specimens/mm<sup>2</sup>~Approx. 10,000,000 specimens/mm<sup>2</sup>The range is.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Training a binary classification model)
The inventors have disclosed a computer-implemented method for training a neural network to identify analytes and associated analyte metadata. The method includes obtaining training data for training the neural network. The training data includes a plurality of training examples and corresponding ground truth data to be generated by the neural network by processing the training examples. Each training example includes image data from a sequence of image sets. Each image in the sequence of image sets covers a tile of a flow cell and shows intensity emission of analytes on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. Each ground truth data identifies a center of an analyte whose intensity emission is indicated by the image data of the corresponding training example. The ground truth data has a plurality of sub-pixels, and each sub-pixel in the plurality of sub-pixels is classified as either an analyte center or a non-center. The method includes training a neural network to generate outputs of the training examples, including iteratively optimizing a loss function that minimizes an error between the output and ground truth data, and updating parameters of the neural network based on the error; training a neural network to generate outputs of the training examples, and updating parameters of the neural network based on the error.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the method includes storing updated parameters of the neural network in memory upon error convergence after the last iteration to apply to further neural network based template generation and base calling. In one embodiment, in the ground truth data, all sub-pixels classified as analyte centers are assigned the same first predefined class score and all sub-pixels classified as non-centers are assigned the same second predefined class score. In one embodiment, in each output, each sub-pixel has a first predicted score that is an analyte center and a second predicted score that is non-center. In one embodiment, the loss function is a custom weighted binary cross entropy loss, which is minimized on a sub-pixel basis between the predicted scores and class scores of corresponding sub-pixels in the output and the ground truth. In one embodiment, the ground truth data identifies centers at the centroids of corresponding analytes among the analytes. In one embodiment, in the ground truth, all sub-pixels classified as analyte centers are assigned the same first predefined value and all sub-pixels classified as non-centers are assigned the same second predefined value. In one embodiment, the first and second predetermined values are intensity values, hi another embodiment, the first and second predetermined values are continuous values.
In one embodiment, the ground truth data identifies the spatial distribution of the analytes including at least one of analyte shape, analyte size, and/or analyte boundary as part of the associated analyte metadata. In one embodiment, the image data includes images in a sequence of image sets, the images having a resolution of 1800x1800. In one embodiment, the image data includes at least one image patch from each of the images in the sequence of image sets, the image patch covering a portion of a tile, and having a resolution of 20x20. In one embodiment, the image data includes an upsampled representation of an image patch from each of the images in the sequence of image sets, the upsampled representation of the image patch having a resolution of 80x80. In one embodiment, in the training data, the multiple training examples are each represented as a different image patch of image data from each image in the sequence of image sets of the same tile, and at least a portion of the different image patches overlap each other. In one embodiment, the ground truth data has an upsampled resolution of 80x80. In one embodiment, the training data includes training examples of multiple tiles of a flow cell. In one embodiment, the training data includes training examples of various flow cells, sequencing installations, sequencing protocols, sequencing chemistries, sequencing reagents, and analyte densities. In one embodiment, the neural network is a deep fully convolutional segmentation neural network having an encoder sub-network and a corresponding decoder network followed by a classification layer, where the encoder sub-network includes a hierarchy of encoders and the decoder sub-network includes a hierarchy of decoders that map low resolution encoder feature maps to full input resolution feature maps for sub-pixel-by-subpixel classification by the classification layer.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Three-way classification model)
The inventors have disclosed a computer-implemented method for identifying analytes on a tile of a flow cell and associated analyte metadata. The method includes processing input image data from a sequence of image sets through a neural network and generating alternative representations of the image data. Each image in the sequence of image sets covers a tile and shows the intensity emission of analytes on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. The method includes processing the alternative representations through a classification layer and generating an output that identifies the spatial distribution of the analytes represented by the input image data and their surrounding background, including at least one of analyte center, analyte shape, analyte size, and/or analyte boundary. The output has a plurality of subpixels, each subpixel in the plurality of subpixels is classified as either background, analyte center, or analyte interior.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the classification layer assigns each subpixel in the output a first likelihood score that is background, a second likelihood score that is analyte center, and a third likelihood score that is analyte interior. In one embodiment, the first, second, and third likelihood scores are determined based on a softmax function and are exponentially normalized between zero and one. In one embodiment, each subpixel in the output is classified as either background, analyte center, or analyte interior based on which one of the first, second, and third likelihood scores is highest. In one embodiment, each subpixel in the output is classified as either background, analyte center, or analyte interior based on whether the first, second, and third likelihood scores are above a predetermined threshold likelihood score. In one embodiment, the output identifies an analyte center at the center of mass of a corresponding one of the analytes. In one embodiment, at the output, all subpixels classified as background are assigned the same first predetermined value, all subpixels classified as analyte center are assigned the same second predetermined value, and all subpixels classified as analyte interior are assigned the same third predetermined value. In one embodiment, the first, second, and third predetermined values are intensity values. In one embodiment, the first, second, and third predetermined values are continuous values.
In one embodiment, the method includes determining location coordinates of sub-pixels classified as analyte centers on an analyte basis, downscaling the location coordinates by an upsampling factor used to prepare the input image data, and storing the downscaled location coordinates in an analyte-based memory by an analyte basis for use on an analyte calling basis. In one embodiment, the method includes determining location coordinates of sub-pixels classified as analyte interior on an analyte basis, downscaling the location coordinates by an upsampling factor used to prepare the input image data, and storing the downscaled location coordinates in an analyte-based memory by an analyte basis for use on an analyte calling basis. In one embodiment, the method includes determining a distance of a sub-pixel classified as analyte interior from a corresponding one of the sub-pixels classified as analyte centers on an analyte basis, and storing the distance in an analyte-based memory by an analyte basis for use on an analyte calling basis. In one embodiment, the method includes extracting intensities from subpixels classified as analyte-internal on an analyte basis, including using at least one of nearest neighbor intensity extraction, Gaussian intensity extraction, intensity extraction based on average 2x2 subpixel area, intensity extraction based on brightest test of 2x2 subpixel areas, intensity extraction based on average 3x3 subpixel area, bilinear intensity extraction, quadratic intensity extraction, and/or intensity extraction based on weighted area coverage, and/or intensity extraction.
In one embodiment, the input image data includes images in a sequence of an image set, the images having a resolution of 3000x3000. In one embodiment, the input image data includes at least one image patch from each of the images in the sequence of an image set, the image patch covering a portion of a tile and having a resolution of 20x20. In one embodiment, the input image data includes an upsampled representation of an image patch from each of the images in the sequence of an image set, the upsampled representation having a resolution of 80x80. In one embodiment, the output has an upsampled resolution of 80x80. In one embodiment, the neural network is a deep full convolutional segmentation neural network having an encoder sub-network and a corresponding decoder network, followed by a classification layer, the encoder sub-network including a hierarchy of encoders and the decoder sub-network including a hierarchy of decoders that map the low resolution encoder feature maps to full input resolution feature maps for sub-pixel-by-subpixel classification by the classification layer. In one embodiment, the density of the specimens is about 100,000 specimens/mm<sup>2</sup>~Approx. 1,000,000 specimens/mm<sup>2</sup>In another embodiment, the density of the specimens is in the range of about 1,000,000 specimens/mm<sup>2</sup>~Approx. 10,000,000 specimens/mm<sup>2</sup>The range is.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Training a terminal class classification model)
The inventors have disclosed a computer-implemented method for training a neural network to identify analytes and associated analyte metadata. The method includes obtaining training data for training the neural network. The training data includes a plurality of training examples and corresponding ground truth data to be generated by the neural network by processing the training examples. Each training example includes image data from a sequence of image sets. Each image in the sequence of image sets covers a tile of a flow cell and shows intensity emission of analytes on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. Each ground truth data specifies a spatial distribution of analytes and their surrounding background represented by the input image data, including analyte center, analyte shape, analyte size, and analyte boundary. The ground truth data has a plurality of sub-pixels, and each sub-pixel in the plurality of sub-pixels is classified as either background, analyte center, or analyte interior. The method includes training a neural network to generate outputs of the training examples, including iteratively optimizing a loss function that minimizes an error between the output and ground truth data, and updating parameters of the neural network based on the error; training a neural network to generate outputs of the training examples, and updating parameters of the neural network based on the error.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the method includes, upon error convergence after the last iteration, storing the updated parameters of the neural network in the memory for application to further neural network based template generation and base calling. In one embodiment, in the ground truth data, all subpixels classified as background are assigned the same first predetermined class score, all subpixels classified as analyte center are assigned the same second predetermined class score, and all subpixels classified as analyte interior are assigned the same third predetermined class score.
In one embodiment, in each output, each subpixel has a first predicted score that is background, a second predicted score that is analyte center, and a third predicted score that is analyte interior. In one embodiment, the loss function is a custom weighted ternary cross entropy loss that is minimized on a subpixel basis between the predicted scores and the class scores of corresponding subpixels in the output and the ground truth. In one embodiment, the ground truth data identifies analyte centers at the center of mass of corresponding analytes in the analytes. In one embodiment, in the ground truth, all subpixels classified as background are assigned the same first predetermined value, all subpixels classified as analyte center are assigned the same second predetermined value, and all subpixels classified as analyte interior are assigned the same third predetermined value. In one embodiment, the first, second, and third predetermined values are intensity values. In one embodiment, the first, second, and third predetermined values are continuous values. In one embodiment, the image data includes images in a sequence of image sets, and the images have a resolution of 1800x1800. In one embodiment, the image data includes images in a sequence of an image set, and the images have a resolution of 1800x1800.
In one embodiment, the image data includes at least one image patch from each of the images in the sequence of the image set, the image patch covering a portion of a tile, and having a resolution of 20x20. In one embodiment, the image data includes an upsampled representation of an image patch from each of the images in the sequence of the image set, the upsampled representation of the image patch having a resolution of 80x80. In one embodiment, in the training data, the multiple training examples are each represented as a different image patch of image data from each image in the sequence of the image set of the same tile, and at least a portion of the different image patches overlap each other. In one embodiment, the ground truth data has an upsampled resolution of 80x80. In one embodiment, the training data includes training examples of multiple tiles of a flow cell. In one embodiment, the training data includes training examples of various flow cells, sequencing installations, sequencing protocols, sequencing chemistries, sequencing reagents, and analyte densities. In one embodiment, the neural network is a deep fully convolutional segmentation neural network having an encoder sub-network and a corresponding decoder network followed by a classification layer, where the encoder sub-network includes a hierarchy of encoders and the decoder sub-network includes a hierarchy of decoders that map low-resolution encoder feature maps to full input resolution feature maps for sub-pixel classification by the classification layer.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Segmentation)
We disclose a computer-implemented method for determining analyte metadata. The method includes processing input image data derived from a set of sequential images through a neural network and generating alternative representations of the input image data. The input image data has an array of units depicting analytes and their the output values of the units and classifying a first subset of the units as background units depicting the surrounding background. The method includes locating peaks in the output values of the units and classifying a second subset of the units as central units containing the center of the analyte. The method includes applying a divider to the output values of the units and determining the shape of the analyte as a non-overlapping region of consecutive units centered on the central unit, separated by background units. The segmentation starts from the central unit and for each central unit determines a group of consecutively consecutive units indicative of the same analyte whose centers are contained in the central unit.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one implementation, the units are pixels. In another implementation, the units are subpixels. In yet another implementation, the units are superpixels. In one implementation, the output value is a continuous value. In another implementation, the output value is a softmax score. In one implementation, consecutive units in corresponding ones of the non-overlapping regions have output values weighted according to the distance of the consecutive units from a central unit in the non-overlapping region to which the adjacent units belong. In one implementation, the central unit has the highest output value in its respective one of the non-overlapping regions.
In one embodiment, the non-overlapping regions have an irregular contour and the unit is a sub-pixel. In such an embodiment, the method includes determining the analyte intensity of the given analyte by identifying sub-pixels that contribute to the analyte intensity of the given analyte based on corresponding non-overlapping regions of contiguous sub-pixels that identify the shape of the given analyte, locating the identified sub-pixels in one or more optical pixel resolution images generated for one or more image channels in a current sequencing cycle, interpolating the intensities of the identified sub-pixels in each of the images, combining the interpolated intensities and normalizing the combined interpolated intensities to generate an analyte intensity per image for the given analyte in each of the images, and combining the analyte intensities per image for each of the images to determine the analyte intensity of the given analyte in the current sequencing cycle. In one embodiment, the normalization is based on a normalization factor, which is the number of identified sub-pixels. In one embodiment, the method includes calling the given analyte based on the analyte intensity in the current sequencing cycle.
In one embodiment, the non-overlapping regions have an irregular contour and the unit is a sub-pixel. In such an embodiment, the method includes determining an analyte intensity of a given analyte by identifying sub-pixels that contribute to the analyte intensity of the given analyte based on corresponding non-overlapping regions of contiguous sub-pixels that identify the shape of the given analyte, locating the identified sub-pixels in one or more sub-pixel resolution images upsampled from the corresponding optics, combining intensities of the identified sub-pixels in each of the upsampled images and normalizing the combined intensities to generate an analyte intensity per image for the given analyte in each of the upsampled images, and combining the analyte intensity per image for each of the upsampled images to determine an analyte intensity of the given analyte in the current sequencing cycle. In one embodiment, the normalization is based on a normalization factor, which is the number of identified sub-pixels. In one embodiment, the method includes calling the given analyte based on the analyte intensity in the current sequencing cycle.
In one embodiment, each image in the sequence of image sets covers a tile and shows the intensity emission of the analytes on the tile and their surrounding background captured for a particular image channel at a particular one of multiple sequencing cycles of a sequencing run performed on the flow cell. In one embodiment, the input image data includes at least one image patch from each of the images in the sequence of image sets, the image patch covering a portion of a tile and having a resolution of 20x20. In one embodiment, the input image data includes an upsampled sub-pixel resolution representation of the image patch from each of the images in the sequence of image sets, the upsampled sub-pixel representation having a resolution of 80x80.
In one embodiment, the neural network is a convolutional neural network. In another embodiment, the neural network is a recurrent neural network. In yet another embodiment, the neural network is a residual neural network with residual Bock and residual connections. In yet another embodiment, the neural network is a deep full convolutional segmentation neural network having an encoder sub-network and a corresponding decoder network, where the encoder sub-network includes a hierarchy of encoders and the decoder sub-network includes a hierarchy of decoders that map low resolution encoder feature maps to full input resolution feature maps.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Peak detection)
We disclose a computer-implemented method for determining analyte metadata. The method includes processing input image data derived from a set of sequential images through a neural network and generating alternative representations of the input image data. The input image data has an array of units that depict analytes and their surrounding background. The method includes processing the alternative representations through an output layer to generate output values for each unit in the array. The method includes thresholding the output values of the units and classifying a first subset of the units as background units that depict the surrounding background. The method includes locating a peak in the output values of the units and classifying a second subset of the units as center units that include a center of the analyte.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one embodiment, the method includes applying a divider to the output values of the units and determining the shape of the analyte as a non-overlapping region of consecutive units separated by background units and centered on a central unit. The segments start from the central unit and for each central unit, determine a group of consecutively consecutive units whose centers represent the same analyte contained in the central unit.
Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
Neural network based analytical data generator
In one embodiment, the method includes processing image data through a neural network and generating an alternative representation of the image data. The image data indicates the intensity emission of the analytes. The method includes processing the alternative representation through an output layer and generating an output that identifies metadata about the analytes, including at least one of the spatial distribution of the analytes, the shape of the analytes, the center of the analytes, and/or the boundaries between the analytes. Each of the features described in the specific embodiment section for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the set of base features specified in other embodiments. Other embodiments of the methods described in this section can include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another embodiment of the methods described in this section can include a system including a memory and one or more processors operable to execute instructions stored in the memory, and can perform any of the methods described above.
(Unit-based regression model)
The inventors have disclosed a computer-implemented method for identifying analytes on tiles of a flow cell and associated analyte metadata. The method includes processing input image data from a sequence of image sets through a neural network to generate alternative representations of the input image data. Each image in the sequence of image sets covers a tile and shows the intensity emission of analytes on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. The method includes processing the alternative representations through an output layer and generating an output in which the intensity emission is represented by the input image data as discontinuous regions of adjacent units, the centers of the analytes as central units at the center of mass of each one of the junction regions, and their surrounding background as background units that do not belong to any of the discontinuous regions.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one implementation, the unit is a pixel. In another implementation, the unit is a subpixel. In yet another embodiment, the unit is a superpixel. Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Still other implementations of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Unit-based joint classification model)
The inventors disclose a computer-implemented method for identifying analytes on a tile of a flow cell and associated analyte metadata. The method includes processing input image data from a sequence of image sets through a neural network and generating alternative representations of the image data. Each image in the sequence of image sets covers a tile and shows the intensity emission of an analyte on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. The method includes processing the alternative representations through a classification layer and generating an output that identifies a center of an analyte whose intensity emission is indicated by the input image data. The output has a plurality of units, each unit in the plurality of units being classified as either an analyte center or a non-center.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one implementation, the unit is a pixel. In another implementation, the unit is a subpixel. In yet another embodiment, the unit is a superpixel. Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Still other implementations of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
(Unit-based ternary classification model)
The inventors disclose a computer-implemented method for identifying analytes on a tile of a flow cell and associated analyte metadata. The method includes processing input image data from a sequence of image sets through a neural network and generating alternative representations of the image data. Each image in the sequence of image sets covers a tile and shows the intensity emission of analytes on the tile and their surrounding background captured for a particular image channel at a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell. The method includes processing the alternative representations through a classification layer and generating an output that identifies the spatial distribution of the analytes represented by the input image data and their surrounding background, including at least one of analyte centers, analyte shapes, analyte sizes, and/or analyte boundaries. The output has a plurality of units, each unit in the plurality of units being classified as either background, analyte center, or analyte interior.
Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features specified in these embodiments can be easily combined with the base feature sets specified in other embodiments.
In one implementation, the unit is a pixel. In another implementation, the unit is a subpixel. In yet another embodiment, the unit is a superpixel. Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Still other implementations of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory, and may perform any of the methods described above.
item
The present inventors disclose the following items:
Item Set 1
1. A computer-implemented method for determining image regions indicative of analytes on a tile of a flow cell, the method comprising: accessing a series of image sets generated during a sequencing run, each image set being a series generated during a respective sequencing cycle of the sequencing run, each image in the series depicting analytes and their surrounding background, each image in the series having a plurality of sub-pixels; obtaining base calls from the base calls that classify each of the sub-pixels, thereby generating a base call sequence for each of the sub-pixels over a plurality of sequencing cycles of the sequencing run; determining a plurality of discontinuous regions of contiguous sub-pixels that share substantially matching base call sequences; and generating an analyte map identifying the determined discontinuous regions.
2. The computer-implemented method of claim 1, further comprising training a classifier based on the determined plurality of discontinuous regions of contiguous subpixels, the classifier being a neural network-based template generator for processing the input image data to generate an attenuation map, a ternary map representing one or more characteristics of each of a plurality of analytes represented in the input image data for base calling by a neural network-based base station, or preferably a neural network-based binary map for increasing the level of throughput in high-throughput nucleic acid sequencing techniques.
3. The computer-implemented method of any one of clauses 1-2, further comprising generating an analyte map by identifying subpixels that do not belong to any of the discontinuous regions as background.
4. The computer-implemented method of any one of items 1-3, wherein the analyte map identifies an analyte boundary between two consecutive subpixels where the base call sequences do not substantially match.
5. The computer-implemented method of any one of items 1-4, wherein determining a plurality of discontinuous regions of contiguous subpixels further comprises identifying an origin subpixel in preliminary center coordinates of the specimen determined by the base caller, and performing a breadth-first search for a substantially matching base call sequence by starting from the origin subpixel and continuing through consecutive contiguous non-origin subpixels.
6. The computer-implemented method of any one of items 1-5, further comprising determining the hyperlocation center coordinates of the analyte by calculating the center of mass of the discontinuous region of the analyte map as the average of the coordinates of each contiguous sub-pixel that forms the discontinuous region, and storing the hyperlocation center coordinates of the analyte in memory for use as ground truth for training the classifier.
7. The computer-implemented method of claim 6, further comprising identifying centers of mass subpixels within the discontinuous regions of the analyte map at analyte hyperlocation center coordinates; upsampling the analyte map using interpolation and storing the upsampled analyte map in memory for use as ground truth for training a classifier; and in the upsampled analyte map, assigning a value to each contiguous subpixel within the discontinuous region based on a decay coefficient proportional to the distance of the contiguous subpixel from the center of mass subpixel within the discontinuous region to which the adjacent subpixel belongs.
8. The computer-implemented method of item 7, wherein the method further preferably includes generating an attenuation map from the upsampled analyte map representing contiguous subpixels within the isolated region and representing subpixels identified as background based on their respective assigned values, and storing the attenuation map in memory for training the classifier.
9. The computer-implemented method of claim 8, further comprising the steps of: classifying, on an analyte basis, in the upsampled analyte map, contiguous subpixels within discontinuous regions as analyte interior subpixels belonging to the same analyte; classifying the center of mass subpixels as analyte center subpixels; subpixels containing analyte boundary portions as boundary subpixels; and subpixels identified as background as background subpixels; and storing the classifications in memory for training the classifier.
10. The computer-implemented method of any one of items 1-9, further comprising: storing in memory, based on the analyte, analyte interior subpixels, analyte center subpixels, boundary subpixels, and background subpixels for use as ground truth for training the classifier; downscaling the coordinates by a factor used to upsample the analyte map; and storing the downscaled coordinates in memory based on the analyte for use as ground truth for training the classifier.
11. The computer-implemented method of any one of items 1-10, further comprising: in binary ground truth data generated from the upsampled analyte map, labeling analyte center subpixels as belonging to an analyte center class and all other subpixels as belonging to non-center classes; and storing the binary ground truth data in memory for training a classifier.
12. A computer-implemented method according to any one of items 1-11, further comprising labeling background subpixels as belonging to a background class, labeling analyte center subpixels as belonging to an analyte center class, and analyte interior subpixels as belonging to an analyte interior class in ternary ground truth data generated from the upsampled analyte map, and storing the ternary ground truth data in memory for use as ground truth for training the classifier.
13. The computer-implemented method of any one of items 1-12, further comprising: generating an analyte map for a plurality of tiles of a flow cell; storing the analyte map in a memory; determining a spatial distribution of analytes within the tiles based on the analyte map, including their shapes and sizes; classifying, in the upsampled analyte map of analytes in the tiles, analyte interior subpixels belonging to the same analyte, analyte center subpixels, boundary subpixels, and background subpixels on an analyte by analyte basis; storing the classification in memory to train a classifier; storing the analyte interior subpixels, analyte center subpixels, boundary subpixels, and background subpixels on an analyte by analyte basis to train a classifier; downscaling the coordinates by a factor used to upsample the analyte map; and storing the downscaled coordinates across the tiles based on analyte in memory to train a classifier.
14. The computer-implemented method of any of items 1-13, wherein the base call sequences substantially match when predetermined portions of the base calls match per sequence position.
15. The computer-implemented method of any of items 1-14, wherein determining a plurality of discontinuous regions of contiguous subpixels that share substantially matching base call sequences is based on a predetermined minimum number of subpixels for the discontinuous regions.
16. The computer-implemented method of any of items 1-15, wherein the flow cell has at least one patterned surface having an array of wells that occupy analytes, and further comprising determining, based on the determined shapes and sizes of the analytes, that one of the wells is substantially occupied by at least one analyte, one of the wells is minimally occupied, and one of the wells is co-occupied by multiple analytes.
17. A computer-implemented method for determining metadata regarding analytes on tiles of a flow cell, the method comprising: accessing a set of images of the tile captured during a sequencing run; accessing preliminary centroid coordinates of the analytes determined by a base caller; for each set of images, obtaining a base call classification from the base caller as one of four bases; generating a base call sequence for an origin subpixel that includes the preliminary centroid coordinate and a predetermined neighborhood of contiguous subpixels that are contiguous to a corresponding one of the origin subpixels, thereby generating an analyte map that identifies the analyte as a discontinuous region of contiguous subpixels that are contiguous to at least a portion of each of the origin subpixels and share a substantially matching base call sequence of one of the four bases with at least a portion of each one of the origin subpixels; storing the analyte map in a memory; and determining a shape and size of the analyte based on the discontinuous regions in the analyte map.
18. A computer-implemented method for generating training data for neural network-based template generation and base calling, comprising: accessing a multiplicity of images of a flow cell captured over multiple cycles of a sequencing run, where the flow cell has a multiplicity of tiles, where each of the tiles in the multiplicity of images has a series of image sets generated over a multiplicity of cycles, constructing a training set having a multiplicity of training examples, where each image in the sequence of image sets shows an intensity emission of an analyte of a particular one of the tiles and their surrounding background in a particular one of the cycles, where each training example corresponds to a particular one of the tiles and includes image data from at least some of the image sets in the sequence of image sets of the particular one of the tiles; and generating at least one ground truth data representation for each of the training examples, where the ground truth representation is indicated by the image data and is determined, at least in part, using the method of any of items 1-17.
19. The computer-implemented method of item 18, wherein the at least one characteristic of the analyte is selected from the group consisting of a spatial distribution of the analyte on the tile, an analyte shape, an analyte size, an analyte boundary, and a center of a contiguous area containing a single analyte.
20. A computer-implemented method according to any of items 18-19, wherein the image data includes each image of at least some of the image sets in the sequence of the image set of a particular one of the tiles.
21. The computer-implemented method of any one of items 18-20, wherein the image data includes at least one image patch from each of the images.
22. The computer-implemented method of any of items 18-21, wherein the image data comprises an upsampled representation of the image patch.
23. A computer-implemented method according to any of items 18-22, wherein a plurality of training examples correspond to the same particular one of the tiles and each include different image patches from respective images of at least some of the image sets within the sequence of image sets of the same particular one of the tiles, and at least some of the different image patches overlap each other.
24. The computer-implemented method of any one of items 18-23, wherein the ground truth data representation identifies analytes as discontinuous regions of adjacent subpixels, with centers of analytes as centers of mass subpixels in corresponding ones of the discontinuous regions, and their surrounding background, as subpixels that do not belong to any of the discontinuous regions.
25. The computer-implemented method of any one of items 18-24, further comprising storing the training examples in the training set and associated ground truth data representation as training data for neural network-based template generation and base calling.
26. A computer-implemented method, comprising: accessing sequence images of a specimen generated by a sequencer; generating training data from the sequence images; and using the training data to train a neural network to generate metadata about the analyte.
27. A computer-implemented method, comprising: accessing sequence images of analytes generated by a sequencer; generating training data from the sequence images; and using the training data to train a neural network to call the analytes into a base.
28. A computer-implemented method for determining image regions indicative of analytes on a tile of a flow cell, the method comprising: accessing a series of image sets generated during a sequencing run, each image set being a series generated during a respective sequencing cycle of the sequencing run, each image in the series depicting analytes and their surrounding background, each image in the series having a plurality of sub-pixels; obtaining base calls from the base calls that classify each of the sub-pixels, thereby generating a base call sequence for each of the sub-pixels over a plurality of sequencing cycles of the sequencing run; and determining a plurality of discontinuous regions of adjacent sub-pixels that share substantially matching base call sequences.
Item Set 2
1. A computer-implemented method of generating ground truth training data to train a neural network-based template generator for a cluster metadata determination task, comprising: accessing a series of image sets generated during a sequencing run, each image set generated during a respective sequencing cycle of the sequencing run, the series of images depicting clusters and their surrounding background, each of the pixels being divided into a plurality of sub-pixels in a sub-pixel domain, obtaining base calls classifying each of the sub-pixels as one of four bases (A, C, T, and G), thereby generating base call sequences for each of the sub-pixels over a plurality of sequencing cycles of the sequencing operation; generating a cluster map that identifies clusters as discontinuous regions of adjacent sub-pixels that share substantially matching base call sequences; and based on the discontinuous regions in the cluster map, generating a base call sequence for each of the sub-pixels. and determining cluster metadata using the cluster metadata to generate ground truth training data for training a neural network-based template generator for a cluster metadata determination task, the ground truth training data including attenuation maps, ternary maps, or binary maps, the neural network-based template generator being trained based on the ground truth training data to generate the attenuation maps, ternary maps, or binary maps as output, and upon performance of the cluster metadata determination task during inference, the cluster metadata is then determined from the attenuation maps, ternary maps, or binary maps generated as output by the trained neural network-based template generator.
2. The computer-implemented method of claim 1, further comprising using cluster metadata derived from an attenuation map, a ternary map, or a binary map generated as output by a neural network-based template generator for base calling by a neural network-based base caller to increase throughput in high-throughput nucleic acid sequencing techniques.
3. The computer-implemented method of claim 1, further comprising generating the cluster map by identifying sub-pixels that do not belong to any of the discontinuous regions as background.
4. The computer-implemented method of claim 1, wherein the cluster map identifies cluster boundaries between two consecutive subpixels where the base call sequences do not substantially match.
5. The computer-implemented method of claim 1, wherein the cluster map is generated by identifying an origin subpixel at preliminary center coordinates of a cluster determined by the base caller, and performing a breadth-first search for substantially matching base call sequences by starting from the origin subpixel and continuing through consecutive non-origin subpixels.
6. The computer-implemented method of claim 1, further comprising: determining the hyperlocation center coordinates of the clusters by calculating the center of mass of the discontinuous regions of the cluster map as the average of the coordinates of each contiguous sub-pixel that forms the discontinuous region; and storing the hyperlocation center coordinates of the clusters in memory for use as ground truth training data for training a neural network-based template generator.
7. The computer-implemented method of claim 6, further comprising: identifying centers of mass subpixels in non-connected regions of the cluster map at the hyper-location center coordinates of the clusters; up-sampling the cluster map using interpolation and storing the up-sampled cluster map in memory for use as ground truth training data for training a neural network based template generator; and in the up-sampled cluster map, assigning a value to each contiguous subpixel in the discontinuous region based on a decay coefficient proportional to the distance of the adjacent subpixel from the center of the mass subpixel in the discontinuous region to which the adjacent subpixel belongs.
8. The computer-implemented method of claim 7, further comprising: generating an attenuation map from the upsampled cluster map, the attenuation map representing contiguous subpixels in discontinuous regions, the subpixels being identified as background based on their assigned values; and storing the attenuation map in memory for use as ground truth training data for training a neural network-based template generator.
9. The computer-implemented method of claim 8, further comprising: classifying, for each cluster in the upsampled cluster map, consecutive subpixels within isolated regions as cluster-interior subpixels belonging to the same cluster; classifying center of mass subpixels as cluster central subpixels, subpixels containing cluster boundary portions, and subpixels identified as background as background subpixels; and storing the classifications in memory for use as ground truth training data for training a neural network-based template generator.
10. The computer-implemented method of claim 9, further comprising: for each cluster, storing cluster interior subpixels, cluster center subpixels, boundary subpixels, and background subpixels in memory for use as ground truth training data for training a neural network based template generator; downscaling the coordinates by a factor used to upsample the cluster map; and for each cluster, storing the downscaled coordinates in memory for use as ground truth training data for training the neural network based template generator.
11. The computer-implemented method of claim 10, comprising: generating a cluster map for a plurality of tiles of a flow cell; storing the cluster map in a memory; determining cluster metadata for clusters within the tiles based on the cluster map, including cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and/or cluster boundaries; classifying sub-pixels per cluster in the upsampled cluster map for clusters within the tiles into sub-pixels as cluster interior sub-pixels, cluster center sub-pixels, boundary sub-pixels, and background sub-pixels that belong to the same cluster; and training a neural network-based template generator. storing the classification in memory for use as ground truth training data; for each cluster, storing coordinates of cluster interior subpixels, cluster center subpixels, boundary subpixels, and background subpixels in memory for use as ground truth training data for training a neural network based template generator; downscaling the coordinates by a factor used to upsample the cluster map; and for each cluster across the tiles, storing the downscaled coordinates in memory for use as ground truth training data for training a neural network based template generator.
12. The computer-implemented method of claim 11, wherein the base call sequences substantially match when predetermined portions of the base calls match sequence position by sequence position.
13. The computer-implemented method of claim 1, wherein the cluster map is generated based on a predetermined minimum number of subpixels for discontinuous regions.
14. The computer-implemented method of claim 1, wherein the flow cell has at least one patterned surface having an array of wells occupied by clusters, and further comprising determining, based on the determined shapes and sizes of the clusters, one of the wells is substantially occupied by at least one cluster, one of the wells is minimally occupied, and one of the wells is co-occupied by multiple populations.
15. A computer-implemented method for determining metadata for clusters on a tile of a flow cell, the method comprising: accessing a set of images of the tile captured during a sequencing run; accessing preliminary center coordinates of the clusters determined by a base caller; for each set of images, obtaining from the base caller a base call classification as one of four bases; generating a base call sequence for an origin subpixel that includes the preliminary center coordinate and a predetermined neighborhood of contiguous subpixels that are contiguous to a corresponding one of the origin subpixels, thereby generating a cluster map that identifies clusters as discontinuous regions of adjacent subpixels that are contiguous to at least a portion of each of the origin subpixels and share a substantially matching base call sequence of one of the four bases with at least a portion of each one of the origin subpixels; storing the cluster map in a memory; and determining a shape and size of the cluster based on the discontinuous regions in the cluster map.
16. A computer-implemented method for generating training data for neural network-based template generation and base calling, comprising: accessing a number of images of a flow cell captured over multiple cycles of a sequencing run, where the flow cell has a number of tiles, where each of the tiles in the number of images has a series of image sets generated over a number of cycles, where each image in the sequence of image sets shows intensity emissions of clusters and their surrounding background at a particular cycle, on a particular one of the tiles at the particular cycle; constructing a training set having a number of training examples, where each training example corresponds to a particular one of the tiles and includes image data from at least some of the image sets in the sequence of image sets for the particular one of the tiles; and generating at least one ground truth data representation for each of the training examples, where the ground truth data representation identifies a characteristic of an analyte for the particular one of the tiles whose intensity emissions are depicted by the image data.
17. The computer-implemented method of claim 16, wherein at least one characteristic of the clusters is selected from the group consisting of: spatial distribution of the clusters on the tile, cluster shape, cluster size, cluster boundary, and center of a contiguous region containing a single cluster.
18. The computer-implemented method of claim 16, wherein the image data includes images of each of at least some of the image sets in the sequence of image sets for a particular one of the tiles.
19. The computer-implemented method of claim 18, wherein the image data includes at least one image patch from each of the images.
20. The computer-implemented method of claim 19, wherein the image data comprises an upsampled representation of the image patch.
21. The computer-implemented method of claim 16, wherein the multiple training examples correspond to the same particular one of the tiles and each include a different image patch from a respective image of at least some of the image sets in the sequence of image sets of the same particular one of the tiles, and at least some of the different image patches overlap each other.
22. The computer-implemented method of claim 16, wherein the ground truth data representation identifies clusters as discontinuous regions of adjacent sub-pixels, with cluster centers identified as centers of mass sub-pixels within a corresponding one of the discontinuous regions, and surrounding background as sub-pixels that do not belong to any of the discontinuous regions.
23. The computer-implemented method of claim 16, further comprising storing the training examples in the training set and associated ground truth data representations as training data for neural network-based template generation and base calling.
24. A computer-implemented method, comprising: accessing sequence images of clusters generated by a sequencer; generating training data from the sequence images; and using the training data to train a neural network to generate metadata about the clusters.
25. A computer-implemented method, comprising: accessing sequence images of clusters generated by a sequencer; generating training data from the sequence images; and using the training data to train a neural network, calling the clusters.
26. A computer-implemented method for determining image regions indicative of analytes on a tile of a flow cell, the method comprising: accessing a series of image sets generated during a sequencing run, each image set being a series generated during a respective sequencing cycle of the sequencing run, each image in the series depicting analytes and their surrounding background, each image in the series having a plurality of sub-pixels; obtaining base calls from the base calls that classify each of the sub-pixels, thereby generating a base call sequence for each of the sub-pixels over a plurality of sequencing cycles of the sequencing run; determining a plurality of discontinuous regions of contiguous sub-pixels that share substantially matching base call sequences; and generating a cluster map identifying the determined discontinuous regions.
1510 Training equipment
1512 Neural Network Based Template Generator
1514 Neural network based base caller
1802 Thresholder
1806 Peak Locator
1810 Divider
1814 Aftertreatment Device
1902 Intensity Extractor
1904 Subpixel Locator
1906 Interpolator and Subpixel Intensity Combiner
1908 Normalizer
1910 Cross-Channel Sub-Pixel Intensity Accumulator
2002 Intensity Extractor
2004 Subpixel Locator
2006 Subpixel Intensity Combiner
2008 Normalizer
2010 Cross-Channel Sub-Pixel Intensity Accumulator
2302 Upsampler
2600 Regression Model
3102 Watershed Divider
4600 Binary Classification Models
5400 Three-way classification model
7804 Temperature Control System
7806 System Controller
7808 Fluid Control Systems
7812 Biosensors
7814 Fluid Storage System
7816 Lighting System
7818 User Interface
7824 Main Control Module
7826 Lighting Module
7828 Fluid Control Module
7830 Fluid Storage Module
7832 Temperature Control Module
7836 Equipment Module
7838 Identification Module
7840 SBS Module
7842 Amplification Module
7844 Analysis Module
7846 Configurable Processor
7848 Memory
7848A Memory (bit files, base caller parameters)
7848B memory (image data, base call reads)
7848C Memory (tile data, model parameters)
7852 CPU (execution time)
7908 Data Flow Logic
7914 Multi-cycle execution cluster 1
8200 Computer Systems
8210 Storage Subsystem
8222 Memory Subsystem
8236 File Storage Subsystem
8238 User Interface Input Devices
8255 Bus Subsystem
8274 Network Interface Subsystem
8276 User Interface Output Device
8278 Deep Learning Processors (GPU, FPGA, CGRA)
88 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| US20120020537A1 | Cites | United States of America |
| WO2018129314A1 | Cites | World Intellectual Property Organization (WIPO) |
| WO2018068014A1 | Cites | World Intellectual Property Organization (WIPO) |
| “MiSeq: Imaging and Base Calling”,[online],2013年,インターネット<URL: https://support.illumina.com/content/dam/illumina-support/courses/MiSeq_Imaging_and_Base_Calling/story_content/external_files/MiSeq%20Imaging%20and%20Base%20Calling%20Script.pdf>,[検索日:2024.05.08] | Non-patent | – |
100 members in 14 offices
Priority claims31
| Document | Office | Kind | Date |
|---|---|---|---|
| 62821602 | United States of America | – | |
| 62821618 | United States of America | – | |
| 62821681 | United States of America | – | |
| 62821724 | United States of America | – | |
| 62821766 | United States of America | – | |
| 201962821602 | United States of America | P | |
| 201962821618 | United States of America | P | |
| 201962821681 | United States of America | P | |
| 201962821724 | United States of America | P | |
| 201962821766 | United States of America | P | |
| 2023310 | Netherlands (Kingdom of the) | – | |
| 2023311 | Netherlands (Kingdom of the) | – | |
| 2023312 | Netherlands (Kingdom of the) | – | |
| 2023314 | Netherlands (Kingdom of the) | – | |
| 2023316 | Netherlands (Kingdom of the) | – | |
| 2023310 | Netherlands (Kingdom of the) | A | |
| 2023311 | Netherlands (Kingdom of the) | A | |
| 2023312 | Netherlands (Kingdom of the) | A | |
| 2023314 | Netherlands (Kingdom of the) | A | |
| 2023316 | Netherlands (Kingdom of the) | A | |
| 16825987 | United States of America | – | |
| 16825991 | United States of America | – | |
| 16826126 | United States of America | – | |
| 16826134 | United States of America | – | |
| 202016825987 | United States of America | A | |
| 202016825991 | United States of America | A | |
| 202016826126 | United States of America | A | |
| 202016826134 | United States of America | A | |
| 16826168 | United States of America | – | |
| 202016826168 | United States of America | A | |
| 2020024090 | United States of America | W |
Members100
| Document | Office | Kind | |
|---|---|---|---|
| CA3104951A1 | Canada | A1 | |
| US2020302223A1 | United States of America | A1 | |
| US2020302224A1 | United States of America | A1 | |
| US2020302225A1 | United States of America | A1 | |
| US2020302297A1 | United States of America | A1 | |
| WO2020191387A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2020191389A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2020191390A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2020191391A2 | World Intellectual Property Organization (WIPO) | A2 | |
| NL2023310B1 | Netherlands (Kingdom of the) | B1 | |
| NL2023311B1 | Netherlands (Kingdom of the) | B1 | |
| NL2023312B1 | Netherlands (Kingdom of the) | B1 | |
| NL2023314B1 | Netherlands (Kingdom of the) | B1 | |
| NL2023316B1 | Netherlands (Kingdom of the) | B1 | |
| WO2020205296A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2020327377A1 | United States of America | A1 | |
| WO2020191390A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2020191391A3 | World Intellectual Property Organization (WIPO) | A3 | |
| AU2020241905A1 | Australia | A1 | |
| AU2020256047A1 | Australia | A1 | |
| AU2020240141A1 | Australia | A1 | |
| AU2020241586A1 | Australia | A1 | |
| SG11202012441QA | Singapore | A | |
| SG11202012453PA | Singapore | A | |
| SG11202012461XA | Singapore | A | |
| SG11202012463YA | Singapore | A | |
| IL279522A | Israel | A | |
| IL279522D0 | Israel | D0 | |
| IL279525A | Israel | A | |
| IL279525D0 | Israel | D0 | |
| IL279527A | Israel | A | |
| IL279527D0 | Israel | D0 | |
| IL279533A | Israel | A | |
| IL279533D0 | Israel | D0 | |
| CN112313666A | China | A | |
| CN112334984A | China | A | |
| NL2023311B9 | Netherlands (Kingdom of the) | B9 | |
| BR112020026408A2 | Brazil | A2 | |
| BR112020026426A2 | Brazil | A2 | |
| BR112020026433A2 | Brazil | A2 | |
| BR112020026455A2 | Brazil | A2 | |
| MX2020014293A | Mexico | A | |
| MX2020014299A | Mexico | A | |
| CN112585689A | China | A | |
| AU2020240383A1 | Australia | A1 | |
| CN112689875A | China | A | |
| CN112789680A | China | A | |
| MX2020014288A | Mexico | A | |
| MX2020014302A | Mexico | A | |
| IL281668A | Israel | A | |
| IL281668D0 | Israel | D0 | |
| KR20210142529A | Republic of Korea | A | |
| KR20210143100A | Republic of Korea | A | |
| KR20210143154A | Republic of Korea | A | |
| KR20210145115A | Republic of Korea | A | |
| KR20210145116A | Republic of Korea | A | |
| US11210554B2 | United States of America | B2 | |
| EP3942070A1 | European Patent Office (EPO) | A1 | |
| EP3942071A1 | European Patent Office (EPO) | A1 | |
| EP3942072A1 | European Patent Office (EPO) | A1 | |
| EP3942073A2 | European Patent Office (EPO) | A2 | |
| EP3942074A2 | European Patent Office (EPO) | A2 | |
| JP2022524562A | Japan | A | |
| JP2022525267A | Japan | A | |
| US2022147760A1 | United States of America | A1 | |
| JP2022526470A | Japan | A | |
| US11347965B2 | United States of America | B2 | |
| JP2022532458A | Japan | A | |
| JP2022535306A | Japan | A | |
| US11436429B2 | United States of America | B2 | |
| US2022292297A1 | United States of America | A1 | |
| US2023004749A1 | United States of America | A1 | |
| US11676685B2 | United States of America | B2 | |
| US2023268033A1 | United States of America | A1 | |
| EP3942072B1 | European Patent Office (EPO) | B1 | |
| US11783917B2 | United States of America | B2 | |
| EP4276769A2 | European Patent Office (EPO) | A2 | |
| EP4276769A3 | European Patent Office (EPO) | A3 | |
| US11908548B2 | United States of America | B2 | |
| US2024071573A1 | United States of America | A1 | |
| US11961593B2 | United States of America | B2 | |
| IL279533B1 | Israel | B1 | |
| CN112313666B | China | B | |
| JP7566638B2 | Japan | B2 | |
| US12119088B2 | United States of America | B2 | |
| JP7581190B2 | Japan | B2 | |
| CN112689875B | China | B | |
| JP7604232B2This record | Japan | B2 | |
| IL279533B2 | Israel | B2 | |
| JP7608172B2 | Japan | B2 | |
| JP2025016472A | Japan | A | |
| US12217831B2 | United States of America | B2 | |
| CN112585689B | China | B | |
| US2025069704A1 | United States of America | A1 | |
| CN119626331A | China | A | |
| CN119694395A | China | A | |
| US12277998B2 | United States of America | B2 | |
| CN112789680B | China | B | |
| MY210241A | Malaysia | A | |
| JP7767012B2 | Japan | B2 |
8 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 7604232
- Application
- 572704
Titles2
- Japanese
- 人工知能ベースの配列決定のための訓練データ生成
- English
- Training Data Generation for Artificial Intelligence-Based Sequencing
Classification
- CPC, 38
- G16B30/10
- G16B40/20
- G06F18/23211
- G16B40/00
- G06F16/58
- G06N3/084
- G06V10/454
- G06V10/993
- G16B30/20
- G06V10/82
- G06V10/267
- G06V20/47
- G06V20/69
- G06V2201/03
- G16B40/10
- G06V10/7715
- G06V10/763
- G06V10/7784
- G06V10/764
- G06N3/048
- G06N3/044
- G06N3/045
- G06N3/0455
- G06N3/0464
- G06N3/09
- G06N3/08
- G06F18/24
- G16B20/20
- G06F16/907
- G06N3/04
- G06N5/046
- G06F18/23
- G06F18/214
- G06F18/217
- G06F18/2415
- G06F18/2431
- G06N7/01
- G06V10/751
- IPC, 2
- G16B30 00
- C12Q1 6869
