Methods of analyzing nucleic acids from individual cells or cell populations.
Abstract
Methods, compositions and systems for analyzing individual cells or cell populations through split analysis of the contents of individual cells or cell populations. Individual cells or cell populations are codivided with processing reagents to access cell contents, and to especially identify the contents of a certain cell or cell population and subsequently analyze the cell contents and characterize it as derived from a individual cell or population of cells, including analysis and characterization of cell nucleic acids through sequencing.

Term
8.8 yearsleft in the term
Expires 26 June 2035.
- Priority
- Filed
- Granted
- Today
- Expires
88 claims: 4 independent, 84 dependent
- 1REIVINDICACIONES 1. Un método para analizar ácidos nucleicos de células que comprende:(a) proporcionar ácidos nucleicos derivados de una célula individual en una partición discreta;(b) generar una o más secuencias de ácido nucleico primarias derivadas de los ácidos nucleicos dentro de la partición discreta, las cuales una o más secuencias de ácido nucleico primarias tienen oligonucleótidos unidos a las mismas que comprenden una secuencia de código de barras de ácido nucleico común;;(c) generar una caracterización de la o las secuencias de ácido nucleico primarias o la o las secuencias de ácido nucleico secundarias derivadas de la o las secuencias de ácido nucleico primaras, las cuales una o más secuencias de ácido nucleico secundarias comprenden la secuencia de código de barras común;y (d) identificar la o las secuencias de ácido nucleico primarias o la o las secuencias de ácido nucleico secundarias como derivadas de la célula individual en base, al menos en parte, a una presencia de la secuencia de código de barras de ácido nucleico común en la caracterización generada en (c).
- 2El método de la reivindicación 1 en donde la partición 151 discreta es una gotita discreta.
- 3El método de la reivindicación 1 en donde, en (a), los oligonucleótidos se codividen con los ácidos nucleicos derivados de la célula individual en la partición discreta.
- 4El método de la reivindicación 3 en donde, en (a), al menos 10.000 de los oligonucleótidos se codividen con los ácidos nucleicos derivados de la célula individual en la partición discreta.
- 5El método de la reivindicación 4 en donde, en (a), al menos 100.000 de los oligonucleótidos se codividen con los ácidos nucleicos derivados de la célula individual en la partición discreta.
- 6El método de la reivindicación 5 en donde, en (a), al menos 500.000 de los oligonucleótidos se codividen con los ácidos nucleicos derivados de la célula individual en la partición discreta.
- 7El método de la reivindicación 1 en donde, en (a), los oligonucleótidos son proporcionados unidos a una perla, en donde cada oligonucleótido en una perla comprende la misma secuencia de código de barras, y la perla está codividida con la célula individual en la partición discreta.
- 8El método de la reivindicación 7 en donde los oligonucleótidos se unen de forma desprendióle a la perla. 152
- 9El método de la reivindicación 8 en donde la perla comprende una perla degradable.
- 10El método de la reivindicación 9 que además comprende, antes o durante (b), liberar los oligonucleótidos de la perla mediante degradación de la perla.
- 11El método de la reivindicación 1 que además comprende, antes de (c) , liberar la o las secuencias de ácido nucleico primarias de la partición discreta.
- 12El método de la reivindicación 1 en donde (c) comprende secuenciar la o las secuencias de ácido nucleico primarias o la o las secuencias de ácido nucleico secundarias.
- 13El método de la reivindicación 12 que además comprende ensamblar una secuencia de ácido nucleico contigua para al menos una porción de un genoma de la célula individual de secuencias de la o las secuencias de ácido nucleico primaras o la o las secuencias de ácido nucleico secundarias.
- 14El método de la reivindicación 13 en donde la célula individual se caracteriza en función de la secuencia de ácido nucleico para al menos una parte del genoma de la célula individual.
- 15El método de la reivindicación 1 en donde los ácidos nucleicos se liberan de la célula individual en la partición discreta. 153
- 16El método de la reivindicación 1 en donde los ácidos nucleicos comprenden ácido ribonucleico (ARN).
- 17El método de la reivindicación 16 donde el ARN es ARN mensajero (ARNm). :
- 18El método de la reivindicación 16 en donde (b) además comprende someter los ácidos nucleicos a transcripción inversa en condiciones que producen la o las secuencias :de ácido nucleico primarias.
- 19El método de la reivindicación 18 en donde la transcripción inversa se produce en la partición discreta.
- 20El método de la reivindicación 18 en donde los oligonucleótidos se proporcionan en la partición discreta y además comprenden una secuencia poli-T.
- 21El método de la reivindicación 20 en donde la transcripción inversa comprende hibridar la secuencia poli-T con al menos una parte de cada uno de los ácidos nucleicos y extender la secuencia poli-T de forma dirigida por una plantilla.
- 22El método de la reivindicación 21 en donde los oligonucleótidos además comprenden una secuencia de anclaje que facilita la hibridación de la secuencia poli-T.
- 23El método de la reivindicación 20 en donde los oligonucleótidos además comprenden una secuencia cebadora 154 aleatoria.
- 24El método de la reivindicación 23 en donde la secuencia cebadora aleatoria es un hexámero aleatorio.
- 25El método de la reivindicación 24 en donde . la transcripción inversa comprende hibridar la secuencia cebadora aleatoria con al menos una parte de cada uno de los ácidos nucleicos y extender la secuencia cebadora aleatoria de forma dirigida por una plantilla. .
- 26El método de la reivindicación 1 en donde una secuencia determinada de la o las secuencias de ácido nucleico primarias tiene complementariedad de secuencia con al menos una parte de uno de los ácidos nucleicos. ·
- 27EL método de la reivindicación 1 en donde la partición discreta como máximo incluye la célula individual entre múltiples células.
- 28El método de la reivindicación 1 en donde los oligonucleótidos además comprenden un segmento de secuencia molecular único.
- 29El método de la reivindicación 28 que además comprende identificar una secuencia de ácido nucleico individual de la o las secuencias de ácido nucleico primarias o de la o las secuencias de ácido nucleico secundarias derivadas de ’ un ácido nucleico determinado de los ácidos nucleicos basándose, 155 al menos en parte, en una presencia del segmento de secuencia molecular exclusivo.
- 30El método de la reivindicación 29 que además comprende determinar una cantidad del ácido nucleico determinado basándose en una presencia del segmento de secuencia molecular exclusivo. ;
- 31El método de la reivindicación 1 que además comprende, antes de (c), añadir una o más secuencias adicionales a la o las secuencias de ácido nucleico primarias para generar la o las secuencias de ácido nucleico secundarias.
- 32El método de la reivindicación 31 que además comprende añadir una primera secuencia de ácido nucleico adicional a la o las secuencias de ácido nucleico primarias con la ayuda de un oligonucleótido de conmutación.
- 33El método de la reivindicación 32 en donde el oligonucleótido de conmutación se híbrida con al menos una parte de la o las secuencias de ácido nucleico primarias y se extiende de forma dirigida por una plantilla para acoplar, la primera secuencia de ácido nucleico adicional con la o las secuencias de ácido nucleico primarias.
- 34El método de la reivindicación 33 que además comprende ampliar la o las secuencias de ácido nucleico primarias acopladas a la primera secuencia de ácido nucleico adicional. 156
- 35El método de la reivindicación 34 en donde la amplificación se produce en la partición discreta.
- 36El método de la reivindicación 34 en donde la ampliación se produce después de liberar la o las secuencias de ácido nucleico primarias acopladas a la primera secuencia de ácido nucleico adicional de la partición discreta.
- 37El método de la reivindicación 34 que además comprende, antes de amplificar, añadir una o más secuencias de ácido nucleico adicionales secundarias a la o las secuencias de ácido nucleico primarias acopladas a la primera secuencia adicional para generar la o las secuencias de ácido nucleico secundarias.
- 38El método de la reivindicación 37 en donde el añadido de una o más secuencias adicionales secundarias comprende retirar una parte de cada una de la o las secuencias de ácido nucleico primarias acopladas a la primera secuencia de ácido nucleico adicional y acoplar allí la o las secuencias de ácido nucleico adicionales secundarias.
- 39El método de la reivindicación 38 en donde la retirada se realiza mediante el cizallamiento de la o las secuencias de ácido nucleico primarias acopladas a la primera secuencia de ácido nucleico adicional.
- 40El método de la reivindicación 39 en donde el 157 acoplamiento se completa mediante ligación.
- 41El método de la reivindicación 18 que además comprende, antes de (c) , someter la o las secuencias de ácido nucleico primarias a transcripción para generar uno o más fragmentos de ARN. _
- 42El método de la reivindicación 41 en donde la transcripción se produce después de liberar la o las secuencias de ácido nucleico primarias de la partición discreta.
- 43El método de la reivindicación 41 en donde los oligonucleótidos además comprenden una secuencia promotora T7. :
- 44El método de la reivindicación 43 que además comprende, antes de (c) , retirar una parte de cada una de la o las secuencias de ARN y acoplar una secuencia adicional a la o las secuencias de ARN.
- 45El método de la reivindicación 44 que además comprende, antes de (c) , someter la o las secuencias de ARN acopladas a la secuencia adicional a transcripción inversa para generar la o las secuencias de ácido nucleico secundarias.
- 46El método de la reivindicación 45 que además comprende, antes de (c), amplificar la o las secuencias de ácido nucleico secundarias. 158
- 47El método de la reivindicación 41 que además comprende, antes de (c) , someter la o las secuencias de ARN a transcripción inversa para generar una o más secuencias - de ADN.
- 48El método de la reivindicación 47 que además comprende, antes de (c) , retirar una parte de cada una de la o las secuencias de ADN y acoplar una o más secuencias adicionales a la o las secuencias de ADN para generar la o las secuencias de ácido nucleico secundarias.
- 49El método de la reivindicación 48 que además comprende, antes de (c), amplificar la o las secuencias de ácido nucleico secundarias.
- 50El método de la reivindicación 1 en donde los ácidos nucleicos comprenden complementario (ADNc) generado a partir de transcripción inversa de ARN de la célula individual.
- 51El método de la reivindicación 50 en donde los oligonucleótidos además comprenden una secuencia cebadora y se proporcionan en la partición discreta.
- 52El método de la reivindicación 51 en donde la secuencia cebadora comprende un N-mer aleatorio.
- 53El método de la reivindicación 51 en donde (b) comprende hibridar la secuencia cebadora con el ADNc y extender la secuencia cebadora de una forma dirigida por una plantilla. 159
- 54El método de la reivindicación 1 en donde la partición discreta comprende oligonucleótidos de conmutación que comprenden una secuencia complementaria de los oligonucleótidos.
- 55El método de la reivindicación 54 en donde (b) comprende hibridar los oligonucleótidos de conmutación con al menos una parte de los fragmentos de ácido nucleico derivados de los ácidos nucleicos y extender los oligonucleótidos de conmutación de forma dirigida por una plantilla.
- 56El método de la reivindicación 1 en donde (b) comprende unir los oligonucleótidos a la o las secuencias de ácido nucleico primarias.
- 57El método de la reivindicación 1 en donde la o las secuencias de ácido nucleico primarias son fragmentos de ácido nucleico derivados de los ácidos nucleicos.
- 58El método de la reivindicación 1 en donde (b) comprende acoplar los oligonucleótidos a los ácidos nucleicos.
- 59El método de la reivindicación 58 en donde el acoplamiento comprende ligación.
- 60El método de la reivindicación 1 en donde múltiples particiones comprenden la partición discreta.
- 61El método de la reivindicación 60 en donde, en promedio, las múltiples particiones comprenden menos de una célula por 160 partición.
- 62El método de la reivindicación 60 en donde menos del 25 % de las particiones de las múltiples particiones no comprenden una célula.
- 63El método de la reivindicación 60 en donde las múltiples particiones comprenden particiones discretas que tienen, cada una, al menos una célula dividida.
- 64El método de la reivindicación 63 en donde menos del 25 % de las particiones discretas comprenden más de .una célula.
- 65El método de la reivindicación 64 en donde al menos un subconjunto de las particiones discretas comprende una perla.
- 66El método de la reivindicación 65 en donde al menos, el 75 % de las particiones discretas comprenden al menos una célula y al menos una perla.
- 67El método de la reivindicación 63 en donde las particiones discretas además comprenden secuencias de código de barras de ácido nucleico divididas.
- 68El método de la reivindicación 67 en donde las secuencias discretas comprenden al menos 1.000 secuencias de código de barras de ácido nucleico divididas diferentes. .
- 69El método de la reivindicación 68 en donde las secuencias discretas comprenden al menos 10.000 secuencias, de 161 código de barras de ácido nucleico divididas diferentes.
- 70El método de la reivindicación 69 en donde las secuencias discretas comprenden al menos 100.000 secuencias de código de barras de ácido nucleico divididas diferentes.
- 71El método de la reivindicación 60 en donde las múltiples particiones comprenden al menos 1.000 particiones. ’
- 72El método de la reivindicación 71 en donde las múltiples particiones comprenden al menos 10.000 particiones.
- 73El método de la reivindicación 72 en donde las múltiples particiones comprenden al menos 100.000 particiones.
- 74Un método para caracterizar células en una población de múltiples tipos de células diferentes que comprende:(a) proporcionar ácidos nucleicos de células individuales en la población a particiones discretas;(b) unir los oligonucleótidos que comprenden una secuencia de código de barras de ácido nucleico común con uno o más fragmentos de los ácidos nucleicos de las células individuales dentro de las particiones discretas, en donde múltiples particiones diferentes comprenden secuencias de código de barras de ácido nucleico comunes diferentes;(c) caracterizar el o los fragmentos de los ácidos nucleicos de las múltiples particiones discretas y atribuir el o los fragmentos a las células individuales basándose, al menos en 162 parte, en la presencia de una secuencia de código de barras común;y (d) caracterizar múltiples células individuales en la población basándose en la caracterización del o de los 5 fragmentos en las múltiples particiones discretas. :
- 75El método de la reivindicación 74 que además comprende fragmentar los ácidos nucleicos.
- 76El método de la reivindicación 74 en donde las particiones discretas son gotitas. 10
- 77El método de la reivindicación 74 en donde la caracterización del o de los fragmentos de los ácidos nucleicos comprende secuenciar el ácido desoxirribonucleico ribosómico de las células individuales y la caracterización de las células comprende identificar un género, especie, cepa 15 o variante celular.
- 78El método de la reivindicación 77 en donde las células individuales derivan de una muestra de microbioma.
- 79El método de la reivindicación 74 en donde las células individuales derivan de una muestra de tejido humano. 20
- 80El método de la reivindicación 74 en donde las células individuales derivan de células circulantes en un mamífero.
- 81El método de la reivindicación 74 en donde las células individuales derivan de una muestra forense. 163
- 82El método de la reivindicación 74 en donde los ácidos nucleicos se liberan de las células individuales en :las particiones discretas.
- 83Un método para caracterizar una célula individual o población de células que comprende:(a) incubar una célula con múltiples tipos de grupos de unión de características de superficie celular diferentes, : en donde cada tipo de grupo de unión de superficie celular diferente es capaz de unirse a una característica de superficie celular diferente y en donde cada tipo de grupo: de unión de superficie celular diferente comprende : un oligonucleótido informante asociado con este, en condiciones que permiten la unión entre uno o más grupos de unión de características de superficie celular y su característica de superficie celular respectiva, si está presente;(b) dividir la célula en una partición que comprende múltiples oligonucleótidos que comprenden una secuencia de código de barras;(c) unir la secuencia de código de barras con los grupos informantes de oligonucleótidos presentes en la partición;(d) secuenciar los grupos informantes de oligonucleótidos y los códigos de barras unidos;y : (e) caracterizar las características de superficie celular 164 presentes en la célula basándose en los oligonucleótidos informantes que se secuencian.
- 84Una composición que comprende múltiples particiones, en donde cada una de las múltiples particiones comprende una (i) célula individual y una (ii) población de oligonucleótidos que comprende una secuencia de código de barras de ácido nucleico común. .
- 85La composición de la reivindicación 84 en donde las múltiples particiones comprenden gotitas en una emulsión. '
- 86La composición de la reivindicación 84 en donde la población de oligonucleótidos dentro de cada una de las múltiples particiones está acoplada con una perla dispuesta dentro de cada una de las múltiples particiones.
- 87La composición de la reivindicación 84 en donde la célula individual tiene asociada a ella múltiples grupos de unión de características de superficie celular diferentes asociados con sus características de superficie celular respectivas, donde cada tipo diferente de grupo de unión de característica de superficie celular comprende un grupo informante de oligonucleótidos que comprende una secuencia de nucleótidos diferente.
- 88La composición de la reivindicación 87 en donde los múltiples grupos de unión de características de superficie 165 celular diferentes comprenden múltiples anticuerpos diferentes o fragmentos de anticuerpo que tienen una afinidad de unión hacia múltiples características de superficie celular diferentes. 166
Independent claims88
308 paragraphs in 7 sections, as filed
(54) Title: METHODS FOR ANALYZING NUCLEIC ACIDS FROM INDIVIDUAL CELLS OR POPULATIONS OF CELLS.
(54) Title: METHODS OF ANALYZING NUCLEIC ACIDS FROM INDIVIDUAL CELLS OR CELL POPULATIONS.
(57) Summary
Methods, compositions and systems for analyzing individual cells or cell populations through split analysis of the contents of individual cells or cell populations. Individual cells or cell populations are codivided with processing reagents to access cell contents, and to especially identify the contents of a certain cell or cell population and subsequently analyze the cell contents and characterize it as derived from a individual cell or population of cells, including analysis and characterization of cell nucleic acids through sequencing.
(57) Abstract
Methods, compositions and systems for analyzing individual cells or cell populations through the partitioned analysis of contents of individual cells or cell populations. Individual cells or cell populations are co-partitioned with processing reagents for accessing cellular contents, and for uniquely identifying the contents of a given cell or cell population, and subsequently analyzing the cell's contents and characterizing it as having derived from an individual cell or cell population , including analysis and characterization of the cell's nucleic acids through sequencing.
METHODS FOR ANALYZING NUCLEIC ACIDS IN CELLS
INDIVIDUALS OR CELL POPULATIONS
CROSS REFERENCE TO RELATED REQUESTS
This application claims priority from US Provisional Patent Application No. 62 / 017,558, filed June 26, 2014, and US Provisional Patent Application No. 62 / 061,567, filed October 8, 2014 , each of which is incorporated herein in its entirety by reference for all purposes. :
BACKGROUND.
Significant advances in the analysis and characterization of biological and biochemical materials and systems have led to unprecedented advances in understanding the mechanisms of life, health, disease, and treatment. Among these advances, technologies that target and characterize the genomic makeup of biological systems have produced some of the most innovative results, including advancements in the use and exploitation of gene amplification technologies, and nucleic acid sequencing technologies.
Nucleic acid sequencing can be used to obtain information in a wide variety of biomedical contexts, including diagnostic, prognostic, biotechnology, and forensic biology. Sequencing may involve basic methods including MaxamGilbert and chain termination sequencing methods, or de novo sequencing methods including random sequencing and bridging PCR, or next generation methods including polony sequencing, 454 pyrosequencing, sequencing from Illumina, SOLiD sequencing, Ion Torrent semiconductor sequencing, HeliScope single molecule sequencing, SMRT® sequencing, and more. '
Despite these advances in biological characterization, many challenges have not yet been addressed, or have been relatively poorly addressed through the solutions that are offered today. The present description provides novel solutions and approaches to address many of the shortcomings of existing technologies.
COMPENDIUM
Provided herein are methods, compositions, and systems for analyzing individual cells or small cell populations, including the analysis and attribution of nucleic acids from and to these individual cells and cell populations.
One aspect of the disclosure provides a method of analyzing cell nucleic acids that includes providing nucleic acids derived from an individual cell in a discrete partition; generating one or more primary nucleic acid sequences derived from the nucleic acids within the discrete partition, which one or more primary nucleic acid sequences have oligonucleotides attached thereto that comprise a common nucleic acid barcode sequence; generate a characterization of the primary nucleic acid sequence (s) or the secondary nucleic acid sequence (s) derived from the primary nucleic acid sequence (s), which one or more secondary nucleic acid sequences comprise the barcode sequence common; and identifying the primary nucleic acid sequence (s) or the secondary nucleic acid sequence (s) as derived from the individual cell based, at least in part, on a presence of the common nucleic acid barcode sequence in the characterization. generated.
In some embodiments, the discrete partition is a discrete droplet. In some embodiments, the oligonucleotides are codivided with the nucleic acids derived from the individual cell in the discrete partition. In some embodiments, at least 10,000, at least 100,000, or at least 500,000 of the oligonucleotides are codivided with the nucleic acids derived from the single cell in the discrete partition.
In some embodiments, the oligonucleotides are provided attached to a bead, where each oligonucleotide in a bead comprises the same barcode sequence, and the bead is codivided with the individual cell in the discrete partition. In some embodiments, the oligonucleotides are loosely attached to the bead. In some embodiments, the pearl comprises a degraded pearl. In some embodiments, before or during the generation of the primary nucleic acid sequence (s), the method includes leading the bead oligonucleotides by degradation of the bead. In some embodiments, prior to generating the characterization, the method includes leading the primary nucleic acid sequence (s) from the discrete partition.
In some embodiments, generating the characterization comprises sequencing the primary nucleic acid sequence (s) or the secondary nucleic acid sequence (s). The method may also include assembling a contiguous nucleic acid sequence for at least a portion of an individual cell genome from sequences of the primary nucleic acid sequence (s) or the secondary nucleic acid sequence (s). Furthermore, the method may also include characterizing the individual cell based on the nucleic acid sequence for at least a part of the genome of the individual cell.
In some embodiments, nucleic acids are released from the individual cell at the discrete partition. In some embodiments, the nucleic acids comprise ribonucleic acid (RNA), such as, for example, messenger RNA (mRNA). In some embodiments, generating one or more primary nucleic acid sequences includes subjecting the nucleic acids to reverse transcription under conditions that produce the or, the primary nucleic acid sequences. In some embodiments, reverse transcription occurs in the discrete partition. In some embodiments, the oligonucleotides are provided in the discrete partition and include a poly-T sequence. In some embodiments, reverse transcription comprises hybridizing the poly-T sequence to at least a portion of each of the nucleic acids and extending the poly-T sequence in a template-directed fashion. In some embodiments, the oligonucleotides include an anchor sequence that facilitates hybridization of the poly-T sequence. In some embodiments, the oligonucleotides include a random primer sequence which can be, for example, a random hexamer. In some embodiments, reverse transcription comprises hybridizing the random primer sequence to at least a portion of each of the nucleic acids and extending the random primer sequence in a template-directed manner.
In some embodiments, a particular sequence of the primary nucleic acid sequence (s) has sequence complementarity with at least a portion of one of the nucleic acids. In some embodiments, the discrete partition at most includes the single cell among multiple cells. In some embodiments, the oligonucleotides include a unique molecular sequence segment. In some embodiments, the method may include identifying an individual nucleic acid sequence from the primary nucleic acid sequence (s) or from the secondary nucleic acid sequence (s) derived from a particular nucleic acid of the nucleic acids on the basis, at least in part , in a presence of the unique molecular sequence segment. In some embodiments, the method includes determining an amount of the determined nucleic acid based on a presence of the unique molecular sequence segment.
In some embodiments, the method includes, prior to generating the characterization, adding one or more additional sequence (s) to the primary nucleic acid sequence (s) to generate the secondary nucleic acid sequence (s). In some embodiments, the method includes adding an additional first nucleic acid sequence to the primary nucleic acid sequence (s) with the aid of a switch oligonucleotide. In some embodiments, the switch oligonucleotide hybridizes to at least a portion of the primary nucleic acid sequence (s) and is directed by a template to couple the additional first nucleic acid sequence with the nucleic acid sequence (s). primary. In some embodiments, the method includes extending the primary nucleic acid sequence (s) coupled to the first additional nucleic acid sequence. In some embodiments, the amplification occurs in the discrete partition. In some embodiments, enlargement occurs after releasing the primary nucleic acid sequence (s) coupled to the first additional nucleic acid sequence from the discrete partition.
In some embodiments, after extension, the method includes adding one or more additional secondary nucleic acid sequence (s) to the primary nucleic acid sequence (s) coupled to the first additional sequence to generate the secondary nucleic acid sequence (s). In some embodiments, adding one or more additional secondary sequence (s) includes removing a portion of each of the primary nucleic acid sequence (s) coupled to the first additional nucleic acid sequence and coupling the additional secondary nucleic acid sequence (s) there. . In some embodiments, removal is accomplished by shearing the coupled primary nucleic acid sequence (s) (eg. , linked) to the first additional nucleic acid sequence. In some embodiments, prior to generating the characterization, the method includes subjecting the primary nucleic acid sequence (s) to transcription to generate one or more RNA fragments. In some embodiments, transcription occurs after releasing the primary nucleic acid sequence (s) from the discrete partition. In some embodiments, the oligonucleotides include a T7 promoter sequence. In some embodiments, prior to generating the characterization, the method includes removing a portion of each of the RNA sequence (s) and coupling an additional sequence to the RNA sequence (s). In some embodiments, prior to generating the characterization, the method includes subjecting the RNA sequence (s) coupled to the additional sequence to reverse transcription to generate the secondary nucleic acid sequence (s). In some embodiments, prior to generating the characterization, the method includes amplifying the secondary nucleic acid sequence (s). In some embodiments, before generating the characterization, the method includes submitting the o. reverse transcribed RNA sequences to generate one or more DNA sequences. In some embodiments, prior to generating the characterization, the method includes removing a portion of each of the DNA sequence (s) and coupling one or more additional sequence (s) to the DNA sequence (s) to generate the secondary nucleic acid sequence (s). . In some embodiments, prior to generating the characterization, the method includes amplifying the o: secondary nucleic acid sequences.
In some embodiments, the nucleic acids include complementary (cDNA) generated from reverse transcription of individual cell RNA. In some embodiments, the oligonucleotides include a primer sequence and are provided in the discrete partition. In some embodiments, the primer sequence includes a random N-mer. In some embodiments, generating the primary nucleic acid sequence (s) includes hybridizing the primer sequence to the cDNA and extending the primer sequence in a template-directed manner.
In some embodiments, the discrete partition includes switch oligonucleotides that comprise a sequence complementary to the oligonucleotides. In some embodiments, generating the primary nucleic acid sequence (s) includes hybridizing the switch oligonucleotides to at least a portion of the nucleic acid-derived nucleic acid fragments and extending the switch oligonucleotides in a template-directed fashion. In some embodiments, generating the primary nucleic acid sequence (s) includes linking the oligonucleotides to the primary nucleic acid sequence (s). In some embodiments, the primary nucleic acid sequence (s) are nucleic acid fragments derived from the nucleic acids. In some embodiments, generating the primary nucleic acid sequence (s) includes coupling (eg, ligating) the oligonucleotides to the nucleic acids.
In some embodiments, multiple partitions comprise the discrete partition. In some embodiments, multiple partitions, on average, comprise less than one cell per partition. In some embodiments, less than 25% of the partitions of the multiple partitions do not comprise a cell. In some embodiments, multiple partitions comprise discrete partitions each having at least one divided cell. In some embodiments, less: 25%, less than 20%, less than 15%, less than 10%, less than 5%, or less than 1% of the discrete partitions comprise more than one cell. In some embodiments, at least a subset of the discrete partitions comprises a bead. In some embodiments, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% of the discrete partitions comprise at least one cell and at least minus one pearl. In some embodiments,: the discrete partitions include divided nucleic acid barcode sequences. In some embodiments, the discrete partitions include at least 1,000, at least 10,000, or at least 100,000 different divided nucleic acid barcode sequences. In some embodiments, the multiple partitions comprise at least 1,000, at least 10,000, or at least 100,000 partitions.
In another aspect, the disclosure provides a method of characterizing cells in a population of multiple different cell types that includes providing nucleic acids from individual cells in the population to discrete partitions; binding the oligonucleotides they comprise. a common nucleic acid barcode sequence with one or more individual cell nucleic acid fragments within the discrete partitions, wherein multiple different partitions comprise different common nucleic acid barcode sequences; and characterizing the nucleic acid fragment (s) from the multiple discrete partitions and attributing the fragment (s) to individual cells based, at least in part, on the presence of a common barcode sequence; and characterizing multiple individual cells in the population based on the characterization of the fragment (s) in the multiple discrete partitions.
In some embodiments, the method includes fragmenting the nucleic acids. In some embodiments, the discrete partitions are droplets. In some embodiments, characterization of the nucleic acid fragment (s) includes sequencing the ribosomal deoxyribonucleic acid of individual cells and characterization of cells comprises identifying a genus, species, strain, or cell variant. In some embodiments, individual cells are derived from a sample of the microbiome. In some embodiments, individual cells are derived from a sample of human tissue. In some embodiments, the individual cells are derived from cells circulating in a mammal. In some embodiments, individual cells are derived from a forensic sample. In some embodiments, nucleic acids are released from individual cells at discrete partitions.
A further aspect of the disclosure provides a method for characterizing an individual cell or a population of cells that includes incubating a cell with multiple types of different cell surface characteristics binding groups, wherein each type of different cell surface binding group is able to join <sub>:</sub> a different cell surface feature and wherein each type of different cell surface binding group comprises a reporter oligonucleotide associated therewith, under conditions that allow binding between one or more cell surface feature binding groups and their surface feature respective cell phone, if present; dividing the cell into a partition comprising multiple oligonucleotides comprising a barcode sequence; joining the barcode sequence with the oligonucleotide reporter groups present in the partition; sequencing the oligonucleotide reporter pools and attached barcodes; and characterizing the cell surface characteristics present in the cell based on the reporter oligonucleotides that are sequenced.
Another aspect of the disclosure provides a composition comprising multiple partitions, wherein each of the multiple partitions comprises an individual cell and a population of oligonucleotides comprising a common nucleic acid barcode sequence. In some embodiments, multiple partitions comprise droplets in an emulsion. In some embodiments, the population of oligonucleotides within each of the multiple partitions is coupled with a bead disposed within each of the multiple partitions. In some embodiments, the individual cell has associated with it multiple different cell surface feature binding groups associated with its respective cell surface features and each different type of cell surface feature binding group includes a reporter group of oligonucleotides comprising a different nucleotide sequence. In some embodiments, the multiple feature union groups<sup>:</sup> Different cell surface characteristics include multiple different antibodies or antibody fragments that have a binding affinity for multiple different cell surface characteristics.
Certain additional aspects and advantages of the present disclosure will be apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be understood, the present description is capable of other different embodiments, and its various details are capable of modification in various obvious respects, all without departing from the description. Accordingly, the figures and description should be understood to be illustrative and not restrictive in nature.
INCORPORATION BY REFERENCE
All publications, patents and patent applications mentioned in the present specification are incorporated herein by this reference to the same extent as if it were indicated that each individual publication, patent or patent application is specifically and individually incorporated by reference . To the extent that publications and patents or patent applications incorporated by reference contradict the description present in the specification, it is intended that the specification prevail over any of these inconsistent materials.
BRIEF DESCRIPTION OF THE DRAWINGS
The novel features of the invention are set forth in detail in the appended claims. A better understanding of the features and advantages of the present invention may be obtained with reference to the following detailed description which sets out illustrative embodiments, in which the principles of the invention are used, and the accompanying drawings (also contained herein) at which:<sub>:</sub>
Figure 1 schematically illustrates a microfluidic channel structure for dividing individual or small cell groups.
Figure 2 schematically illustrates a microfluidic channel structure for codivating cells and beads or microcapsules comprising additional reagents.
Figure 3 schematically illustrates an exemplary process for the amplification and barcode creation of nucleic acids from cells.
Figure 4 provides a schematic illustration of the use of cell nucleic acid barcoding to attribute sequence data to individual cells or groups of cells for use in characterization. .
Figure 5 provides a schematic illustrating cells associated with labeled cell-binding ligands.
Figure 6 provides a schematic illustration of an exemplary workflow for performing RNA analysis using the methods described herein.
Figure 7 provides a schematic illustration of. an exemplary barcode oligonucleotide framework for use in ribonucleic (RNA) analysis using the methods described herein.
Figure 8 provides an image of codivided individual cells together with beads having individual barcodes
Figure 9A-9E provides a schematic illustration of exemplary barcode oligonucleotide structures for use in RNA analysis and exemplary operations to perform RNA analysis.
Figure 10 provides a schematic illustration of an exemplary barcode oligonucleotide structure for use in exemplary RNA analysis and use of a sequence for in vitro transcription.
Figure 11 provides a schematic illustration of an exemplary barcode oligonucleotide structure for use in RNA analysis and exemplary operations to perform RNA analysis.
Figure 12A-12B provides a schematic illustration of an exemplary barcode oligonucleotide structure for use in RNA analysis.
Figure 13A-13C provides illustrations of exemplary performances of template swap reverse transcription and PCR in the partitions.
Figure 14A-14B provides illustrations of exemplary performances of reverse transcription and cDNA amplification in partitions with various numbers of cells.
Figure 15 provides an illustration of exemplary yields of cDNA synthesis and quantitative real-time PCR with various concentrations of input cells and also the effect of varying primer concentration on performance with a fixed input cell concentration.
Figure 16 provides an illustration of exemplary performances of in vitro transcription.
Figure 17 shows an exemplary computer control system that is programmed or otherwise configured to implement the methods provided herein.
DETAILED DESCRIPTION
Although various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It will be apparent to those skilled in the art that various variations, changes, and substitutions can be made without departing from the spirit of the invention. It will be understood that various alternatives to the embodiments of the invention described herein may be employed.
When values are described as ranges, such description shall be understood to include the description of all possible sub-ranges within those ranges, as well as specific numerical values that lie within such ranges regardless of whether a specific numerical value or a specific sub-range is used. expressly mentioned.
I. Single cell analysis
Advanced nucleic acid sequencing technologies have provided tremendous results in sequencing biological materials, including providing substantial sequence information on individual organisms and relatively pure biological samples. However, these systems have not been effective in successfully identifying and characterizing subpopulations of cells in biological samples that may represent a smaller minority of the overall conformation of the sample, but for which individualized sequence information could be even more valuable.
Most nucleic acid sequencing technologies derive the nucleic acids they sequence from collections of cells derived from tissue or other samples. Cells can be processed, typically, in bulk, to extract genetic material representing an average of the cell population, which can then be processed into sequencing ready-DNA libraries that are configured for a particular sequencing technology. As will be appreciated, although often stated in terms of DNA or nucleic acids, nucleic acids derived from cells can include DNA, or RNA, inclusive, e.g. ex. , MRNA, total RNA, or the like, which can be processed to produce cDNA for sequencing, e.g. eg, using any of a number of RNA sequencing methods. Following this processing, without a specific cellular marker, the attribution of genetic material as contributed by a subset of cells or all cells in a sample is practically impossible in such an ensemble approach. '
In addition to the inability to attribute characteristics to particular subsets of cell populations, such pooled sample preparation methods are also, from the outset, predisposed to primarily identify and characterize most of the constituents in the cell sample and are not designed to be able to distinguish minority constituents, p. eg, genetic material contributed by a cell, some cells, or a small percentage of total cells in the sample. Also, when expression levels are analyzed, p. For example, from mRNA, a pool approach would be predisposed to present possibly very inaccurate data from cell populations that are not homogeneous in terms of expression levels. In some cases, when expression is high in a small minority of cells in a population tested and is absent in most cells in the population, a set method indicates a low level expression for the entire population.
This original majority bias is further magnified, and even overloaded, by processing operations used to augment the sequencing libraries of these samples. In particular, most of the latest generation sequencing technologies rely on the geometric amplification of nucleic acid fragments, such as the polymerase chain reaction, to produce sufficient DNA for the sequencing library. However, such geometric amplification is predisposed to amplification of the major constituents in a sample and may not preserve the starting proportions of said minority and major constituents. By way of example, if a sample includes 95% DNA from a particular cell type in a sample, e.g. ex. , cells from host tissue, and 5% DNA from another cell type, e.g. ex. , cancer cells, PCR-based amplification can preferentially amplify majority DNA rather than minority DNA, both based on comparative exponential amplification (repeated doubling of the highest concentration rapidly exceeds that of the smallest fraction) and based on of the sequestration of the reagents and the amplification resources (as the larger fraction is amplified, it preferably uses primers and other amplification reagents).
Although some of these difficulties can be addressed using different sequencing systems, such as single molecule systems that do not require amplification, single molecule systems, as well as the ensemble sequencing methods of other state-of-the-art sequencing systems, may also have requirements for entry DNA requirements large enough. In particular, single molecule sequencing systems such as the Pacific Biosciences SMRT sequencing system may have sample input DNA requirements of between 500 nanograms (ng) to more than 10 micrograms (pg), which is much higher than expected. which can be derived from individual cells or even small subpopulations of cells. Likewise, other NGS systems can be optimized for starting amounts of sample DNA in the sample of between about 50 ng and about 1 pg.
II. Cell compartmentalization and characterization Described herein, however, are methods and systems for characterizing nucleic acids from small populations of cells and, in some cases, for characterizing nucleic acids from individual cells, especially in; the context of larger populations of cells. The methods and systems provide the advantages of being able to provide the attribution advantages of unamplified single molecule methods with the high throughput of the other next generation systems, with the additional advantages of being able to process and sequence extremely low amounts nucleic acids: input derived from individual cells or small collections of cells.
In particular, the methods described herein compartmentalize the analysis of single cells or small populations of cells, including, e.g. g., nucleic acids from individual cells or small groups of cells and then allow the analysis to be reattributed to the individual cell or small group of cells from which: the nucleic acids were derived. This can be achieved regardless of whether the cell population represents a 50/50 mix of cell types, a 90/10 mix of cell types, or virtually any ratio of cell types, as well as a complete heterogeneous mix of different cell types. cells or any mixture of these. Different types of cells can include biological cells or organisms from different types of tissues from an individual, from different individuals, from different genera, species, strains, variants, or any combination of the above. For example, different cell types can include normal and tumor tissue from an individual, multiple different bacterial species, strains, and / or variants: from environmental, forensic, microbiome, or other samples, or any of several different mixtures of cell types. cells.
In one aspect, the methods and systems described herein provide for the compartmentalization, deposition, or division of the nucleic acid content of individual cells of a cell-containing sample material, into discrete compartments or partitions (interchangeably referred to herein partitions), where each partition maintains the separation of its own content from the content of other partitions. Unique identifiers, eg. ex. Bar codes can be supplied before, after or simultaneously to the partitions containing the compartmentalized or divided cells, to allow later attribution of the characteristics of the individual cells to the particular compartment.
As used herein, in some respects, partitions refer to containers (such as wells, microwells, tubes, through ports in nanoarray substrates, eg, BioTrove nanoarrays or other containers). In many aspects, however, the compartments or partitions comprise partitions that can flow within fluid streams. These partitions can be composed, eg. ex. , by microcapsules or microvesicles that have an external barrier that surrounds an internal fluid center or core, or they can be a porous matrix that is capable of entraining and / or retaining materials within its matrix. However, in some aspects, these partitions comprise aqueous fluid droplets within a continuous non-aqueous phase, e.g. ex. , an oil phase. A number of various containers are described, for example, in US patent application; No. 13 / 966,150, filed August 13, 2013, the full description of which is incorporated herein by this reference in its entirety for all purposes. Also, some emulsion systems for creating stable droplets in continuous non-aqueous or oily phases are described in detail, e.g. eg, in US Patent Publication No. 2010/0105112, the full disclosure of which is incorporated herein in its entirety by this reference for all purposes.
In the case of droplets in an emulsion, the assignment of individual cells to discrete partitions can generally be achieved by introducing a flow stream of cells in an aqueous fluid into a flow stream of a non-aqueous fluid, so that the droplets are generated at the junction of the two streams. By providing the stream containing aqueous cells at a given cell concentration level, the occupancy level of the resulting partitions can be controlled in terms of cell numbers. In some cases, when single cell partitions are desired, it may be desirable to control the relative flow rates of the fluids so that, on average,<sup>:</sup> Partitions contain less than one cell per partition to ensure that partitions that are occupied are mostly uniquely occupied. Also, you may want to control the flow rate to provide that a higher percentage of partitions are occupied, e.g. ex. , leaving only a small percentage of partitions unoccupied. In some respects, channel and flow architectures are controlled to ensure a desired number of uniquely occupied partitions, less than a certain level of unoccupied partitions, and less than a certain level of multi-occupied partitions. .
In many cases, systems and methods are used to ensure that the vast majority of occupied partitions (partitions containing one or more microcapsules) include no more than 1 cell per occupied partition. In some cases, the division process is controlled so that less than 25% of the occupied partitions contain more than one cell and, in many cases, less than 20% of the occupied partitions have more than one cell, while in some In these cases, less than 10% or even less than 5% of occupied partitions include more than one cell per partition. Additionally or alternatively, in many cases, it is desirable to avoid creating an excessive number of empty partitions. Although this can be achieved by providing sufficient numbers of cells to the division zone, the Poisson distribution would hopefully increase the number of partitions that would include multiple cells. Thus, in accordance with the aspects described herein, the flow of one or more of the cells or other fluids directed to the division zone is controlled so that, in many cases, no more than 50% of the partitions generated they are unoccupied, that is, they include less than 1 cell, no more than% of the generated partitions, no more than 10% of the generated partitions, can be unoccupied. Furthermore, in some respects, these flows are controlled to present a non-Poissonian distribution of uniquely occupied partitions while providing lower levels of unoccupied partitions. By restatement, in some respects, the above-stated ranges of unoccupied partitions can be achieved by providing any of the unique occupancy rates described above. For example, in some cases, using the systems and methods described herein creates resulting partitions that have multiple occupancy rates of less than 25%, less than 20%, less than 15%, less than 10%, and, in some cases, less than 5%, while they have unoccupied partitions of between less than 50%, less than 40%, less than 30%, less than 20%, less than 10% and, in some cases, less than 5 %.
As will be appreciated, the above mentioned occupancy rates can also be applied to partitions that include cells and beads containing the barcode oligonucleotides. In particular, in some aspects, a substantial percentage of the overall occupied partitions will include a bead and a cell. In particular, it may be desirable to provide that at least 50% of the partitions are occupied by at least one cell and at least one bead or at least 75% of the partitions may be occupied thereby or even at least 80% or at least 90% of the partitions are occupied. Furthermore, in cases where it is desired to provide a single cell and a single bead within a partition, at least 50% of the partitions may be occupied thereby, at least 60%, at least%, at least 80% or even at least 90% of the partitions can be occupied that way.
Although described in terms of providing substantially uniquely occupied partitions above, in certain cases, it is desirable to provide multi-occupied partitions, e.g. eg, containing two, three, four or more cells and / or beads within a single partition. Consequently, as indicated above, the flow characteristics of the cell and / or bead containing fluids and partition fluids can be controlled to provide such multiply occupied partitions. In particular, flow parameters can be controlled to provide a desired occupancy rate of more than 50% of the partitions, more than 75%, and in some cases more than 80%, 90%, 95% or more. .
Furthermore, in many cases, the multiple beads within a single partition may comprise different reagents associated with them. In such cases, it may be advantageous to introduce different beads into a common channel or droplet generation junction, from different bead sources, that is, containing different associated reagents, through different channel inputs into said common channel or junction. droplet generation. In these cases, the flow and frequency of the 10 different beads in the channel or junction can be controlled to provide the desired proportion of microcapsules from each source, ensuring the desired pairing or combination of said beads in a partition with the amount desired cell.
The partitions described herein are often characterized by extremely low volumes, e.g. ex. , less than 10 pL, less than 5 pL, less than 1 pL, less than 900 picoliters (pL), less than 800 pL, less than 700 pL, less than 600 pL, less than 500 pL, less than 400 pL, less than 300 pL, 20 less than 200 pL, less than 100 pL, less than 50 pL, less than 20 pL, less than 10 pL, less than 1 pL, less than 500 nanoliters (nL) or even less than 100 nL, 50 nL or even less.
For example, in the case of droplet-based partitions, the droplets may have overall volumes less than 1000 pL, less than 900 pL, less than 800 pL, less than 700 pL, less than 600 pL, less than 500 pL, less than 400 pL, less than 300 pL, less than 200 pL, less than 100 pL, less than 50 pL, less than 20 pL, less than 10 pL, or even less than 1 pL. When codivided with beads, it will be appreciated that the fluid volume of sample, e.g. For example, which includes codivided cells, within the partitions it can be less than 90% of: the volumes described above, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, or even less than 10% of the volumes described above. :
As described elsewhere herein, the division of species can generate a population of partitions. In these cases, any suitable number of partitions can be generated to generate the partition population. For example, in a method described herein, a partition population can be generated comprising at least about 1,000 partitions, at least about 5,000 partitions, at least about 10 .. 000 partitions, at least apr about 50,000 partitions, at least about 100,000 partitions, at least about 500,000 partitions, at least about 1,000,000 partitions, at least about 5,000,000 partitions at least about 10,000,000 partitions, at least about 50,000,000 partitions, at least about 100,000,000 partitions, at least about 500,000,000 partitions, or at least about 1,000,000,000 partitions. Further, <sub>:</sub> the partition population can comprise unoccupied partitions (eg, empty partitions) and occupied partitions.
In certain cases, microfluidic channel networks are particularly suitable for generating partitions as described herein. Examples of such microfluidic devices include those described in detail in US Provisional Patent Application<sub>:</sub> No. 61 / 977,804, filed April 4, 2014, the full description of which is incorporated herein by this reference in its entirety for all purposes. Alternative mechanisms can also be employed in the division of individual cells, including porous membranes through which aqueous mixtures of cells are extruded into non-aqueous fluids. These systems are generally supplied eg. eg, by Nanomi, Inc.
An example of a simplified microfluidic channel structure for dividing individual cells is illustrated in Figure 1. As described elsewhere herein, in some cases most occupied partitions include no more than one cell per occupied partition. and in some cases some of the generated partitions are idle. In some cases, however, some of the occupied partitions may include more than one cell. In some cases, the division process can be controlled so that less than 25% of the occupied partitions contain more than one cell, and in many cases less than 20% of the occupied partitions have more than one cell, while in some cases, less than 10% or even less than 5% of the occupied partitions include more than one cell per partition. As shown, the channel structure may include channel segments 102, 104, 106, and 108 that communicate at a junction of channels 110. In operation, a first aqueous fluid 112 including suspended cells 114 may be transported along with channel segment 102 at junction 110, while a second fluid 116 that: is immiscible with aqueous fluid 112 is supplied to junction 110 of channel segments 104 and 106 to create discrete droplets 118 of the aqueous fluid including individual cells 114, which flow into channel segment 108.
In some aspects, this second fluid 116 comprises an oil, such as a fluorinated oil, that includes a fluorosurfactant to stabilize the resulting droplets, e.g. eg, by inhibiting subsequent coalescence of the resulting droplets. Examples of particularly useful fluorosurfactants and cleavage fluids are described, for example, in US Patent Publication No. 2010/0105112, which is incorporated herein in its entirety by this reference for all purposes.
In other aspects, in addition to or as an alternative to droplet-based division, cells can be encapsulated within a microcapsule comprising an outer shell or layer or a porous matrix wherein one or more individual cells or small groups of cells are entrained and may include other reagents. Cell encapsulation can be accomplished through various processes. In general, such processes combine an aqueous fluid containing the cells to be analyzed with a polymeric precursor material that may be capable of being formed into a gel or other solid or semi-solid matrix upon application of a particular stimulus to the polymeric precursor. These stimuli include, p. eg, thermal stimuli (either heating or cooling), photostimuli (eg, by light curing), chemical stimuli (eg. , by crosslinking, initiation of polymerization of the precursor (eg, by added initiators) or the like.
The preparation of microcapsules comprising cells can be carried out by various methods. For example, air knife droplets or aerosol generators can be used to dispense droplets of precursor fluids to gelling solutions to form microcapsules that include individual cells or small groups of cells. Also, membrane-based encapsulation systems, such as those provided, may be used, e.g. ex. , ' by
Nanomi, Inc., to generate microcapsules as described herein. In some aspects, microfluid systems such as those shown in Figure 1 can easily be used for encapsulation of cells as described herein. In particular, and with reference to Figure 1, the aqueous fluid comprising the cells and the polymeric precursor material is caused to flow towards the junction of the channels 110, where it is divided into droplets 118 comprising the individual cells 114, through nonaqueous fluid flow 116. In the case of encapsulation methods, the nonaqueous fluid 116 may also include an initiator to cause polymerization and / or crosslinking of the polymeric precursor to form the microcapsule that includes entrained cells. Examples of particularly useful polymeric initiator / precursor pairs include those described, e.g. For example, U.S. Patent Application No. 61 / 940,318, filed February 7, 2014, 61 / 991,018, filed May 9, 2014, and U.S. Patent Application No. 14 / 316,383, filed on June 26, 2014, which are incorporated herein in their entirety by this reference for all purposes.
For example, in the case where the polymeric precursor material comprises a linear polymeric material, e.g. For example, a linear polyacrylamide, PEG or other linear polymeric material, the activating agent may comprise a crosslinking agent or a chemical that activates a crosslinking agent within the droplets formed. Likewise, for polymeric precursors comprising polymerizable monomers, the activating agent may comprise a polymerization initiator. For example, in certain cases, when the polymeric precursor comprises a mixture of acrylamide monomer with a comonomer of Ν, Ν'-bis (acryloyl) cystamine (BAO), an agent such as tetraethylmethylenediamine (TEMED) can be provided within the second fluid stream in channel segments 104 and 106, which initiates copolymerization of acrylamide and BAC in a cross-linked polymer network or hydrogel.
Upon contact of the second fluid stream 116 with the first fluid stream 112 at junction 110 in droplet formation, the TEMED may diffuse from the second fluid 116 into the first aqueous fluid 112 comprising the linear polyacrylamide, which will activate crosslinking of the polyacrylamide within the droplets, resulting in gel formation, e.g. ex. , hydrogel, microcapsules 118, such as solid or semi-solid beads or particles that entrain the resulting cells 114. Although described in terms of polyacrylamide encapsulation, other activatable encapsulation compositions may also be employed in the context of the methods and compositions described in the present. For example, the formation of alginate droplets followed by exposure to divalent metal ions, e.g. eg, Ca2 +, can be used as an encapsulation process using the processes described. Likewise, agarose droplets can also be capsulated by gelation based on temperature, e.g. eg after cooling or similar. As will be appreciated, in some cases, the encapsulated cells may be selectively releasable from the microcapsule, e.g. eg, by the passage of time, or after the application of a particular stimulus, which degrades the microcapsule sufficiently to allow the cell or its contents to be released from the microcapsule, e.g. eg to an additional partition, such as a droplet. For example, in the case of the polyacrylamide polymer described above, degradation of the microcapsule can be achieved by introducing a suitable reducing agent, such as DTT or the like, to cleave the disulfide bonds that cross-link the polymer matrix. See, p. e.g., U.S. Provisional Patent Application No. 61 / 940,318, filed February 7, 2014, 61 / 991,018, filed May 9, 2014, and U.S. Patent Application No. 14 / 316,383, filed on June 26, 2014, which are incorporated herein in their entirety by this reference for all purposes.
As will be appreciated, encapsulated cells or cell populations provide certain potential storage advantages and being more portable than droplet-based divided cells. In addition, in some cases, it may be desirable to allow cells to be analyzed to incubate for a selected period to characterize changes in said cells over time, either with or without different stimuli. In these cases, 'encapsulation of individual cells may allow a longer incubation than simple division into emulsion droplets, although in some cases, cells divided into droplets can also be incubated for different periods, e.g. e.g. at least 10 seconds, at least 30 seconds, at least 1 minute, at least 5 minutes, at least 10 minutes, at least 30 minutes, at least 1 hour, at least 2 hours, at least 5 hours, or at least 10 hours or more. As indicated above, encapsulation of cells can constitute division of cells where other reagents are codivided. Alternatively, encapsulated cells: can be easily deposited in other partitions, e.g. eg, droplets, as described above.
According to certain aspects, cells can be divided together with lysis reagents to release: the contents of the cells within the partition. In these cases, the lysing agents can be contacted with the cell suspension simultaneously or immediately prior to the introduction of the cells into the droplet generation / division junction zone, e.g. eg, through an additional channel or channels upstream of the channel junction 110. Examples of lysis agents include bioactive reagents, such as lysis enzymes that are used for the lysis of different types of cells, e.g. g. gram positive or negative bacteria, plants, yeast, mammals, etc., such as lysozymes, achromopeptidase, lysostafin, labiase, kitalase, lithicase and various available lysis enzymes, e.g. g., from Sigma-Aldrich, Inc. (St Louis, MO), as well as other commercially available lysis enzymes. Other lysing agents can be co-divided further or alternatively. with the cells to cause the release of the contents of the cells in the partitions. For example, in some cases, surfactant-based lysis solutions can be used to lyse cells, although these may be less desirable for emulsion-based systems where surfactants can interfere with stable emulsions. In some cases, the lysis solutions may include non-ionic surfactants such as TritonX-100 and Tween 20. In some cases, the lysis solutions may include ionic surfactants such as sarcosyl and sodium dodecyl sulfate (SDS) . Similarly, lysis methods employing other methods such as electroporation, thermal, acoustic or mechanical cell disruption can also be used in certain cases, e.g. e.g., non-emulsion based division as cell encapsulation which may be in addition to or in place of droplet division, where any pore size of the encapsulation is small enough to retain nucleic acid fragments of a desired size, after cell disruption.
In addition to the cell-cleaved lysis agents described above, other reagents can also be co-cleaved with the cells, including, for example, DNase and RNase inactivating agents or inhibitors, such as proteinase K, chelating agents, such as EDTA, and other reagents. used to remove or otherwise reduce the negative activity or impact of the various components of Cellular Use on the subsequent processing of nucleic acids. Furthermore, in the case of encapsulated cells, the cells can be exposed to a suitable stimulus to release the cells or their contents from a codivided microcapsule. For example; In some cases, a chemical stimulus can be co-divided along with an encapsulated cell to allow degradation of the microcapsule and release of the cell or its contents to the larger partition. In some cases, this stimulus may be the same as the stimulus described elsewhere herein for the release of oligonucleotides from their respective bead or partition. In alternative aspects, this may be a different and non-overlapping stimulus to allow an encapsulated cell to be released to a partition at a different time from the release of the oligonucleotides to the same partition.
Other reagents can also be co-cleaved with cells, such as endonucleases to fragment DNA from cells, DNA polymerase enzymes, and dNTPs used to amplify nucleic acid fragments from cells and to bind barcode oligonucleotides to the amplified fragments. Other reagents may also include reverse transcriptase enzymes, including enzymes with terminal transferase activity, primers, and oligonucleotides and switch oligonucleotides (also referred to herein as switch oligos) that can be used for template exchange. In some cases, template swapping can be used to increase the length of a cDNA. In an example of template exchange, cDNA can be generated from the reverse transcription of a template, e.g. g., cellular mRNA, where a reverse transcriptase with terminal transferase activity can add additional nucleotides, e.g. g., polyC, to cDNAs that are not encoded by the template, such as at one end of the cDNA. Switch oligonucleotides can include sequences complementary to the additional nucleotides, e.g. eg, polyG. The additional nucleotides (p.
g., polyC) in the cDNA can hybridize to sequences complementary to additional nucleotides (eg, polyG) in the switch oligonucleotide, whereby the switch oligonucleotide can be used by reverse transcription as a template to further extend the cDNA. The switch oligonucleotides can comprise deoxyribonucleic acids, ribonucleic acids, modified nucleic acids including blocked nucleic acids (LNAs), or any combination.
In some cases, the length of a switch oligonucleotide can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12,
13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26,27,
28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41,42,
43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56,57,
58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71,72,
73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86,87,
88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101,
102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113,
114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, .125,
<td> 126,</td><td> 127,</td><td> 128,</td><td> 129,</td><td> 130,</td><td> 131,</td><td> 132,</td><td> 133,</td><td> 134,</td><td> 135,</td><td> 136,</td><td> 137,</td>
<td> 138,</td><td> 139,</td><td> 140,</td><td> 141,</td><td> 142,</td><td> 143,</td><td> 144,</td><td> 145,</td><td> 146,</td><td> 147,</td><td> 148,</td><td> 14 9,</td>
<td> 150,</td><td> 151,</td><td> 152,</td><td> 153,</td><td> 154,</td><td> 155,</td><td> 156,</td><td> 157,</td><td> 158,</td><td> 159,</td><td> 160,</td><td> 161,</td>
<td> 162,</td><td> 163,</td><td> 164,</td><td> 165,</td><td> 166,</td><td> 167,</td><td> 168,</td><td> 169,</td><td> 170,</td><td> 171,</td><td> 172,</td><td> 173,</td>
<td> 174,</td><td> 175,</td><td> 176,</td><td> 177,</td><td> 178,</td><td> 179,</td><td> 180,</td><td> 181,</td><td> 182,</td><td> 183,</td><td> 184,</td><td> 185,</td>
<td> 18'6,</td><td> 187,</td><td> 188,</td><td> 189,</td><td> 190,</td><td> 191,</td><td> 192,</td><td> 193,</td><td> 194,</td><td> 195,</td><td> 196,</td><td> 197 ,</td>
<td> 198,</td><td> 199,</td><td> 200,</td><td> 201,</td><td> 202,</td><td> 203,</td><td> 204,</td><td> 205,</td><td> 206,</td><td> 207,</td><td> 208,</td><td> 209,</td>
<td> 210,</td><td> 211,</td><td> 212,</td><td> 213,</td><td> 214,</td><td> 215,</td><td> 216,</td><td> 217,</td><td> 218,</td><td> 219,</td><td> 220,</td><td> 221,</td>
<td> 222,</td><td> 223,</td><td> 224,</td><td> 225,</td><td> 226,</td><td> 227,</td><td> 228,</td><td> 229,</td><td> 230,</td><td> 231,</td><td> 232,</td><td> 233,</td>
<td> 234,</td><td> 235,</td><td> 236,</td><td> 237,</td><td> 238,</td><td> 239,</td><td> 240,</td><td> 241,</td><td> 242,</td><td> 243,</td><td> 244,</td><td> 24 5,</td>
246, 247, 248, 249, 250 nucleotides or longer.
In some cases, the length of a switch oligonucleotide can be at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24,25,
26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39,40,
41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54,55,
56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69,70,
71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84,85,
<td> 86, 87, 88</td><td> , 89,</td><td> 90, 91, 92, 93</td><td> , 94,</td><td> 95, 96, 97, 98,</td><td> 99,</td><td> 100</td>
<td> 101, 102,</td><td> 103,</td><td> 104, 105, 106,</td><td> 107,</td><td> 108, 109, 110,</td><td> 111,</td><td> 112</td>
<td> 113, 114,</td><td> 115,</td><td> 116, 117, 118,</td><td> 119,</td><td> 120, 121, 122,</td><td> 123,</td><td> 124</td>
<td> 125, 126,</td><td> 127,</td><td> 128, 129, 130,</td><td> 131,</td><td> 132, 133, 134,</td><td> 135,</td><td> 136</td>
<td> 137, 138,</td><td> 139,</td><td> 140, 141, 142,</td><td> 143,</td><td> 144, 145, 146,</td><td> 147,</td><td> 148</td>
<td> 149, 150,</td><td> 151, 152,</td><td> 153, 154, 155,</td><td> 156,</td>
<td> 161, 162,</td><td> 163, 164,</td><td> 165, 166, 167,</td><td> 168,</td>
<td> 173, 174,</td><td> 175, 176,</td><td> 177, 178, 179,</td><td> 180,</td>
<td> 185, 186,</td><td> 187, 188,</td><td> 189, 190, 191,</td><td> 192,</td>
<td> 197 , 198,</td><td> 199, 200</td><td> , 201, 202, 203,</td><td> 204,</td>
<td> 209, 210,</td><td> 211, 212,</td><td> 213, 214, 215,</td><td> 216,</td>
<td> 221, 222,</td><td> 223, 224,</td><td> 225, 226, 227,</td><td> 228,</td>
<td> 233, 234,</td><td> 235, 236,</td><td> 237, 238, 239,</td><td> 240,</td>
<td> 245, 246, ;</td><td> 247, 248,</td><td colspan="2">249 or 250 nucleotides</td>
<td colspan="2">In some cases,</td><td>the length of</td><td>> a</td>
<td colspan="3">switching can be a maximum of</td><td><sub>;</sub> 2, :</td>
<td> 10, 11, 12</td><td> , 13, 14,</td><td> 15, 16, 17, 18,</td><td> 19,</td>
<td> 25, 26, 27</td><td> , 28, 29,</td><td> 30, 31, 32, 33,</td><td> 34,</td>
<td> 40, 41, 42</td><td> , 43, 44,</td><td> 45, 46, 47, 48,</td><td> 49,</td>
<td> 55, 56, 57</td><td> , 58, 59,</td><td> 60, 61, 62, 63,</td><td> 64,</td>
<td> 70, 71, 72</td><td> , 73, 74,</td><td> 75, 76, 77, 78,</td><td> 79,</td>
<td> 85, 86, 87</td><td> , 88, 89,</td><td> 90, 91, 92, 93,</td><td> 94,</td>
<td> 100, 101,</td><td> 102, 103,</td><td> 104, 105, 106,</td><td> 107,</td>
<td> 112, 113,</td><td> 114, 115,</td><td> 116, 117, 118,</td><td> 119,</td>
<td> 124, 125,</td><td> 126, 127,</td><td> 128, 129, 130,</td><td> 131,</td>
<td> 136, 137,</td><td> 138, 139,</td><td> 140, 141, 142,</td><td> 143,</td>
<td> 148, 149,</td><td> 150, 151,</td><td> 152, 153, 154,</td><td> 155,</td>
<td> 160, 161,</td><td> 162, 163,</td><td> 164, 165, 166,</td><td> 167,</td>
157, 158, 159, 160, 169, 170, 171, 172, 181, 182, 183, 184, 193, 194, 195, 196, 205, 206, 207, 208, 217, 218, 219, 220, 229, 230, 231, 232, 241, 242, 243, 244, or longer. 'oligonucleotide of, 4, 5, 6, 7, 8, 9, 20, 21, 22, 23,24,
35, 36, 37, 38,39,
50, 51, 52, 53,54,
65, 66, 67, 68,.69,
80, 81, 82, 83,84,
95, 96, 97, 98,99,
108, 109, 110, 111, 120, 121, 122, 123, 132, 133, 134, 135, 144, 145, 146, 147, 156, 157, 158, 159, 168, 169, 170, 171,
<td> 172,</td><td> 173,</td><td> 174,</td><td> 175,</td><td> 176, 177,</td><td> 178,</td><td> 179,</td><td> 180,</td><td> 181,</td><td> 182,</td><td> 183</td>
<td> 184,</td><td> 185,</td><td> 186,</td><td> 187,</td><td> 188, 189,</td><td> 190,</td><td> 191,</td><td> 192,</td><td> 193,</td><td> 194,</td><td> 195</td>
<td> 196,</td><td> 197</td><td> , 198</td><td> , 199</td><td> , 200, 201,</td><td> 202</td><td> , 203,</td><td> 204,</td><td> 205,</td><td> 206,</td><td> 207</td>
<td> 208,</td><td> 209,</td><td> 210,</td><td> 211,</td><td> 212, 213,</td><td> 214,</td><td> 215,</td><td> 216,</td><td> 217,</td><td> 218,</td><td> 219</td>
<td> 220,</td><td> 221,</td><td> 222,</td><td> 223,</td><td> 224, 225,</td><td> 226,</td><td> 227,</td><td> 228,</td><td> 229,</td><td> 230,</td><td> 231</td>
<td> 232,</td><td> 233,</td><td> 234,</td><td> 235,</td><td> 236, 237,</td><td> 238,</td><td> 239,</td><td> 240,</td><td> 241,</td><td> 242,</td><td> 243</td>
<td> 244,</td><td> 245,</td><td> 246,</td><td> 247,</td><td>248, 249 or</td><td> 250</td><td colspan="2">nucleotides</td><td></td><td></td><td></td>
Once the contents of the cells are released into their respective partitions, the nucleic acids present there can be further processed within the partitions. According to the methods and systems described herein, the content of the nucleic acids of individual cells is generally provided with unique identifiers so that after characterization of these nucleic acids they can be attributed as derived from the same cell or cells. . The ability to attribute characteristics to individual cells or groups of cells is provided by the assignment of unique identifiers specifically to an individual cell or groups of cells, which is another advantageous aspect of the methods and systems described herein. In particular, unique identifiers, eg. For example, in the form of nucleic acid barcodes they are assigned or associated with individual cells or populations of cells to label the components of the cells (and as a consequence, their characteristics) with the unique identifiers. These unique identifiers are used to attribute; the components and characteristics of cells to an individual cell or group of cells. In some aspects, this is accomplished by codivision of individual cells or groups of cells with the unique identifiers. In some aspects, unique identifiers are provided in the form of oligonucleotides comprising nucleic acid barcode sequences that may be linked or otherwise associated with the nucleic acid content of individual cells or other components of cells and particularly to fragments of these nucleic acids. The oligonucleotides are divided such that as between the oligonucleotides in a given partition, the nucleic acid barcode sequences present there are the same, but as between different partitions, the oligonucleotides can and do have different barcode sequences or at least they represent a large number of different barcode sequences across all partitions in a given scan. In some aspects, only one nucleic acid barcode sequence may be associated with a given partition, although in some cases, two or more different barcode sequences may be present. .
Nucleic acid barcode sequences can include between 6 and about 20 or more nucleotides within the oligonucleotide sequence. In some cases, the length of a barcode sequence can be 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19/20 nucleotides or longer. In some cases, the length of a barcode sequence can be at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or more. long. In some cases, the length of a barcode sequence can be a maximum of 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or more short. These nucleotides can be completely contiguous, ie, in a single stretch of adjacent nucleotides or they can be separated into two or more separate subsequences that are separated by 1 or more nucleotides. In some cases, the separate barcode subsequences can be between about 4 and about 16 nucleotides in length. In some cases, the barcode subsequence can be 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, nucleotides or longer. In some cases, the barcode subsequence can be at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some cases, the barcode subsequence can be a maximum of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter. :
The codivided oligonucleotides may also comprise other functional sequences useful for the processing of nucleic acids from codivided cells. These sequences include, e.g. e.g., target or random / universal amplification primer sequences to amplify genomic DNA of individual cells within partitions while binding associated barcode sequences, sequencing primers or primer recognition sites, hybridization sequences or priming, e.g. eg, for the identification of the presence of the sequences or to decrease barcoded nucleic acids or any of a number of other possible functional sequences. Again, the codivision of oligonucleotides and associated barcodes and other functional sequences, along with sample materials, is described, for example, in U.S. Patent Application No. 61 / 940,318, filed February 7,
2014, 61 / 991,018, filed May 9, 2014, and U.S. Patent Application No. 14 / 316,383, filed June 26, 2014, as well as U.S. Patent Application No. 14 / 175,935, filed 7 February 2014, which are incorporated herein in their entirety by this reference for all purposes. As will be appreciated, other oligonucleotide codivision mechanisms may also be employed, including, e.g. eg, coalescence of two or more droplets, where one droplet contains oligonucleotides or microdispense of oligonucleotides in partitions, e.g. eg, droplets within microfluidic systems.
Briefly, in one example, beads, microparticles, or microcapsules are provided that each include large amounts of the oligonucleotides described above loosely bound to the beads, wherein all oligonucleotides bound to a particular bead include the same sequence of nucleic acid barcode, but where a large number of diverse barcode sequences are represented among the population of beads used. In particularly useful examples, hydrogel beads, e.g. For example, comprising polyacrylamide polymer matrices, they are used as a solid support and delivery vehicle for the oligonucleotides to the partitions, as they are capable of transporting large amounts of oligonucleotide molecules and can be configured to release these oligonucleotides upon removal. exposure to a particular stimulus, as described elsewhere herein. In some cases, the bead population will provide a diverse barcode sequence library that includes at least 1,000 different barcode sequences, at least 5,000 different barcode sequences, at least 10,000 different barcode sequences, at least 50,000 different barcode sequences, at least 100,000 different barcode sequences, at least 1,000,000 different barcode sequences, at least 5,000,000 different barcode sequences or at least 10,000,000 different barcode sequences. Additionally, each bead can be provided with large amounts of attached oligonucleotide molecules. In particular, the number of oligonucleotide molecules including the barcode sequence in a single bead can be at least 1,000 oligonucleotide molecules, at least 5,000 oligonucleotide molecules, at least 10,000 oligonucleotide molecules, at least 50,000 oligonucleotide molecules , at least 100,000 oligonucleotide molecules, at least 500,000 oligonucleotides, at least 1,000,000 oligonucleotide molecules, at least 5,000,000 oligonucleotide molecules, at least 10,000,000 oligonucleotide molecules, at least 50,000,000 oligonucleotide molecules, at least 100,000,000 oligonucleotide molecules, and in some cases at least 1 billion molecules of oligonucleotides.
Furthermore, when the population of beads is divided, the resulting population of partitions can also include a diverse barcode library that includes at least 1,000 different barcode sequences, at least 5,000 different barcode sequences, at least 10,000 different barcode sequences, at least 50,000 different barcode sequences, at least 100,000 different barcode sequences, at least 1,000,000 different barcode sequences, at least 5,000,000 different barcode sequences, or at least 10,000,000 different barcode sequences. In addition, each population partition may include at least 1,000 oligonucleotide molecules, at least 5,000 oligonucleotide molecules, at least 10,000 oligonucleotide molecules, at least 50,000 oligonucleotide molecules, at least 100,000 oligonucleotide molecules, at least 500,000 oligonucleotides, at least less than 1,000,000 oligonucleotide molecules, at least 5,000,000 oligonucleotide molecules, at least 10,000,000 oligonucleotide molecules, at least 50,000,000 molecules, of oligonucleotides, at least 100,000,000 molecules of oligonucleotides, and, in some cases, at least 1 billion molecules of oligonucleotides.
In some cases, it may be desirable to incorporate multiple different barcodes within a given partition, either attached to a single bead or to multiple beads within the partition. For example, in some cases, a mixed but known set of barcode sequences can provide greater assurance of identification in post-processing, e.g. For example, by providing stronger addressing or attribution of barcodes to a particular partition, such as a duplicate or independent confirmation of the output of a particular partition. ·
The oligonucleotides are shed from the beads upon application of a particular stimulus to the beads. In some cases, the stimulus may be a photostimulus, e.g. eg, by cleavage of a photolabile junction that releases the oligonucleotides. In other cases, a thermal stimulus can be used, where raising the temperature of the environment of the beads will result in cleavage of a junction or other release of the oligonucleotide form from the beads. In other cases, a chemical stimulus is used which clears a junction of the oligonucleotides to the beads or otherwise causes the release of the oligonucleotides from the beads. Examples of this type of system are described in US Patent Application No. 13 / 966,150, filed August 13, 2013, as well as US Provisional Patent Application No. 61 / 940,318, filed February 7, 2013. 2014, 61 / 991,018, filed May 9, 2014 and U.S. Patent Application No. 14 / 316,383, filed June 26, 2014, the full descriptions of which are incorporated herein in their entirety by reference for all purposes. In one case, these compositions include the polyacrylamide matrices described above for encapsulation of cells and can be degraded to release bound oligonucleotides through exposure to a reducing agent, such as DTT.
In accordance with the methods and systems described herein, the beads including the linked oligonucleotides are codivided with the individual cells, such that a single bead and a single cell are found within a single partition. As stated above, although the single cell / single bead occupancy is the state<sup>:</sup> More desired, it will be appreciated that multiply occupied partitions (either in terms of cells, beads, or both) or unoccupied partitions (either in terms of cells, beads, or both) will often be present. An example of a microfluidic channel structure for codivating cells and beads comprising barcode oligonucleotides is schematically illustrated in Figure 2. As described elsewhere herein, in some respects a substantial percentage of the overall occupied partitions will include a bead and a cell, and in some cases some of the partitions that are generated will be unoccupied. In some cases, some of the partitions may have beads and cells that do not divide 1: 1. In some cases, it may be desirable to provide multi-occupied partitions, eg. eg, containing two, three, four or more cells and / or beads within a single partition. As shown, channel segments 202, 204, 206, 208, and 210 are provided in fluid communication at the junction of channels 212. An aqueous stream comprising individual cells 214 is flowed through channel segment 202 toward the junction of channels 212. As described above, these cells can be suspended within an aqueous fluid or they can have been preencapsulated prior to the division process.
Simultaneously, an aqueous stream comprising the barcode-containing beads 216 is flowed through the channel segment 204 towards the junction of: the channels 212. A non-aqueous dividing fluid 216 is introduced at the junction of the channels 212 from each of the side channels 206 and 208 and the combined streams are caused to flow into the outlet channel 210. Within the junction of. channels 212, the two combined aqueous streams from channel segments 202 and 204 are combined and divided into droplets 218, which include codivided cells 214 and beads 216. As noted above, by controlling the flow characteristics of each of the fluids combining at the junction of the channels 212 and by controlling the geometry of the junction of the channels, the combination and division can be optimized to achieve a desired level of bead, cell, or both occupancy within partitions 218 that are generated. ·
In some cases, the lysing agents, e.g. eg, cell lysis enzymes, can be introduced into the partition with the bead stream, eg. eg, flowing through channel segment 204, so that cell lysis begins only at or after division. Other reagents can also be added to the partition in this configuration, such as endonucleases to fragment DNA from cells, DNA polymerase enzyme, and dNTPs used to amplify nucleic acid fragments from cells and to bind oligonucleotides from cells. barcode to amplified fragments. As stated above, in many cases, a chemical stimulus, such as DTT, can be used to release the barcodes from their respective beads to the partition. In these cases, it may be particularly desirable to provide the chemical stimulus along with the cell-containing stream in the channel segment 202, such that the release of. barcodes are only produced after the two streams have been combined, e.g. ex. , within partitions 218. When cells are encapsulated, however, the introduction of a common chemical stimulus, e.g. For example, releasing the oligonucleotides from the beads and releasing the cells from the microcapsules can generally be provided from a separate additional side channel (not shown) upstream or connected to the junction of channel 212.
As will be appreciated, a number of other reagents can be co-cleaved together with cells, beads, lysing agents, and chemical stimuli, including, for example, protective reagents, such as proteinase K, chelators, nucleic acid extension, replication, transcription or amplification reagents such as polymerases, reverse transcriptases, transposases that can be used for transposon-based methods (eg. , Nextera), nucleoside triphosphates or NTP analogs, primer sequences and additional cofactors such as bivalent metal ions used in such reactions, ligation reaction reagents, such as ligase enzymes and ligation sequences, inks, labels or other labeling reagents.
Channel networks, p. ex. , as described herein, may be fluidly coupled with suitable fluid components. For example, input channel segments, eg. For example, channel segments 202, 204 ,: 206 and 208 are fluidly coupled with suitable sources of the materials to be supplied to the junction of channels 212. For example, channel segment 202 will be fluidly coupled with a source of an aqueous suspension of cells 214 for analysis, while channel segment 204 will be fluidly coupled with a source of an aqueous suspension of beads 216. Channel segments 206 and 208 will be fluidly connected to one or more non-aqueous fluid sources. These sources may include any of a number of different fluid components, from simple reservoirs defined or connected to a body structure of a microfluidic device, to fluid conduits supplying fluids from sources outside the device, manifolds, or the like. Also, the outlet channel segment 210 may be fluidly coupled with a receiving container or conduit for the divided cells. Again, this can be a defined reservoir in the body of a microfluidic device or it can be a fluid conduit for supplying the divided cells to a downstream operation, instrument or component.
Figure 8 shows images of individual Jurkat cells codivided together with barcode oligonucleotide containing beads in aqueous droplets in a water-in-oil emulsion. As illustrated, individual cells can easily be co-divided with individual beads. As will be appreciated, optimization of individual cell loading can be accomplished through various methods, including providing dilutions of cell populations to the microfluidic system to achieve the desired cell loading per partition as described elsewhere herein. .
In operation, once lysed, the nucleic acid content of individual cells is available for further processing within the partitions, including, e.g. eg, fragmentation, amplification, and barcode creation, as well as binding of other functional sequences. As indicated above, fragmentation can be achieved by codivision of shear enzymes, such as endonucleases, to fragment nucleic acids' into smaller fragments. These endonucleases can include restriction endonucleases, including type II and type II restriction endonucleases, as well as other nucleic acid cleavage enzymes, such as cut endonucleases and the like. In some cases, fragmentation may not be desired and full-length nucleic acids can be retained within the partitions or in the case of encapsulated cells or cell contents, fragmentation can be performed before division , p. eg, through enzymatic methods, eg. g., those described herein or by mechanical methods, eg. eg, mechanical, acoustic or other cutting.
Once the cells have been codivided and lysed to release the nucleic acids, the oligonucleotides arranged on the bead can be used to barcode them and amplify fragments of these nucleic acids. A particularly elegant process for using these barcode oligonucleotides to amplify and barcode the sample nucleic acid fragments is described in detail in US Provisional Patent Application No. 61 / 940,318, filed 7/7. February 2014, 61 / 991,018, filed May 9, 2014 and U.S. Patent Application No. 14 / 316,383 filed June 26, 2014, previously incorporated by this reference. In summary, in one aspect, the oligonucleotides present on the beads that are codivided with the cells are released from their beads upon partition with the nucleic acids of the cells. The oligonucleotides can include, along with the barcode sequence, a primer sequence at their 5 'end. This primer sequence can be a random oligonucleotide sequence intended to randomly prime a number of different regions in cell nucleic acids or it can be a specific primer sequence directed to prime upstream of a specific target region of the cell genome. .
Once released, the primer portion of the oligonucleotide can hybridize to a complementary region of the nucleic acid of the cells. Extension reaction reagents, e.g. e.g. DNA polymerase, nucleoside triphosphates, cofactors (e.g. g., Mg2 + or Mn2 +), which also codivide with cells and beads, then extend the primer sequence using the nucleic acid of the cells as a template, to produce a fragment complementary to the nucleic acid chain of the cells at the same time. which primer was hybridized, wherein the complementary fragment includes the oligonucleotide and its associated barcode sequence. The hybridization and extension of multiple primers to different parts of the nucleic acids of the cells cause a large group of overlapping complementary nucleic acid fragments, each having its own barcode sequence that indicates the partition in which it was created . In some cases, these complementary fragments can be used themselves as a template primed by the oligonucleotides present in the partition to produce a complement.<sup>:</sup> of the plugin that, again, includes the barcode sequence. In some cases, this repeating process is configured so that when the first complement duplicates, it produces two complementary sequences in<sup>:</sup> the ends or near them, to allow the formation of a hairpin structure or partial hairpin structure, reduces the ability of the molecule to be the base to produce other iterative copies. As described herein, the nucleic acids of the cells can include any desired nucleic acids within the cell including, for example, the DNA of the cell, e.g. eg, genomic DNA, RNA, eg. eg, messenger RNA and the like. For example, in some cases, the methods and systems described herein are used for the characterization of expressed mRNA, including, e.g. g., the presence and quantification of said mRNA, and may include RNA sequencing processes such as the characterization process. Additionally or alternatively, the reagents split together with the cells may include reagents for the conversion of mRNA to cDNA, e.g. eg, reverse transcriptase enzymes and reagents, to facilitate sequencing processes where DNA sequencing is used. In some cases, when the nucleic acids to be characterized comprise RNA, e.g. ex. , MRNA, a schematic illustration of an example of this is shown in figure 3.<sup>:</sup>
As shown, oligonucleotides that include a barcode sequence are codivided, e.g. ex. , in a droplet 302 in an emulsion, together with a nucleic acid from sample 304. As highlighted elsewhere herein, oligonucleotides 308 can be provided in a bead 306 that is codivided with nucleic acid from sample 304 , wherein the oligonucleotides are detached from bead 306, as shown in panel A.<sup>:</sup> Oligonucleotides 308 include a 312 barcode sequence in addition to one or more functional sequences, e.g. eg, sequences 310, 314, and 316. For example, oligonucleotide 308 is shown which comprises barcode sequence 312, as well as sequence 310 which can act as a binding or immobilization sequence for a given sequencing system. , p. eg, a · P5 sequence used for binding in flow cells of an Illumina Hiseq® or Miseq® system. As shown, the oligonucleotides also include a primer sequence 316, which may include a random or targeted N-mer for repeat priming of nucleic acid portions of sample 304. Also included within oligonucleotide 308 is a 314 sequence that can provide a sequencing priming region, such as a readl or Rl priming region, which is used to prime template-directed, polymerase-mediated sequencing by synthesis reactions in the sequencing systems. As will be appreciated, functional sequences can be selected to be compatible with a variety of different sequencing systems, e.g. eg 454 sequencing, Ion Torrent Proton or PGM, Illumina X10, etc., and your requirements. In many cases, the 312 barcode sequence, 310 immobilization sequence, and Rl 314 sequence may be common for all oligonucleotides bound to a given bead. The primer sequence 316 can vary for random N-mer primers or it can be common for oligonucleotides in a given bead for certain targeted applications.
As will be appreciated, in some instances, functional sequences may include primer sequences useful for RNA sequencing applications. For example, in some cases, the oligonucleotides may include poly-T primers to prime the reverse transcription of RNA for RNA sequencing. In still other cases, the oligonucleotides in<sub>:</sub> a certain partition, p. E.g., included in a single bead, can include multiple types of primer sequences in addition to the common barcode sequences, eg, DNA sequencing and RNA sequencing primers, e.g. eg, poly-T primer sequences included within bead-coupled oligonucleotides. In such cases, a single sequencing cell can undergo both DNA and RNA sequencing processes.
Based on the presence of primer sequence 316,: oligonucleotides can prime nucleic acid in the sample as shown in panel B, allowing extension of oligonucleotides 308 and 308a using enzymes, polymerase, and other extension reagents also codivided with bead 306 and nucleic acid from sample 304. As shown in panel C, after extension of the oligonucleotides that, for random N-mer primers, would hybridize to multiple different regions of the nucleic acid from sample 304; multiple overlapping nucleic acid complements or fragments are created, e.g. ex. , fragments 318 and 320. Although they include parts of sequences that are complementary to parts of the nucleic acid in the sample, e.g. For example, sequences 322 and 324, these constructs are generally referred to herein as comprising fragments of the nucleic acid from sample 304, having the barcode sequences attached. .
The barcoded nucleic acid fragments can be subjected to characterization, e.g. g., by sequence analysis, or can be further amplified in the process, as shown in panel D. For example, additional oligonucleotides, eg. Eg, oligonucleotide 308b, also released from bead 306, can prime fragments 318 and 320. This is shown in fragment 318. In particular, again, based on the presence of random N-mer primer 316b in oligonucleotide 308b (which in many cases may be different from the other random N-mer in a given partition, eg. , primer sequence 316), the oligonucleotide hybridizes to fragment 318 and is extended to create a complement 326 for at least a portion of fragment 318 that includes sequence 328 comprising a duplicate of a portion of the nucleic acid sequence of the sample. The extension of oligonucleotide 308b continues until it has replicated through the oligonucleotide 308 portion of fragment 318. As noted elsewhere herein, and as illustrated in panel D, oligonucleotides can be configured to cause a stop in repeat by the polymerase at a desired point, e.g. ex. , after repetition through sequences 316 and 314 of oligonucleotide 308 that is included within fragment 318. As described herein, this can be achieved through different methods, including, for example, the incorporation of different nucleotides and / or nucleotide analogs that are not capable of being processed by the polymerase enzyme used. For example, this may comprise including nucleotide-containing uracil within the 312 sequence region to prevent a uracil non-tolerant polymerase from repeating that region. As a result, a fragment 326 is created that includes the full-length oligonucleotide 308b at one end, which includes the barcode sequence 312, the junction sequence 310, the primer region Rl 314, and the random N-mer sequence 316b. At the other end of the sequence, complement 316 'to the random N-mer of the first oligonucleotide 308 can be included, as well as a complement of all or part of the sequence Rl,: shown as sequence 314'. The sequence Rl 314 and its complement 314 'are capable of hybridizing together to form a partial hairpin structure 328. As will be appreciated because the random N-mers differ between different oligonucleotides, no, these sequences and their complements are expected participate in the formation of hairpins, p. For example, the sequence 316 ', which is the complement of random N-mer 316, is not expected to be complementary to the sequence of random N-mer 316b. This would not be the case for other applications, eg. eg, targeted primers,<sup>;</sup> where the N-mer would be common among the oligonucleotides within a given partition.
By forming these partial hairpin structures, it allows the removal of first-level duplicates from the sample sequence of an additional repeat, e.g. eg, prevents iterative copying of copies. The partial hairpin structure also provides useful structure for post-processing of created fragments, e.g. eg, fragment 326.
In general, the amplification of the nucleic acids of the cell is carried out until the overlapping barcoded fragments within the partition constitute at least one IX coverage of the particular portion or the entire genome of the cell, at least 2X , at least 3X, at least 4X, at least 5X, at least 10X, at least 20X, at least 40X or more coverage of the genome or its relevant portion of interest. Once the fragments are produced, they can be directly sequenced in a suitable sequencing system, e.g. ; eg, an Illumina Hiseq®, Miseq®, or X10 system, or may undergo additional processing, eg, additional amplification, binding of other functional sequences, e.g. eg, secondary sequencing primers, for reverse reads, sample index sequences, and the like.
All fragments from multiple different partitions can then be pooled for sequencing in high-throughput sequencers as described herein, where the pooled fragments comprise a large number of fragments derived from the nucleic acids of different cells or small cell populations, but where the nucleic acid fragments of a given cell share the same barcode sequence. In particular, because each fragment is encoded in terms of its partition of origin and, consequently, its unique cell or small population of cells, the sequence of that fragment can again be attributed to that cell or cells based on the presence of the barcode, which also helps to apply the various fragments of sequence from the multiple partitions to the assembly of individual genomes for different cells. This is schematically illustrated in Figure 4. As shown in one example, a first nucleic acid 404 from a first cell 400, and a second nucleic acid 406 from a second cell 402 are codivided along with their own sets of oligonucleotides. barcode as described above. Nucleic acids can comprise a chromosome, a whole genome, or other large nucleic acid in cells.
Within each partition, the nucleic acids from each cell 404 and 406 are processed to separately provide an overlapping set of secondary fragments from the primary fragments, e.g. ex. , sets of secondary fragments 408 and 410. This processing also provides the second fragments with a barcode sequence that is the same for each of the second fragments derived from a particular first fragment. As shown, the barcode sequence for secondary fragment set 408 is indicated by 1, while the barcode sequence for fragment set 410 is indicated by 2. A diverse library of barcodes can be used. used to differentially assign barcodes to large numbers of different chunk sets. However, each set of secondary fragments of a different primary fragment need not be assigned a barcode with different barcode sequences. In fact, in many cases, multiple different primary fragments can be processed simultaneously to include the same barcode sequence. Various bar code libraries are described in detail elsewhere herein.<sup>:</sup>
Barcode snippets, eg. From the fragment sets 408 and 410 can be pooled for sequencing using, for example, sequencing by synthesis technologies available from Illumina or the Ion Torrent division from Thermo-Fisher, Inc. Once sequenced, the 412 sequence reads can be attributed to their respective fragment set, e.g. e.g., as shown in aggregate readings 414 and 416, which are based, at least in part, on the included barcodes and in some cases, in part on<sup>:</sup> the sequence of the fragment itself. The attributed sequence reads for each fragment set are assembled to provide the assembled sequence for each nucleic acid in the cell, e.g. ex. , sequences 418 and 420, which, in turn, can be attributed to individual cells, e.g. eg, cells 400 and 402.
Although described in terms of analysis of genetic material present within cells, the methods and systems described herein may have much broader applicability, including the ability to characterize other aspects of individual cells or cell populations, allowing for the assignment of reagents to individual cells, and providing the attributable analysis or characterization of those cells in response to those reagents. These methods and systems: are particularly valuable because they are capable of characterizing cells, e.g. eg for research, diagnosis, pathogen identification, and many other purposes. By way of example, a wide range of different cell surface characteristics, e.g. ex. Cell surface proteins as a cluster of differentiation proteins or DCs have significant diagnostic relevance in the characterization of diseases such as cancer. :
In a particular useful application, the methods and systems described herein can be used to characterize cellular features, eg, cell surface features, e.g. eg, proteins, receptors, etc. In particular, the methods described herein can be used to bind reporter molecules to those cellular features, which when divided as described above can be barcoded and analyzed, e.g. ex. , using DNA sequencing technologies, to determine the presence and, in some cases, the relative abundance or quantity, of said cellular characteristics within an individual cell or population of cells. :
In a specific example, - a library of potential cell-binding ligands can be provided, e.g. g., antibodies, antibody fragments, cell surface receptor binding molecules, or the like, associated with a first set of nucleic acid reporter molecules, e.g. ex. , where a different reporter oligonucleotide sequence is associated with a specific ligand and is therefore capable of binding to a specific cell surface feature. In some aspects, different members of the library may be characterized by the presence of a different oligonucleotide sequence tag, e.g. Eg, an antibody to a first type of cell surface protein or receptor is associated with a first known reporter oligonucleotide sequence, whereas an antibody to a second receptor protein has. an associated different reporter oligonucleotide sequence. Before co-dividing, cells are incubated with the ligand library, which can represent antibodies to a wide panel of different cell surface characteristics, e.g. g., receptors, proteins, etc., and including their associated reporter oligonucleotides. Unbound ligands are washed from the cells, and the cells are codivided together with the barcode oligonucleotides described above. As a result, the partitions include the cell or cells, as well as the binding ligands and their associated known reporter oligonucleotides. .
Without the need to lyse the cells within the partitions, the reporter oligonucleotides can be subjected to the barcode creation operations described above for cellular nucleic acids, to produce barcoded reporter oligonucleotides, where the presence of the reporter oligonucleotides may indicate the presence of the particular cell surface feature, and the barcode sequence allows the attribution of the range of different cell surface characteristics to a particular individual cell or population of cells based on the barcode sequence that was encoded with that cell or population of cells. As a result, a cell-by-cell profile of cell surface characteristics can be generated within a broader cell population. This aspect of the methods and systems described herein is described in detail below.
This example is illustrated schematically in Figure 5. As shown, a population of cells, represented by cells 502 and 504, is incubated with a library of cell surface associated reagents, e.g. eg, antibodies, cell surface binding proteins, ligands or the like, where each different type of binding group includes an associated nucleic acid reporter molecule, shown as ligands and associated reporter molecules 506, 508, 510 and 512 (where the reporter molecules are indicated by the differently shaded circles). When the cell expresses the surface features that are bound by the library, the ligands and their associated reporter molecules can associate or dock with the cell surface. Then individual cells divide into separate partitions, e.g. eg -, droplets 514 and 516, together with their associated reporter ligands / molecules, as well as a single barcode oligonucleotide bead as described elsewhere herein, e.g. ex. , beads 522 and 524, respectively. As with other examples described herein, the barcoded oligonucleotides are released from the beads and used to match the barcode sequence that the reporter molecules present within each partition with a barcode that is common to a certain partition, but it varies widely between different partitions. For example, as shown in Figure 5, reporter molecules that associate with cell 502 in partition 514 are labeled with the 518 barcode sequence, while reporter molecules associated with cell 504 in partition 516 They are marked with barcode 520. As a result, an oligonucleotide library is provided that reflects the cell's surface ligands, as reflected by the reporter molecule, but is substantially attributable to an individual cell by virtue of a common barcode sequence, allowing profiling to single cell level of cell surface characteristics. As will be appreciated, this process is not limited to cell surface receptors but can be used to identify the presence of a wide variety of cell structures, chemicals, or other specific characteristics.
III. Single cell analysis applications
There are a wide variety of different applications for the single cell analysis and processing methods and systems described herein, including analysis of specific individual cells, analysis of different cell types within populations of different cell types, analysis and characterization of large populations of cells for environmental, human health, forensic epidemiological applications, or any of a wide variety of different applications.
A particularly valuable application of the single cell assay processes described herein is in the sequencing and characterization of cancer cells; In particular, conventional analytical techniques, even the aforementioned co-sequencing processes, are not very efficient at picking up small variations in the genomic composition of cancer cells, particularly when these exist in a sea of normal tissue cells. Furthermore, even between humoral cells, large variations can exist, and can be masked through joint approaches to sequencing (see, eg, Patel, et al., Single-cell RNA-seq highlights intratumoral heterogeneity in primary glioblastoma, Science DOI: 10.1126 / science.1254257 (published online June 12, 2014). Cancer cells can be derived from solid tumors, hematological neoplasms, cell lines, or they can be obtained as circulating tumor cells, and subjected to the division processes described above. Upon analysis, individual cell sequences derived from a single cell or a small group of cells can be identified and distinguished over normal tissue cell sequences. Furthermore, as described in co-pending US Provisional Patent Application No. 62 / 017,808, filed June 26, 2014, the disclosure of which is incorporated herein in its entirety by reference for all purposes, may also be obtain sequence information in phases of each cell, allowing a clearer characterization of the haplotype variants within a cancer cell. The single cell analysis approach is particularly useful for systems and methods that involve low amounts of input nucleic acids, as described in co-pending US Provisional Patent Application No. 62 / 017,580, filed June 26, 2014, the description of which is incorporated herein in its entirety by reference for all purposes, i ··
As with the analysis of cancer cells, the analysis and diagnosis of fetal health or abnormality through <sup>;</sup> The analysis of fetal cells is a difficult task using conventional techniques. In particular, in the absence of relatively invasive procedures, such as amniocentesis that obtain samples of fetal cells, they may employ collecting those cells from the maternal circulation. As you will see, these circulating fetal cells make up an extremely small fraction of the general cell population of that circulation. As a result, complex analyzes are performed to characterize what from the data obtained is likely to be derived from fetal cells rather than maternal cells. By employing the single cell characterization methods and systems described herein, however, individual cells can be attributed genetic makeup and those cells categorized as maternal or fetal based on their respective genetic makeup. Furthermore, the genetic sequence of fetal cells can be used to identify any number of genetic disorders, including eg. ex. , aneuploidy such as Down syndrome, Edwards syndrome and Patau syndrome.
The ability to characterize individual cells from diverse large cell populations is also of significant value in both environmental assessment and forensic analysis, where samples, by their nature, may be composed of various populations of cells and other material that contaminate the sample, with respect to the cells for which the sample is evaluated, e.g. eg, environmental indicator organisms, toxic organisms and the like, eg. eg, for environmental and food safety testing, victim and / or perpetrator cells in forensic analysis for sexual assault, other violent crimes, and the like.
Some additional useful applications of the single cell sequencing and characterization processes described above are in the field of neuroscience research and diagnosis. In particular, neural cells can include long scattered nuclear elements (LINE) or 'jumping' genes that can move around the genome, causing each neuron to differ from its neighboring cells. Research has shown that the amount of LINE in the human brain exceeds that of other tissues, eg. g., heart and liver tissue, between 80 and 300 unique insertions (see, eg, Coufal, NG et al. Nature 460, 1127-1131 (2009)). These differences have been postulated to be related to a person's susceptibility to neurological disorders (see, eg, Muotri, A.
R. et al. Nature 468, 443-446 (2010)), or that they provide the brain with diversity with which to respond to challenges. As such, these methods described herein can be used in the sequencing and characterization of individual neural cells.
The single cell analysis methods described herein are also useful in the analysis of gene expression, as indicated above, in terms of identification of RNA transcripts and their quantification. In particular, using the single cell level analysis methods described herein, RNA transcripts present in individual cells, cell populations or subsets of cell populations can be isolated and analyzed. In particular, in some cases, the barcode oligonucleotides may be configured to prime, replicate, and therefore provide barcoded RNA fragments from individual cells. For example, in some cases,. barcode oligonucleotides can include mRNA-specific primer sequences, e.g. ex. , poly-T primer segments that allow priming and replication of mRNA in a reverse transcription reaction or other targeted primer sequences. Alternatively or additionally, random RNA priming can be carried out using random N-mer primer segments from the barcode oligonucleotides.
Figure 6 provides a schematic of an exemplary method for analysis of RNA expression in single cells using the methods described herein. As shown, in step 602, a sample containing cells is sorted for viable cells, which are quantified and diluted for further division. In run 604, individual cells are codivided separately with gel beads having the barcode creating oligonucleotides as described herein. The cells are lysed and the barcoded oligonucleotides are released into the partitions in run 606, where they interact and hybridize with the mRNA in run 608, e.g. g., by virtue of a poly-T primer sequence, which is complementary to the poly-A tail of the mRNA. Using the poly-T barcode oligonucleotide as the primer sequence, at step 610 a reverse transcription reaction is performed to synthesize; a cDNA transcript of the mRNA that includes the barcode sequence. The barcoded cDNA transcripts are then subjected to further amplification at step 612, e.g. eg, using a PCR process, purification in step 614, before being placed into a nucleic acid sequencing system for the determination of the cDNA sequence and its associated barcode sequences. In some cases, as shown, operations 602 to 608 may occur while the reagents remain in their original droplet or partition, while operations 612 to 616 may occur in bulk (e.g. outside of partition) . In the case where a partition is a droplet in an emulsion, the emulsion can be broken and the contents of the droplet can be grouped together to complete operations from 612 to 616. In some cases, the barcode oligonucleotides can be digested with exonucleases after breaking the emulsion. The exonuclease activity can be inhibited by ethylenediaminetetraacetic acid (EDTA) after digestion of the primer. In some cases, operation 610 can be performed within the partitions based on the codivision of the reverse transcription mix, e.g. ex. , reverse transcriptase and associated reagents, or it can be done in bulk.
As noted elsewhere herein: the structure of barcode oligonucleotides can include a number of sequence elements in addition to the oligonucleotide barcode sequence. An example of a barcode oligonucleotide for use in RNA analysis is shown in Figure 7 as described above. As shown, general oligonucleotide 702 is coupled to a 704 bead via a releasable bond 706, for example a disulfide linker. The oligonucleotide can include functional sequences that are used in further processing, for example the 708 functional sequence, which can include one or more sequencer-specific flow cell binding sequences, e.g. g., to a P5 sequence for Illumina sequencing systems, as well as sequencing primer sequences, e.g. eg, an R1 primer for Illumina sequencing systems. A 710 barcode sequence is included within the framework for use in barcoding the sample RNA. An mRNA-specific primer sequence, eg, poly-T 712 sequence, is also included in the oligonucleotide structure. A 714 anchor sequence segment can be included to ensure that the poly-T sequence hybridizes at the sequence end of the mRNA. This anchor sequence can include a short random nucleotide sequence, e.g. eg, 1-mer,
2-mer, 3-mer, or a larger sequence, which ensures that the poly-T segment is more likely to hybridize to a sequence end of the poly-A tail of the mRNA. An additional sequence segment 716 may be provided within the oligonucleotide sequence. In some cases, this additional sequence provides a unique molecular sequence segment, e.g. ex. , as a random sequence (p. g., such as a random N-mer sequence) that varies across individual oligonucleotides coupled to a single bead, whereas the 710 barcode sequence can be constant between oligonucleotides attached to a single bead. This unique sequence serves to provide a unique identifier for the starting mRNA molecule that was captured, to allow quantification of the number of original expressed RNA. As will be appreciated, although shown as a single oligonucleotide attached to the surface of a bead, the individual bead can include tens to hundreds of thousands or even millions of individual oligonucleotide molecules, where, as noted, the code segment The number of bars can be constant or relatively constant for a given bead, but where the variable or unique sequence segment varies throughout an individual bead. This unique molecular sequence segment can include between 5 and about 8 or more nucleotides within the oligonucleotide sequence. In some cases, the unique molecular sequence segment can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, i 18, 19 or 20 nucleotides in length or longer. In some cases, the unique molecular sequence segment can be at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides in length or longer. In some cases, the single molecular sequence segment can be at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides in length or shorter.
In operation, and referring to Figures 6 and 7, a cell is codivided together with a bead having a barcode and is lysed as the barcoded oligonucleotides are released from the bead. The poly-T portion of the released barcode oligonucleotide then hybridizes to the poly-A tail of the mRNA. The poly-T segment then primes reverse transcription of the mRNA to produce a cDNA transcript of the mRNA, but that includes each of the 708-716 sequence segments of the barcode oligonucleotide. Again, because oligonucleotide 702 includes a 714 anchor sequence, it is more likely to hybridize and prime reverse transcription at the sequence end of the poly-A tail of the mRNA. Within any given partition, all cDNA transcripts of individual mRNA molecules include a common 710 barcode sequence segment. However, but including the unique random N-mer sequence, transcripts made from different molecules of mRNA within a given partition vary in this unique sequence. This provides a quantization characteristic that can be identified even after any subsequent amplification of the contents of a given partition, e.g. For example, the number of unique segments associated with a common barcode can indicate the amount of mRNA that originates from a single partition, and thus a single cell. As indicated above, the transcripts are then amplified, cleaned up, and sequenced to identify the sequence of the cDNA transcript of the mRNA, as well as to sequence the barcode segment and the unique sequence segment. '
As noted elsewhere herein, while a poly-T primer sequence is described, other targeted or random primer sequences can also be used to prime the reverse transcription reaction. Also, while described as releasing the barcoded oligonucleotides into the partition along with the contents of the lysed cells, it will be appreciated that, in some cases, the oligonucleotides attached to gel beads can be used to hybridize and capture the mRNA. in the solid phase of the gel beads, to facilitate the separation of RNA from other cellular contents. <sup>;</sup>
A further example of a barcode oligonucleotide for use in RNA analysis, including analysis of messenger RNA (mRNA, including mRNA obtained from a cell) is shown in Figure 9A. As shown,<sup>:</sup> The general oligonucleotide 902 may be coupled to a bead 904 via a releasable bond 906, for example a disulfide linker. The oligonucleotide can include functional sequences that are used in further processing, for example functional sequence 908, which can include a sequencer-specific flow cell binding sequence, e.g. ex. , a P5 sequence for Illumina sequencing systems, as well as a 910 functional sequence, which may include sequencing primer sequences, e.g. eg, a primer binding site R1 for Illumina sequencing systems. A 912 barcode sequence is included within the framework for use in barcoding the sample RNA. Also included is an RNA-specific primer sequence (eg. , mRNA-specific), eg poly-T 914 sequence, in the oligonucleotide backbone. A segment can be included<sup>:</sup> anchor sequence (not shown) to ensure that the poly-T sequence hybridizes at the sequence end of the mRNA. An additional sequence segment 916 may be provided within the oligonucleotide sequence. This additional sequence can provide a unique molecular sequence segment, e.g. ex. , as a random N-mer sequence, that varies across individual oligonucleotides stacked to a single bead, whereas the 912 bar code sequence can be constant between oligonucleotides attached to a single bead. As described elsewhere herein, this unique sequence serves to provide a unique identifier of the starting mRNA molecule that was captured, to allow quantification of the number of original expressed RNA, e.g. eg, mRNA count. As will be appreciated, although shown as a single oligonucleotide attached to the surface of a bead, individual beads can include tens to hundreds of thousands or even millions of individual oligonucleotide molecules, where, as noted, the code segment Bar can be constant or relatively constant for a given bead, but where the variable sequence segment: or unique varies across an individual bead. .
In an exemplary method of analysis of cellular RNA (eg, mRNA) and with reference to Figure 9A, a cell is codivided together with a bead having barcode, switch oligonucleotide 924, and other reagents such as transcriptase conversely, a reducing agent and dNTP in one partition (eg, a droplet in an emulsion). At step 950, the cell is lysed while the 902 barcode oligonucleotides are released from the bead (eg. , by the action of the reducing agent), and the poly-T 914 segment of the released barcode oligonucleotide then hybridizes with the poly-A tail of the 920 mRNA which is released from the cell. Then, in step 952, the poly-T segment 914 is extended in a reaction; Reverse transcription that uses mRNA as a template to produce a 922 cDNA transcript complementary to the 20 mRNA and also includes each of the sequence segments
908, 912, 910, 916 and 914 of the oligonucleotide code of. The terminal transferase activity of the reverse transcriptase can add additional bases to the transcript of cDNA (eg, polyC). Switch oligonucleotide 924 can then hybridize to the additional bases added to the cDNA transcript and facilitate template exchange. A sequence complementary to the switch oligonucleotide sequence can then be incorporated into the 922 cDNA transcript by extension of the 922 cDNA transcript using the 924 switch oligonucleotide as a template. Within any given partition, all cDNA transcripts of the 10 individual mRNA molecules include a segment: 912 common bar coding sequence. However, but including the unique random N-mer sequence 916, transcripts made from different mRNA molecules within a given partition vary in this unique sequence. As described elsewhere herein, this provides a quantization characteristic that can be identified even after any subsequent amplification of the contents of a given partition, e.g. For example, the number of unique segments associated with a common barcode can indicate the amount of mRNA originating from a single partition and hence a single cell. After run 952, the 922 cDNA transcript is amplified with primers 926 (eg, PCR primers) in run 954. The amplified product is then purified (eg, by solid phase reversible immobilization ( SPRI)) in operation 956. In step 958, the amplified product is cut, ligated to additional functional sequences, and further amplified (eg, by PCR). Functional sequences can include a 930 sequencer-specific flow cell binding sequence, e.g. eg, a P7 sequence for Illunkina sequencing systems, as well as a 928 functional sequence, which may include a sequencing primer binding site, e.g. ex. , for a primer R2 for Illumina sequencing systems, as well as a functional sequence 932, which can include a sample index, e.g. eg, an i7 sample index sequence for Illumina sequencing systems. In some cases, operations 950 and 952 can occur on the partition, while operations 954, 956, and 958 can occur in bulk solution (eg, in a pooled mix outside the partition). In the case where a partition is a droplet in an emulsion, the emulsion can be broken and the contents of the droplet can be grouped together to complete operations 954, 956, and 958. In some cases, operation 954 can be completed in the partition. In some cases, barcode oligonucleotides can be digested with exonucleases after emulsion breakdown. The exonuclease activity can be inhibited by ethylenediaminetetraacetic acid (EDTA) after digestion of the primer. Although they are described in terms of 5 references to specific sequences used for certain sequencing systems, e.g. For example, Illumina systems, it will be understood that reference to these sequences is for illustrative purposes only, and the methods described herein may be configured for use with 10 other sequencing systems that incorporate specific primer, binding, index and other operating sequences used in those systems, e.g. eg, systems available from Ion Torrent, Oxford Nanopore, Genia, Pacific Biosciences, Complete Genomics, and the like.
In an alternative example of a barcode oligonucleotide for use in RNA analysis (eg, cellular RNA) as shown in Figure 9A, functional sequence 908 can be a P7 sequence and functional sequence 910 can be be a binding site for primer R2. In addition, functional sequence 930 can be a P5 sequence, functional sequence 928 can be a primer binding site Rl, and functional sequence 932 can be an i5 sample index sequence for Illumina sequencing systems. The configuration of the constructs generated through such a barcode oligonucleotide can help to minimize (or avoid) the sequencing of the poly-T sequence during sequencing.
Another exemplary method for RNA analysis, including cellular mRNA analysis, is shown in Figure 9B. In this method, the switch oligonucleotide 924 is codivided with the single cell and the barcode bead along with reagents such as reverse transcriptase, a reducing agent, and dNTP in one partition (eg, a droplet in an emulsion). . Switch oligonucleotide 924 can be tagged with an additional ¡tag 934, e.g. ex. biotin. In step 951, the cell is lysed while the 902 barcode oligonucleotides (eg, as shown in Figure 9A) are released from the bead (eg, by the action of the reducing agent ). In some cases, sequence 908 is a P7 sequence and sequence 910 is a primer R2 binding site. In other cases, sequence 908 is a P5 sequence and sequence 910 is a primer binding site R1. Then, the poly-T 914 segment of the released barcode copy oligonucleotide hybridizes with the poly-A tail of 920 mRNA that is released from the cell. At step 953, poly-T segment 914 is then extended at '97.
a reverse transcription reaction that uses the mRNA as a template to produce a 922 cDNA transcript complementary to the mRNA and also includes each of the sequence segments 908, 912, 910, 916, and 914 of the oligonucleotide code for. The terminal transferase activity of the reverse transcriptase can add additional bases to the transcript of cDNA (eg, polyC). Switch oligonucleotide 924 can then hybridize to the cDNA transcript and facilitate template exchange. A sequence complementary to the switch oligonucleotide sequence can then be incorporated into the <(transcript of cDNA 922 by extension of the 922 cDNA transcript using the switch oligonucleotide 924 as a template. A 960 isolation run can then be used to isolate the 922 cDNA transcript from the reagents and oligonucleotides in the partition. Additional label 934, p. ex. biotin, can come into contact with a 936 interaction tag, e.g. g., streptavidin, which can be attached to a magnetic bead 938. In step 960, the cDNA can be isolated with a disassembly step (eg, by magnetic separation, centrifugation) prior to amplification (eg. by PCR) in run 955, followed by purification (e.g., by solid phase reversible immobilization (SPRI)) in run 957 and further processing (cleavage, ligation of; sequences 928, 932 and 930 and subsequent amplification (eg, by PCR)) in run 959. In some cases where sequence 908 is a P7 sequence and sequence 910 is a primer R2 binding site, sequence 930 is a P5 sequence and sequence 928 is a primer R1 binding site and sequence 932 is. a sample index sequence i5. In some cases where sequence 908 is qna sequence P5 and sequence 910 is a primer binding site Rl, sequence 930 is a sequence P7 and sequence 928 is a primer binding site R2 and sequence 932 is a sequence sample index i7. In some cases, as shown, operations 951 and 953 can occur on the partition, while operations 960, 955, 957, and 959 can occur in bulk solution (e.g., in a pooled mix outside the partition ). In the case where a partition is a droplet in an emulsion, the emulsion can be broken and the contents of the droplet can be grouped together to complete step 960. Las; Steps 955, 957 and 959 can be carried out after step 960 after pooling the transcripts for processing.
In Figure 9C it sequesters another exemplary method for RNA analysis, including cellular mRNA analysis. In this method, the switch oligonucleotide 924 is codivided with the single cell and the barcode bead along with reagents such as reverse transcriptase, a reducing agent, and dNTP in one partition (eg, a droplet in an emulsion ). At step 961, the cell is lysed while the 902 barcode oligonucleotides (eg. , as shown in Figure 9A) are released from the bead (eg, by the action of the reducing agent). In some cases, sequence 908 is a P7 sequence and sequence 910 is a primer R2 binding site. In other cases, sequence 908 is a P5 sequence and sequence 910 is a primer binding site R1. The poly-T 914 segment of the released barcode oligonucleotide then hybridizes to the poly-A tail of 920 mRNA that is released from the cell. Then, in run 963, the poly-T segment 914 is extended in a reverse transcription reaction that uses the mRNA 'as a template to produce a 922 cDNA transcript complementary to the mRNA and also includes each of the 908 sequence segments. , 912, 910, 916 and 914 of the oligonucleotide code of. The termiAal transferase activity of reverse transcriptase can
100 add additional bases to the cDNA transcript (eg, polyC). Switch oligonucleotide 924 can then hybridize to the cDNA transcript and facilitate template exchange. A sequence complementary to the switch oligonucleotide sequence can then be incorporated into the 922 cDNA transcript by extension of the 922 cDNA transcript using the 924 switch oligonucleotide as a template. After run 951 and run 963, the 920 mRNA and 922 cDNA transcript are denatured in run
962. In step 964, a second strand extends from a primer 940 that has an additional tag 942, e.g. ex. biotin, and hybridizes with the 922 cDNA transcript. In addition, at operation 964, the second biotin-tagged strand may come into contact with a 936 interaction tag, e.g. ex. , streptavidin, which can be attached to a 938 magnetic bead. The cDNA can be isolated, with a display operation (eg. g., by magnetic separation, centrifugation) before amplification (eg, by polymerase chain reaction (PCR)) in step 965, followed by purification (eg, by phase reversible immobilization solid (SPRI)) in operation 967 and additional processing (cutting, ligation
101 of sequences 928, 932, and 930 and subsequent amplification (eg, by PCR)) in step 969. In some cases where sequence 908 is a P7 sequence and sequence 910 is a primer R2 binding site, the sequence 930 is a P5 sequence and sequence 928 is a primer R1 binding site and sequence 932 is an i5 sample index sequence. In some cases where the sequence 908 is<sub>:</sub> a P5 sequence and sequence 910 is a primer R1 binding site, sequence 930 is a P7 sequence and sequence 928 is a primer R2 binding site and sequence 932 is a sample i7 index sequence. In some cases, operations 961 and 963 can occur on the partition, while operations 962, 964, 965, 967, and 969 can occur in bulk (eg, outside the partition). In the case where a partition is a droplet in. an emulsion, the emulsion can be broken and the droplet contents can be grouped together to complete operations 962, 964, 965, 967 and 969.
Another exemplary method for RNA analysis, including cellular mRNA analysis, is shown in Figure 9D. In this method, the 924 switch oligonucleotide is codivided with the single cell and the barcode bead together. with reagents such as reverse transcriptase, an agent
102 reducer and dNTP. In step 971, the cell is lysed while the 902 barcode oligonucleotides (eg, as shown in Figure 9A) are released from the bead (eg, by the action of the reducing agent) . In some cases, sequence 908 is a P7 sequence and sequence 910 is a primer R2 binding site. In other cases, sequence 908 is a P5 sequence and sequence 910 is a primer binding site R1. The poly-T 914 segment of the released barcode oligonucleotide then hybridizes to the poly-A tail of 920 mRNA that is released from the cell. Then, in run 973, the polyT 914 segment is extended in a reverse transcription reaction that uses the mRNA as a template to produce. a 922 cDNA transcript complementary to the mRNA and also includes each of the sequence segments 908, 912, 910, 916 and 914 of the oligonucleotide coding for. The terminal transferase activity of the reverse transcriptase can add additional bases to the transcript of cDNA (eg, polyC). Switch oligonucleotide 924 can then hybridize to the cDNA transcript and facilitate template exchange. A sequence complementary to the switch oligonucleotide sequence can then be incorporated into the 922 cDNA transcript by means of
103 922 cDNA transcription extension using switch oligonucleotide 924 as a template. In step 966, the 920 mRNA, 922 cDNA transcript, and 924 switch oligonucleotide can be denatured, and 922 cDNA transcript can be hybridized with a 944 capture oligonucleotide tagged with an additional 946 tag, e.g. eg, biotin. In this step, the biotin-tagged capture oligonucleotide 944, which hybridizes with the cDNA transcript, can come into contact with an interaction tag 936, e.g. g. streptavidin, which may be attached to a 938 magnetic bead. After separation from other species (eg, oversized barcoded oligonucleotides) using a disassembly operation (eg. g., by magnetic separation, centrifugation), the cDNA transcription can be amplified (eg, by PCR) with primers 926 in run 975, followed by purification (eg, by phase reversible immobilization solid (SPRI)) in run 977 and further processing (cleavage, ligation of sequences 928, 932, and 930, and subsequent amplification (eg, by PCR)) in run 979. In some cases where sequence 908 is a P7 sequence and sequence 910 is a primer R2 binding site, sequence 930 is a
104 sequence P5 and sequence 928 is a primer binding site Rl and sequence 932 is an index sequence of sample i5. In other cases where sequence 908 is a P5 sequence and sequence 910 is a primer R1 binding site, sequence 930 is a P7 sequence and sequence 928 is a primer R2 binding site and sequence 932 is a sample index i7 .. In some cases, operations 971 and 973 can occur on the partition, while operations 966, 975, 977 (purification), and 979 can occur in bulk (eg, outside of the partition). In the case where a partition is a droplet in an emulsion, the emulsion can be broken and the contents of the droplet can be grouped to complete operations 966, 975, 977 and 979..
Another exemplary method for RNA analysis, including cellular RNA analysis, is shown in Figure 9E. In this method, a single cell is codivided along with a bar-coded bead, a 990 switch oligonucleotide, and other reagents such as reverse transcriptase, a reducing agent, and dNTPs in one partition (eg, a droplet in an emulsion). In run 981, the cell is lysed while the barcoded oligonucleotides (eg. e.g.,: 902 as shown in figure 9A) are released from the bead
105 (eg, by the action of the reducing agent). In some cases, sequence 908 is a P7 sequence and sequence 910 is a primer R2 binding site. In other cases, sequence 908 is a P5 sequence and sequence 910 is a primer binding site R1. The poly-T segment of the released barcode oligonucleotide is then hybridized to the poly-A tail of 920 mRNA released from the cell. Then, in run 983, the poly-T segment is extended in a reverse transcription reaction to produce a 922 cDNA transcript complementary to the mRNA and also includes each of the 908, 912, 910, 916, and 914 sequence segments. of the oligonucleotide code of. The terminal transferase activity of the reverse transcriptase can add additional bases to the cDNA transcript (e.g. Eg, polyCJ. 990 switch oligonucleotide can then hybridize to the cDNA transcript and facilitate template exchange. A sequence complementary to the switch oligonucleotide sequence and including a T7 promoter sequence, can be incorporated into the 922 cDNA transcript. In run 968, a second strand is synthesized and in run 970 the T7 polymerase can use the T7 promoter sequence to produce RNA transcripts in in vitro transcription. In the
106 At run 985, the RNA transcripts can be purified (eg, by solid phase reversible immobilization (SPRI)), reverse transcribed to form DNA transcripts, and a second strand can be synthesized for each of the DNA transcripts. . In some cases, prior to purification, RNA transcripts may come into contact with DNase (eg, DNase I) to break up residual DNA. In run 987, the DNA transcripts are then fragmented and ligated to additional functional sequences, eg sequences 928, 932, and 930, and in some cases further amplified (eg, by PCR). In some cases where sequence 908 is a P7 sequence and sequence 910 is a primer R2 binding site, sequence 930 is a P5 sequence and sequence 928 is a primer R1 binding site and sequence 932 is a sequence sample index i5. In some cases where sequence 908 is a P5 sequence and sequence 910 is a primer Rl binding site, sequence 930 is a P7 sequence and sequence 928 is a primer R2 binding site and sequence 932 is a sample index i7. In some cases, before removing a portion of the DNA transcripts, the DNA transcripts may
107 come into contact with an RNase to break up residual DNA. In some cases, operations 981 and 983 can occur on the partition, while operations 968, 970, 985, and 987 can occur in bulk (eg, outside the partition). In the case where a partition is a droplet in<sup>:</sup> an emulsion, the emulsion can be broken and the contents of the droplet can be grouped together to complete operations 968, 970, 985 and 987.
Another example of a barcode oligonucleotide for use in RNA analysis, including analysis of messenger RNA (mRNA, including mRNA obtained from a cell) is shown in Figure 10. As shown, general oligonucleotide 1002 is coupled to a bead 1004 via a releasable bond 1006, for example a disulfide linker. The oligonucleotide can include functional sequences that are used in further processing, for example the 1008 functional sequence, which can include a sequencer-specific flow cell binding sequence, e.g. g., a P7 sequence, as well as a 1010 functional sequence, which can include sequencing primer sequences, e.g. ex. , a primer R2 binding site. A barcode sequence is included. 1012 within the framework for use in creating sample RNA barcodes. It can
108 include an RNA-specific primer sequence (eg, mRNA-specific), eg, poly-T 1014 sequence, in the oligonucleotide backbone. An anchor sequence segment (not shown) can be included to ensure that the poly-T sequence hybridizes at the sequence-end of the mRNA. An additional sequence segment 1016 may be provided within the oligonucleotide sequence. This additional sequence can provide a unique molecular sequence segment, as described elsewhere herein. An additional functional sequence 1020 can be included for in vitro transcription, e.g. eg, a T7 RNA polymerase promoter sequence. As will be appreciated, although shown as a single oligonucleotide attached to the surface of a bead,: individual beads can include tens to hundreds of thousands or even millions of individual oligonucleotide molecules, where, as noted, the segment of Bar code can be constant or relatively constant for a given bead, but where the unique or variable sequence segment varies across an individual bead.
In an exemplary method of cellular RNA analysis and referring to Figure 10, a cell is codivided together<sup>;</sup> with
109 a bead having a barcode, and other reagents such as reverse transcriptase, reducing agent, and dNTP in a partition (eg, a droplet in an emulsion). In step 1050, the cell is lysed in the meantime. The 1002 barcode oligonucleotides are released (eg, by the action of the reducing agent) from the bead, and the released barcode oligonucleotide 1014 poly-T segment is then hybridized to the poly-A tail of the MRNA 1020. Then, in run 1052, the poly-T segment is extended in a reverse transcription reaction that uses the mRNA as a template to produce a 1022 cDNA transcript of the mRNA and also includes each of: sequence segments 1020, 1008, 1012, 1010, 1016 and 1014 of the oligonucleotide code for. Within any given partition, all cDNA transcripts of individual mRNA molecules include a common 1012 barcode sequence segment. However, but including the unique random N-mer sequence, transcripts made from different mRNA molecules within a given partition vary in this unique sequence. As described elsewhere herein, this provides a quantification characteristic that can be identified even after any amplification.
110 back of the contents of a given partition, e.g. For example, the number of unique segments associated with a common barcode can indicate the amount of mRNA that originates from a single partition, and thus a single cell. In run 1054 a second strand is synthesized and in run 1056 the T7 polymerase can utilize the T7 promoter sequence to produce RNA transcripts in in vitro transcription. At run 1058, the transcripts are fragmented (eg, cut), ligated into additional functional sequences, and reverse transcribed. Functional sequences can include a 1030 sequencer-specific flow cell binding sequence, e.g. ex. , a P5 sequence, as well as: a 1028 functional sequence, which may include sequencing primers, e.g. eg, an Rl primer binding sequence, as well as a 1032 functional sequence, which may include a sample index, e.g. ex. , a sample index sequence i5. In step 1060, the RNA transcripts can be reverse transcribed to DNA, the amplified DNA (e.g., by PCR), and can be sequenced to identify the sequence of the cDNA transcript of the mRNA, as well as to sequence the barcode segment and single sequence segment. In some cases, operations 1050 and
111
1052 they can occur on the partition, while the 1054, 1056, 1058, and 1060 operations can occur in bulk (eg, outside the partition). In the case where a partition is a droplet in an emulsion, the emulsion can be broken and the contents of the droplet can be pooled to complete operations 1054, 1056, 1058 and 1060.
In an alternative example of a barcode oligonucleotide for use in RNA analysis (eg, cellular RNA) as shown in Figure 10, functional sequence 1008 can be a P5 sequence and functional sequence 1010 can be a primer binding site R1. In addition, functional sequence 1030 can be a P7 sequence, functional sequence 1028 can be a primer R2 binding site, and functional sequence 1032 can be an i7 sample index sequence.
An additional example of a barcode oligonucleotide for use in RNA analysis, including analysis of messenger RNA (mRNA, including mRNA obtained from a cell) is shown in Figure 11. As shown, general oligonucleotide 1102 is coupled to a bead 1104 via a releasable linkage 1106, for example a disulfide linker. The oligonucleotide can include functional sequences that are used in further processing,
112 for example functional sequence 1108, which may include a sequencer-specific flow cell binding sequence, e.g. eg, a P5 sequence, as well as: a 1110 functional sequence, which may include sequencing primer sequences, e.g. eg, a primer binding site Rl. In some cases, sequence 1108 is a P7 sequence and sequence 1110 is a primer R2 binding site. A barcode sequence 1112 is included within the framework for use in creating barcodes of the sample RNA. An additional sequence segment 1116 may be provided within the oligonucleotide sequence. In some cases, this additional sequence can provide a unique molecular sequence segment, as described elsewhere herein. An additional sequence 1114 can be included to facilitate template exchange, e.g. ex. , polyG. As will be appreciated, although shown as a single oligonucleotide attached to the surface of a bead, individual beads can include tens to hundreds of thousands or even millions of individual oligonucleotide molecules, where, as noted, the code segment The number of bars may be constant or relatively constant for a given bead, but where the segment of
113 Variable or unique sequence varies across an individual bead.
In an exemplary method of cellular mRNA analysis and referring to Figure 11, a cell is codivided along with a bead having barcode, poly-T sequence, and other reagents such as reverse transcriptase, a reducing agent, and dNTP in a partition (eg, a droplet in an emulsion). At step 1150, the cell is lysed while the barcode oligonucleotides are released from the bead (e.g. g., by the action of the reducing agent) and the poly-T sequence hybridizes with the poly-A tail of the 1120 mRNA released from the cell. Then, in step 1152, the poly-T sequence is extended in a reverse transcription reaction that uses the mRNA as a template to produce a cDNA transcript 1122 complementary to the mRNA. The terminal transferase activity of the reverse transcriptase can add additional bases to the transcript of cDNA (eg, polyC). Additional bases added to the cDNA transcript, e.g. g., polyC, can then hybridize to 1114 of the barcode oligonucleotide. This can facilitate template exchange and a sequence complementary to the barcode oligonucleotide can be incorporated into the cDNA transcript. . The
114 Transcripts can be further processed (eg, amplified, removed portions, added additional sequences, etc.) and characterized as described elsewhere herein, e.g. eg, by sequencing. The configuration of the constructs generated through such a method can help to minimize (or avoid) sequencing of the poly-T sequence during sequencing.
A further example of a barcode oligonucleotide for use in RNA analysis, including cellular RNA analysis, is shown in Figure 12A. As shown, the general oligonucleotide 1202 is coupled to a bead 1204 via a free linkage 1206, for example a disulfide linker. The oligonucleotide can include functional sequences that are used in further processing, for example the 1208 functional sequence, which can include a sequencer-specific flow cell binding sequence, e.g. g., a P5 sequence, as well as a 1210 functional sequence, which can include sequencing primer sequences, e.g. g., a primer binding site Rl. In some cases, sequence 1208 is' a P7 sequence and sequence 1210 is a primer R2 binding site. A 1212 barcode sequence is included
115 within the framework for use in creating sample RNA barcodes. An additional sequence segment 1216 may be provided within the oligonucleotide sequence. In some cases, this additional sequence can provide a unique molecular sequence segment, as described elsewhere herein. As will be appreciated, although shown as a single oligonucleotide attached to the surface of a bead, individual beads can include tens to hundreds of thousands or even millions of individual oligonucleotide molecules, where, as noted, the code segment Bar can be constant or relatively constant for a given bead, but where the variable or unique sequence segment varies across an individual bead. In an exemplary method of cellular RNA analysis using this barcode, a cell is codivided along with a barcode-bearing bead and other reagents such as RNA ligase and a reducing agent in one partition (eg, a droplet in an emulsion). The cell is lysed while the barcode oligonucleotides are released (eg, by the action of the reducing agent) from the bead. The barcoded oligonucleotides can then be ligated to the 5 'end of the
116 mRNA transcripts while in RNA ligase partitions. Subsequent steps can include purification (eg, by solid phase reversible immobilization (SPRI)) and further processing (shearing, functional sequence ligation, and subsequent amplification (eg, by PCR)), and these steps can occur in bulk (eg, outside of partition). In the case where a partition is a droplet in an emulsion, the emulsion can be broken and the contents of the droplet can be pooled for further operations.
A further example of a barcode oligonucleotide for use in RNA analysis, including cellular RNA analysis, is shown in Figure 12B. As shown, the general oligonucleotide 1222 is coupled to a bead 1224 via a releasable linkage 1226, for example a disulfide linker. The oligonucleotide can include functional sequences that are used in further processing, for example the 1228 functional sequence, which can include a sequencer-specific flow cell binding sequence, e.g. ex. , a P5 sequence, as well as a 1230 functional sequence, which may include sequencing primer sequences, e.g. e.g., a binding site for
117 primer R1. In some cases, the 1228 sequence is<sup>:</sup> a P7 sequence and the 1230 sequence is a primer R2 binding site. A 1232 barcode sequence is included within the framework for use in barcoding the sample RNA. A 1234 primer sequence (eg, a random primer sequence) can also be included in the oligonucleotide structure, e.g. ex. , a random hexamer. An additional 1236 sequence segment may be provided within the oligonucleotide sequence. In some cases, this additional sequence provides a unique molecular sequence segment, as described elsewhere herein. As will be appreciated, although shown as a single oligonucleotide attached to the surface of a bead, individual beads can include tens to hundreds of thousands or even millions of individual oligonucleotide molecules, where, as noted, the code segment Bar can be constant or relatively constant for a given bead, but where the variable or unique sequence segment varies across an individual bead. In an exemplary method of cellular mRNA analysis using the barcode oligonucleotide of Figure 12B, a cell is codivided along with a bead having barcode and reagents
118 Additives such as reverse transcriptase, a reducing agent, and dNTPs in one partition (eg, a droplet in * an emulsion). The cell is lysed while the barcode oligonucleotides are released from the bead (eg, by the action of the reducing agent). In some cases, sequence 1228 is a P7 sequence and sequence 1230 is a primer R2 binding site. In other cases, sequence 1228 is a P5 sequence and sequence 1230 is a primer binding site R1. The random hexamer 1234 primer sequence can randomly hybridize to cellular mRNA. The random hexamer sequence can then be extended in a reverse transcription reaction that uses the cell's mRNA as a template to produce a cDNA transcript complementary to the mRNA and also includes each of the 1228, 1232, 1230, sequence segments. 1236 and 1234. of the barcode oligonucleotide. Subsequent operations may include purification (eg. , by solid phase reversible immobilization (SPRI)), further processing (shearing, ligation of functional sequences and subsequent amplification (eg, by PCR)), and these operations can occur in bulk (eg, outside partition). In the case where a
119 Partition is a droplet in an emulsion, the emulsion can be broken and the contents of the droplet can be pooled for further operations. Additional reagents that can be co-cleaved along with the barcoded bead can include oligonucleotides to block ribosomal RNA (rRNA) and nucleases to digest the genomic DNA and cDNA of cells. Alternatively, rRNA scavengers can be applied during further processing operations. The configuration of the constructs generated through this method can help to minimize (or avoid) the sequencing of: the poly-T sequence during sequencing. .
The single cell analysis methods described herein may also be useful in whole transcriptome analysis. Again referring to the barcode in Figure 12B, the primer sequence 1234 can be a random N-mer. In some cases, sequence 1228 is a P7 sequence and sequence 1230 is a primer R2 binding site. In other cases, sequence 1228 is a P5 sequence and sequence 1230 is a primer binding site Rl. In an exemplary method of whole transcriptome analysis using this barcode, the individual cell is codivided along with a bead that has
120 barcode, poly-T sequence, and other reagents such as reverse transcriptase, polymerase, a reducing agent, and dNTP in a partition (eg, droplet in an emulsion). In one operation of this method, the cell is lysed while the barcode oligonucleotides are released from the bead (eg, by the action of the reducing agent) and the poly-T sequence is hybridized to the poly-A tail. of cellular mRNA. In a reverse transcription reaction using mRNA as a template, cDNA transcripts of cellular mRNA can be produced. The RNA can then be degraded with an RNase. The 1234 primer sequence in the barcode oligonucleotide can then be randomly hybridized to the cDNA transcripts. The oligonucleotides can be extended using polymerase enzymes and other extension reagents codivided with the bead and similar cell as shown in Figure 3 to generate amplification products (eg, barcoded fragments), similar to the product. example amplification shown in Figure 3 (panel F). Bar-coded nucleic acid fragments, in some cases subjected to additional processing (eg. g., amplification, addition of additional sequences, cleanup processes, etc., as described elsewhere herein), can be
121 characterize, p. eg, through sequence analysis. In this operation, the sequencing signals can come from full-length RNA.
Although operations with various barcode designs have been considered individually, individual beads can include barcode oligonucleotides of various designs for simultaneous use.
In addition to characterizing individual cells or subpopulations of cells from large populations, the processes and systems described herein can also be used to characterize individual cells to provide a general profile of a cell population or other population of organisms. Various applications require the evaluation of the presence and quantification of different types of cells or organisms within a population of cells, including, for example, microbiome analysis and characterization, environmental assessment, food safety assessment, epidemiological analysis, e.g. eg, in tracking contamination or the like. In particular, the assay processes described above can be used to individually characterize, sequence, and / or identify large numbers of individual cells within a population. This characterization can then be used to put together a
122 general profile of the source population, which can provide important diagnostic and prognostic information.
For example, it has been identified that changes in human microbiomes, including, p. For example, intestinal, oral, epidermal microbiomes, etc., are both diagnostic and prognostic of different conditions or general health states. By using the single cell analysis methods and systems described herein, individual cells in a general population can again be characterized, sequenced and identified, and changes within that population can be identified that may indicate factors relevant from the point of view. diagnosis. By way of example, bacterial 16S ribosomal RNA gene sequencing has been used as a highly accurate method for taxonomic classification of bacteria. Utilizing the targeted amplification and sequencing processes described above can provide identification of individual cells within a cell population. The numbers of different cells within a population can be further quantified to identify current states or changes of states over time. See, p. eg, Morgan et al, PLoS Comput. Biol.,
123
Ch. 12, December 2012, 8 (12): el002808, and Ram et al., Syst. Biol. Reprod. Med. 2011 Jun, 57 (3): 162-170, each of which is incorporated herein in its entirety by reference for all purposes. Likewise, the identification and diagnosis of infection or potential infection can also benefit from the single cell assay described herein, e.g. g., to identify microbial species present in large mixtures of other cells or other biological material, cells and / or nucleic acids, including the environments described above, as well as any other environment relevant from a diagnostic point of view, e.g. ex. , cerebrospinal fluid, blood, fecal or intestinal samples, or the like. .
The above analysis may also be particularly useful in characterizing potential drug resistance of different cells, e.g. eg, cancer cells, bacterial pathogens, etc., through analysis of the distribution and profiling of different resistance markers / mutations across cell populations in a given sample. Additionally, characterizing changes in these markers / mutations across cell populations over time can provide valuable information on the progression, alteration, prevention, and treatment of
124 various diseases characterized by such drug resistance problems.
Although described in terms of cells, it will be appreciated that any of various individual biological organisms, or components of organisms, are encompassed within the present disclosure, including, for example, cells, viruses, organelles, cell inclusions, vesicles, or the like. Additionally, when referring to cells,<sup>:</sup> It will be appreciated that such reference includes any cell type, including, and without limitation, prokaryotic cells, eukaryotic cells, bacteria, fungal, plant, mammalian, or other types of animal cells, mycoplasmas, normal tissue cells, tumor cells. , or any other type of cell, whether derived from multicellular organisms or single cell.
Similarly, the analysis of different environmental samples to profile the microbial organisms, viruses or other biological contaminants that are present within these samples, can provide important information on the epidemiology of the disease and, potentially, help prevent disease outbreaks, epidemics and pandemics.
As described above, the methods, systems and
125 Compositions described herein can also be used for the analysis and characterization of other aspects of individual cells or cell populations. In an exemplary process, a sample is provided containing cells that are to be analyzed and characterized for their cell surface proteins. Also provided is a library of antibodies, antibody fragments, or other molecules that have a binding affinity for the cell surface antigens or proteins (or other cell characteristics) for which the cell is to be characterized (also referred to herein as as cell surface feature binding groups). For ease of discussion, these affinity groups are referred to herein as linking groups. The linking groups can include a reporter molecule that indicates the cell surface characteristic to which the linking group binds. In particular, a type of binding group that is specific for one type of cell surface feature will comprise a first reporter molecule, while a type of binding group that is specific for a different cell surface feature will have a different reporter molecule associated with it. . In some aspects, these reporter molecules will comprise sequences of
126 oligonucleotides. Oligonucleotide-based reporter molecules provide the advantages of being able to generate significant diversity in terms of sequence, while they can also easily bind to most biomolecules, e.g. g., antibodies, etc., and furthermore easily detected, eg. eg, using sequencing or array technologies. In the exemplary process, the linking groups include oligonucleotides attached thereto. Therefore, a first type of linking group, e.g. ex. , antibodies to a first type<sub>:</sub> of cell surface characteristic, will have associated a reporter oligonucleotide having a first nucleotide sequence. The different types of linking groups, e.g. .ej. , Antibodies that have binding affinity for other different cell surface characteristics will have associated with them reporter oligonucleotides comprising different nucleotide sequences, e.g. ex. , which have a partially or completely different nucleotide sequence. In some cases, for each type of cell surface feature binding group, e.g. eg, the antibody or the antibody fragment, the reporter oligonucleotide sequence can be known and readily identifiable as associated with the binding group<sup>r</sup> of
127 known cell surface characteristic. These oligonucleotides may be directly coupled to the linking group or they may be attached to a bead. molecular network, e.g. eg, a linear, globular, cross-linked or other polymer or framework that is linked or otherwise associated with the linking group, allowing the attachment of multiple reporter oligonucleotides with a single linking group.<sub>:</sub>
In the case of multiple reporter molecules coupled to a single binding group, these reporter molecules may comprise the same sequence or a particular binding group will include a known set of reporter oligonucleotide sequences. As between different linking groups, e.g. For example, specific for different cell surface characteristics, the reporter molecules may be different and attributable to the particular binding group.<sub>:</sub>
The linking of the reporter groups to the linking groups can be achieved by any of a number of direct or indirect, covalent or non-covalent associations or linkages.
For example, in the case of reporter groups of oligonucleotides associated with antibody-based binding groups, these oligonucleotides can be covalently linked to a part of an antibody or fragment of
128 antibody using chemical conjugation techniques (e.g.
g., Lightning-Link® antibody labeling kits supplied by Innova Biosciences), as well as other non-covalent binding mechanisms, e.g. eg, using biotinylated antibodies and oligonucleotides (or beads: including one or more biotinylated linkers, coupled with oligonucleotides) with an avidin or streptavidin linker. Antibody and oligonucleotide biotinylation techniques are available (See, eg, Fang et al.
Fluoride-Cleavable Biotinylation Phosphoramidite for 5'-endLabeling and Affinity Purification of Synthetic Oligonucleotides, Nucleic Acids Res., 2003 Jan 15; 31 (2): 708-715, DNA 3 'End Biotinylation Kit, available from Thermo Scientific, which are incorporated herein in their entirety by this reference for all purposes). Also, protein and peptide biotinylation techniques have been developed and are readily available (see, eg. , the US patent. No. 6,265,552, which is incorporated herein in its entirety by this reference for all purposes).
Reporter oligonucleotides having any of a range of different lengths can be provided, depending on the diversity of the reporter molecules desired or
129 a determined analysis, the sequence detection scheme employed, and the like. In some cases, these reporter sequences may be more than about 5 nucleotides in length, more than about 10 nucleotides in length, more than about 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, or even 200 nucleotides in length '. In some cases, these reporter nucleotides may be less than about 250 nucleotides in length, less than about 200, 180, 150, 120, 100, 90, 80, 70, 60, 50, 40, or even 30 nucleotides in length. In many cases, the reporter oligonucleotides can be selected to provide barcoded products that are already measured and that are otherwise configured for analysis in a sequencing system. For example, these sequences can be provided at a length that ideally creates sequenced products of a desired length for particular sequencing systems. Also, these reporter oligonucleotides can include additional sequence elements, in addition to the reporter sequence, such as sequencer binding sequences, sequencing primer sequences, amplification primer sequences, or complements to any of these.
In operation, a sample containing cells is incubated
130 with the binding molecules and their associated reporter oligonucleotides, for any of the cell surface characteristics to be analyzed. After incubation, cells are washed to remove unbound binding groups. After washing, cells are divided into separate partitions, e.g. ex. , droplets, together with the barcode-containing beads described above, wherein each partition includes a limited number of cells, e.g. eg, in some cases, a single cell. Upon release of the barcodes from the beads, these will prime the amplification and barcode creation of the reporter oligonucleotides. As indicated above, the barcode replicas of the reporter molecules may additionally include functional sequences, such as primer sequences, binding sequences, or the like.
The barcoded reporter oligonucleotides are subjected to sequence analysis to identify which reporter oligonucleotides bind to the cells within the partitions. Furthermore, by also sequencing the associated barcode sequence, it can be identified that a given cell surface feature likely comes from the same cell as another,
131 different cell surface features, the reporter sequences of which include the same barcode sequence, that is, they are derived from the same partition.
Based on the reporter molecules arising from a single partition based on the presence of the barcode sequence, a cell surface profile of individual cells from a population of cells can be created. The profiles of individual cells or cell populations can be compared to the profiles of other cells, e.g. g., normal cells, to identify variations in cell surface characteristics, which can provide information that is relevant from a diagnostic point of view. In particular, these profiles can be particularly useful for the diagnosis of multiple disorders that are characterized by variations in cell surface receptors, such as cancer and other disorders.
IV. Devices and systems
Also provided herein are microfluidic devices used to divide cells as described above. These microfluidic devices can comprise networks of channels to carry out the division process such as those presented in Figures 1 and 2. Examples
132 Particularly useful microfluidic devices are described in US Provisional Patent Application No. 61 / 977,804, filed April 4, 2014, and incorporated herein by this reference in its entirety for all purposes. Briefly, these microfluidic devices may comprise channel networks, such as those described herein, to divide cells into separate partitions and codivide said cells with members of oligonucleotide barcode libraries, e.g. ex. , arranged in pearls. These channel networks can be arranged within a solid body, e.g. eg, a glass, semiconductor, or polymer body structure where the channels are defined, where these channels communicate at their ends with reservoirs to receive the various input fluids and for the final deposition of the divided cells, etc. , from the output of the channel networks. By way of example and with reference to FIG. 2, a reservoir fluidly attached to a channel 202 may be provided with an aqueous suspension of cells 214, while a reservoir coupled to a channel 204 may be provided with an aqueous suspension of cells. beads 216 that carry the oligonucleotides. Channel segments 206 and 208 can be provided with a solution not
133 aqueous, e.g. g., an oil, where the aqueous fluids divide as droplets at the junction of channels 212. Lastly, an outlet reservoir can be fluidly coupled with channel 210 where divided cells and beads can be supply and from where they can be collected. As will be appreciated, although described as reservoirs, it will be appreciated that the channel segments may be coupled with any of a number of different fluid sources or receiving components, including pipes, manifolds, or fluid components of other systems.
Systems are also provided that control the flow of these fluids through the channel networks, e.g. ex. , by applied pressure differentials, centrifugal force, electrokinetic pumping, gravity or capillary flow or similar.
V. Kits
Kits for analyzing single cells or small populations of cells are also provided herein. Kits can include one, two, three, four, five or more, up to all of the cleavage fluids, including aqueous buffers and non-aqueous cleavage fluids or oils, nucleic acid barcode libraries that are closely associated removable with pearls, according to
134 Described herein are microfluidic devices, reagents for disrupting cells by amplifying nucleic acids and providing additional functional sequences on cellular nucleic acid fragments or replicas thereof, as well as instructions for using any of the above in the methods described herein. .
SAW. Computer control systems<sup>:</sup>
The present description provides computer control systems that are programmed to implement the methods of the description. Figure 17 shows a 1701 computer system that is programmed or otherwise configured to implement the methods of the disclosure including nucleic acid sequencing methods, interpretation of nucleic acid sequencing data, and analysis of cellular nucleic acids, such as RNA (p. g., mRNA), and characterization of cells from sequencing data. Computer system 1701 can be a user's electronic device or a computer system that is remotely located from the electronic device. The electronic device can be a mobile electronic device.
The 1701 computer system includes a central processing unit (CPU, also processor and
135 computer processor herein) 1705, which may be a single or multi-core processor or multiple processors for parallel processing. The 1701 computer system also includes memory or memory location 1710 (eg, random access memory, read-only memory, flash memory), electronic storage unit 1715 (eg, hard drive), communication interface 1720 (p. g., network adapter) to communicate with one or more different systems and 1725 peripheral devices, such as cache, other memory, data storage, and / or electronic display adapters. The 1710 memory, 1715 storage unit, 1720 interface, and: 1725 peripheral devices communicate with the 1705 CPU through a communication bus (solid lines), such as a motherboard. Storage unit 1715 may be a data storage unit (or data repository) for storing data. The computer system 1701 can be operatively coupled to a computer network (network) 1730 with the aid of the communication interface 1720. The network 1730 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is communicated with the Internet. The 1730 network, in some cases, is a telecommunications and / or data network. The 1730 network can include one or more servers
136 IT, which can enable distributed computing, such as cloud computing. The 1730 network, in some cases, with the collaboration of the 1701 computer system, may implement a peer-to-peer network, which may allow. the devices attached to the computer system 1701. behave as a client or a server.
The CPU 1705 can execute a sequence of machine-readable instructions that can be incorporated into a program or software. Instructions can be stored in. a memory location, such as memory 1710. Instructions can be directed to the CPU 1705, which can later program or otherwise configure the CPU 1705 to implement the methods of the present disclosure. Examples of operations performed by the CPU 1705 may include retrieve, decode, execute, and rewrite.
The CPU 1705 can be part of a circuit, like an integrated circuit. One or more other components of the 1701 system may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
Storage unit 1715 can store files, such as drivers, libraries, and saved programs. Storage unit 1715 can store data: user, eg. e.g., user preferences and software
137 user. The computer system 1701, in some cases, may include one or more additional data storage units that are external to the computer system 1701: for example, located on a remote server that is communicated with the computer system, 1701 through a intranet or Internet.
Computer system 1701 may communicate with one or more remote computer systems over network 1730. · For example, computer system 1701 may communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (eg, laptop PC), whiteboard or tablet PC (eg, Apple® iPad, Samsung® Galaxy Tab), phones, smartphones (eg. , Apple® iPhone, Android devices, Blackberry®) or personal digital assistants. The user can access the 1701 computer system through the 1730 network.
The methods described herein may be implemented by machine-executable code (eg, computer processor) stored in an electronic storage location of computer system 1701, such as memory 1710 or storage unit. Electronic 1715. The executable or machine-readable code can be
138 provide in the form of software. During use, code can be executed by processor 1705. In some cases, code can be retrieved from storage unit 1715 and stored in memory 1710 for quick access: by processor 1705. In some situations, the Electronic storage unit 1715 can be bypassed and machine-executable instructions are stored in memory 1710.
The code can be precompiled and configured for use with a machine that has a processor adapted to run the code, or it can be compiled at run time. The code can be supplied in a programming language that can be selected to allow the code to run precompiled or as compiled.
Some aspects of the systems and methods provided herein, such as the 1701 computer system, can be incorporated into programming. Various aspects of technology can be considered as products or articles of manufacture typically in the form of machine (or processor) executable code and / or associated data that is transported or incorporated into a type of machine compatible medium. Machine-executable code can be stored in an electronic storage unit, such as
139 a memory (eg, read-only memory, random access memory, flash memory) or a hard disk. Storage-type media can include either or. the entirety of. the tangible memory of computers, processors or the like, or modules associated with them, for example, various semiconductor memories, tape drives, disk drives, and the like, which can provide non-transient storage at any time for software programming. All or part of the software may, at times, communicate over the Internet or other telecommunications networks. These communications, for example, may allow the uploading of software from one computer or processor to another, for example, from a management server or host computer to the computing platform of an application server. Therefore, another type of medium that may contain the software elements includes optical, electrical, and electromagnetic waves, such as those used between physical interfaces between local devices, across wired and optical ground networks, and between various air links. The physical elements that transport these waves, such as wired or wireless links, optical links or the like, can also be considered as means that
140 contain the software. As used herein, unless limited to tangible, non-transitory storage media, terms such as machine-readable or computer-readable media refer to any media that is involved in supplying instructions to a processor for execution. .
Therefore, a machine-readable medium, such as computer executable code, can take many forms, including but not limited to a readable storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media includes, for example, optical or magnetic discs, like any of the storage devices in any computer or the like, like those that can be used to implement the databases, etc., shown in the drawings. The volatile storage media includes dynamic memory, as the main memory of said computing platform. Tangible transmission media include coaxial cables; copper and fiber optic cable, which includes cables that make up a bus within a computer system. Carrier wave transmission media can take the form of electromagnetic or electrical signals or acoustic or light waves, such as those generated during communications from
141 infrared (IR) and radio frequency (RF) data. Therefore, common forms of computer-readable media include, for example: a floppy disk, floppy disk, hard disk, magnetic tape, any other magnetic media, a CDROM, DVD or DVD-ROM, any other optical media, punched cards, paper tape, any other physical storage media with hole patterns, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave that carries data or instructions, cables or links that carry said carrier wave, or any other means from which a computer can read programming code and / or data. Many of these forms of computer-readable media can be involved in moving one or more sequences of one or more instructions to a processor for execution.
Computer system 1701 may include or be communicated with an electronic display 1735 comprising a user interface (UI) 1740 to provide, for example, nucleic acid sequencing results, analysis of nucleic acid sequencing data, characterization of samples of nucleic acid sequencing, cell characterizations, etc. Some UI examples
142 They include, but are not limited to, a graphical user interface (GUI) and a web-based user interface.
The methods and systems of the present description can be implemented through one or more algorithms. An algorithm can be implemented through software upon execution by central processing unit 1705. The algorithm can, for example, initiate nucleic acid sequencing, process nucleic acid sequencing data, interpret nucleic acid sequencing results, characterize nucleic acid samples, characterize cells, etc.
VII. Examples
Example I Analysis of cellular RNA using emulsions
In one example, template-switched reverse transcription and cDNA amplification (by PCR) are performed in emulsion droplets with operations as shown in Figure 9A. The reaction mix that is split for reverse transcription and cDNA amplification (by PCR) includes 1,000 cells or 10,000 cells or 10 ng RNA, beads containing barcode oligonucleotides / 0.2% Tx100 / 5x Kapa buffer, 2x Kapa HS HiFi Ready Mix, 4 μΜ switch oligonucleotide and Smartscribe. When cells are present, the mixture is divided so that most
143 or all of the droplets comprise a single cell or a single bead. Cells are lysed while releasing the barcode oligonucleotides from the bead and the poly-T segment of the barcode oligonucleotide hybridizes to the poly-A tail of mRNA that is released from the cell as in step 950. The poly-T segment is extended in a reverse transcription reaction as in step 952 and the cDNA transcription is amplified as in step 954. Thermal cycling conditions are 42 ° C for 130 minutes; 98 ° C for 2 min; and 35 cycles of the following 98 ° C for 15 sec, 60 ° C for 20 sec and 72 ° C for 6 min. After thermal cycling, the emulsion is broken and the transcripts are purified with Dynabeads and 0.6x SPRI as in run 956.
The performance of template swap reverse transcription and PCR in the emulsions is shown for 1,000 cells in Figure 13A and 10,000 cells in Figure 13C and 10 ng of RNA in Figure 13B (Smartscribe line). RT and PCR cDNA transcripts made in emulsions for 10 ng of RNA are cut and ligated into functional sequences, cleaned with 0.8x SPRI and further amplified by PCR as in step 958. The amplification product is cleaned with 0.8x SPRI. He
144 Performance of this processing is shown in Figure 13B (line SSII).
Example II Analysis of cellular RNA using emulsions
In another example, template-switched reverse transcription and cDNA amplification (by PCR) are performed in emulsion droplets with operations as shown in Figure 9A. The reaction mix that is split for reverse transcription and cDNA amplification (by PCR) includes Jurkat cells, beads containing barcode oligonucleotides / 0.2% TritonX-100 / 5x Kapa buffer, 2x Kapa HS HiFi Ready Mix, 4 μΜ switch oligonucleotide and Smartscribe. The mixture is divided so that most or all of the droplets comprise a single cell or a single bead. Cells are lysed while releasing the barcode oligonucleotides from the bead and the poly-T segment of the barcode oligonucleotide hybridizes to the poly-A tail of mRNA that is released from the cell as in step 950. The poly-T segment is extended in a reverse transcription reaction as in run 952 and the cDNA transcription is amplified as in run 954. Thermal cycling conditions <sub>;</sub> it is 42 ° C for 130 minutes; 98 ° C for 2 min; and 35 cycles of the next 98 ° C for 15 sec, 60 ° C for 20 sec
145 and 72 ° C for 6 min. After thermal cycling, the emulsion is broken and the transcripts are cleaned with Dynabeads and 0.6x SPRI as in run 956. The yield of reactions with various numbers of cells (625 cells, 5 1,250 cells, 2,500 cells, 5,000 cells and 10,000 cells) is shown in Figure 14A. These yields are confirmed by the results of the GADPH qPCR assays shown in Figure 14B. :
Example III RNA analysis using emulsions.
In another example, reverse transcription is performed in emulsion droplets and bulk cDNA amplification is performed in a manner similar to that shown in Figure 9C. The reaction mix that is split for reverse transcription includes beads containing barcode oligonucleotides, 10 ng of Jurkat RNA (eg, Jurkat mRNA), 5x First-Strand buffer, and Smartscribe. The bead barcode oligonucleotides are released and the poly-T segment of the barcode oligonucleotide is hybridized to the poly-A tail of the RNA as in step '961.
The poly-T segment is extended in a reaction. reverse transcription as in run 963. Thermal cycling conditions for reverse transcription are one cycle at 42 ° C for 2 hours and one cycle at 70 ° C
146 for 10 min. After thermal cycling, the emulsion is broken and the RNA and cDNA transcripts are denatured as in run 962. A second strand is synthesized by primer extension with a primer having a biotin tag as in run 964. Conditions Reactions for this primer extension include the transcription of cDNA as the first strand and the biotinylated extension primer with a concentration ranging from 0.5-3.0 µΜ. Thermal cycling conditions are a cycle at 98 ° C for 3 min and a cycle of 98 ° C for 15 sec, 60 ° C for 20 sec and 72 ° C for 30 min. After primer extension, the second strand is decanted with streptavidin Dynabeads MyOne C1 and TI and cleaned: with Agilent SureSelect XT buffers. The second strand is preamplified by PCR as in run 965 with the following cycling conditions - one cycle at 98 ° C for 3 min and one cycle at 98 ° C for 15 sec, 60 ° C for 20 sec and 72 ° C for 30 min. The performance for various concentrations of biotinylated primer (0.5 μΜ, 1.0 μΜ, 2.0 μΜ, and 3.0 μΜ) is shown in Figure 15.
Example IV RNA analysis using emulsions
In another example, in vitro transcription by T7 polymerase is used to produce RNA transcripts.
147 as shown in Figure 10. The mixture that is divided for reverse transcription includes beads containing barcode oligonucleotides that also include a promoter sequence of T7 RNA polymerase, 10 ng of human RNA (eg, human mRNA ), 5x First-Strand and Smartscribe shock. The mixture is divided so that most or all of the droplets comprise a single bead. The barcode oligonucleotides are released from the bead and the poly-T segment of the barcode oligonucleotide hybridizes to the poly-A tail of the RNA as in step 1050. The poly-T segment is extended in a reaction of reverse transcription as in run 1052. Thermal cycling conditions are one cycle at 42 ° C for 2 hours and one cycle at 70 ° C for 10 min. After thermal cycling, the emulsion is broken and the remaining operations are carried out in bulk. A second strand is synthesized per primer extension as in step 1054. Reaction conditions for this primer extension include transcription of cDNA as template and extension primer. Thermal cycling conditions are a cycle at 98 ° C for 3 min and a cycle of 98 ° C for 15 sec, 60 ° C for 20 sec and 72 ° C for 30 min. After this primer extension, the second strand is purified with 0.6x
148
SPRI. As in run 1056, in vitro transcription is performed to produce RNA transcripts. In vitro transcription is performed overnight and the transcripts are purified with 0.6x SPRI. In vitro transcriptional RNA yields are shown in Figure 16. While some embodiments of the invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The invention is not intended to be limited by the specific examples provided within the specification. While the invention has been described with reference to the above-mentioned specification, the descriptions and illustrations of the embodiments herein are not to be construed in a limiting sense. It will be apparent to those skilled in the art that various variations, changes, and substitutions can be made without departing from the spirit of the invention. Furthermore, it will be understood that all aspects of the invention are not limited to the specific descriptions, configurations, or relative proportions presented herein that depend on a number of conditions and variables. It will be understood that, in the practice of the invention, various alternatives to the embodiments of the invention may be employed.
149 that were described herein. Therefore, it is considered that the invention should also cover such alternatives, modifications, variations or equivalents. The following claims are intended to define the scope of the invention and that the methods and structures within the scope of these claims and their equivalents be encompassed by these claims.
150
Contents7
21 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21
282 members in 13 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 62017558 | United States of America | – | |
| 201462017558 | United States of America | P | |
| 62061567 | United States of America | – | |
| 201462061567 | United States of America | P | |
| 2015038178 | United States of America | W |
Members282
| Document | Office | Kind | |
|---|---|---|---|
| CA2881685A1 | Canada | A1 | |
| CA3216609A1 | Canada | A1 | |
| WO2014028537A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2014155295A1 | United States of America | A1 | |
| CA2894694A1 | Canada | A1 | |
| WO2014093676A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2014206554A1 | United States of America | A1 | |
| CA2900481A1 | Canada | A1 | |
| CA2900543A1 | Canada | A1 | |
| US2014227684A1 | United States of America | A1 | |
| US2014228255A1 | United States of America | A1 | |
| WO2014124336A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2014124338A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2014235506A1 | United States of America | A1 | |
| US2014287963A1 | United States of America | A1 | |
| WO2014124336A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2014378322A1 | United States of America | A1 | |
| US2014378345A1 | United States of America | A1 | |
| US2014378349A1 | United States of America | A1 | |
| US2014378350A1 | United States of America | A1 | |
| CA2915499A1 | Canada | A1 | |
| WO2014210353A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2015005199A1 | United States of America | A1 | |
| US2015005200A1 | United States of America | A1 | |
| AU2013302756A1 | Australia | A1 | |
| IL237156A0 | Israel | A0 | |
| IL237156D0 | Israel | D0 | |
| KR20150048158A | Republic of Korea | A | |
| EP2885418A1 | European Patent Office (EPO) | A1 | |
| IN1126DEN2015A | India | A | |
| AU2013359165A1 | Australia | A1 | |
| CN104769127A | China | A | |
| WO2014210353A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2015218633A1 | United States of America | A1 | |
| AU2014214682A1 | Australia | A1 | |
| US2015224466A1 | United States of America | A1 | |
| US2015225777A1 | United States of America | A1 | |
| US2015225778A1 | United States of America | A1 | |
| MX2015001939A | Mexico | A | |
| JP2015528283A | Japan | A | |
| EP2931919A1 | European Patent Office (EPO) | A1 | |
| KR20150119047A | Republic of Korea | A | |
| CN105102697A | China | A | |
| EP2954065A2 | European Patent Office (EPO) | A2 | |
| EP2954104A1 | European Patent Office (EPO) | A1 | |
| AU2014302277A1 | Australia | A1 | |
| CA2953374A1 | Canada | A1 | |
| WO2015200893A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2015376609A1 | United States of America | A1 | |
| EP2885418A4 | European Patent Office (EPO) | A4 | |
| WO2015200893A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20160032723A | Republic of Korea | A | |
| CN105492607A | China | A | |
| JP2016511243A | Japan | A | |
| WO2015200893A9 | World Intellectual Property Organization (WIPO) | A9 | |
| EP3013957A2 | European Patent Office (EPO) | A2 | |
| US9388465B2 | United States of America | B2 | |
| MX2015016968A | Mexico | A | |
| US9410201B2 | United States of America | B2 | |
| EP2931919A4 | European Patent Office (EPO) | A4 | |
| EP2954065A4 | European Patent Office (EPO) | A4 | |
| EP2954104A4 | European Patent Office (EPO) | A4 | |
| EP3013957A4 | European Patent Office (EPO) | A4 | |
| JP2016526889A | Japan | A | |
| US2016304860A1 | United States of America | A1 | |
| AU2015279548A1 | Australia | A1 | |
| US9567631B2 | United States of America | B2 | |
| KR20170020704A | Republic of Korea | A | |
| IL249617A0 | Israel | A0 | |
| IL249617D0 | Israel | D0 | |
| MX2016016902AThis record | Mexico | A | |
| US2017114390A1 | United States of America | A1 | |
| EP3161160A2 | European Patent Office (EPO) | A2 | |
| US9644204B2 | United States of America | B2 | |
| CN106795553A | China | A | |
| US9689024B2 | United States of America | B2 | |
| US9695468B2 | United States of America | B2 | |
| US9701998B2 | United States of America | B2 | |
| BR112015019159A2 | Brazil | A2 | |
| BR112015032512A2 | Brazil | A2 | |
| BR112015003354A2 | Brazil | A2 | |
| JP2017522867A | Japan | A | |
| US2017247757A1 | United States of America | A1 | |
| US2017321252A1 | United States of America | A1 | |
| US2017335385A1 | United States of America | A1 | |
| US2017342404A1 | United States of America | A1 | |
| US2017356027A1 | United States of America | A1 | |
| US2017362587A1 | United States of America | A1 | |
| US9856530B2 | United States of America | B2 | |
| EP3161160A4 | European Patent Office (EPO) | A4 | |
| BR112015003354A8 | Brazil | A8 | |
| US2018016634A1 | United States of America | A1 | |
| US2018030512A1 | United States of America | A1 | |
| AU2013302756B2 | Australia | B2 | |
| US2018051321A1 | United States of America | A1 | |
| US2018094298A1 | United States of America | A1 | |
| US2018094312A1 | United States of America | A1 | |
| US2018094313A1 | United States of America | A1 | |
| US2018094314A1 | United States of America | A1 | |
| US2018094315A1 | United States of America | A1 |
Numbers
- Publication
- 2016016902
- Application
- 16902
Titles2
- Spanish
- METODOS PARA ANALIZAR ACIDOS NUCLEICOS DE CELULAS INDIVIDUALES O POBLACIONES DE CELULAS.
- English
- METHODS FOR ANALYZING NUCLEIC ACIDS FROM INDIVIDUAL CELLS OR POPULATIONS OF CELLS.
Classification
- CPC, 16
- C12N15/1065
- C12Q1/6816
- C40B20/04
- C40B50/16
- C12Q1/6806
- C12Q1/6874
- C12Q2563/179
- C12Q2563/149
- C12Q2563/159
- C12Q2535/122
- C12Q2565/629
- C12Q1/6804
- C12Q1/683
- C12Q2525/191
- C12Q2537/143
- C12Q2537/149
- IPC, 3
- C12N15 10
- C12P19 34
- C12Q1 68