Computer graphical user interface with genomic workflow.
Abstract
<?abstract ?><p num="0000">Methods and computer devices disclosed to process genomic data in at least partially automated workflows modules. A method comprising: specifying a source of one or more nucleotide sequences may be obtained; Selecting one or more modules for processing data, including at least one module for processing the one or more nucleotide sequences; Displaying graphical components in a graphical user interface, the and or represent the source modules as nodes in a workspace; Receiving input via the graphical user interface that the source and or arrange the modules as a workflow that includes a series of nodes, wherein the sequence for each particular module indicating that an output from either the source or another specific module in the particular module must be entered; Generating an output for the workflow, based on the one or more nucleotide sequences, by each module is executed in an order indicated by the sequence.</p><p><img file="DE102014103482A1_0001.tif" he="165" img-content="drawing" img-format="tif" inline="no" orientation="portrait" wi="109" /></p>

Term
7.5 yearsto projected expiry
Projected expiry 14 March 2034, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
20 claims: 10 independent, 10 dependent
- 1Verfahren, das Folgendes umfasst:Empfangen einer ersten Eingabe, die eine Quelle angibt, von der eine oder mehrere Nukleotidsequenzen erhalten werden sollen;Empfangen einer oder mehrerer zweiter Eingaben, die ein oder mehrere Module zum Verarbeiten von Daten auswählen, einschließlich mindestens ein Modul, um die eine oder mehreren Nukleotidsequenzen zu verarbeiten;Darstellen graphischer Komponenten, die die Quelle und das eine oder die mehreren Module als Knoten in einem Arbeitsbereich repräsentieren, in einer graphischen Benutzeroberfläche;Empfangen einer oder mehrerer dritter Eingaben über die graphische Benutzeroberfläche, die die Quelle und das eine oder die mehreren Module als ein Arbeitsablauf anordnen, der eine Abfolge von Knoten aufweist, wobei die Abfolge für jedes bestimmte Modul der ausgewählten Module anzeigt, dass die Ausgabe von entweder der Quelle oder einem anderen bestimmten Modul in das bestimmte Modul eingegeben werden soll;Erzeugen einer Ausgabe für den Arbeitsablauf, gestützt auf die eine oder mehreren Nukleotidsequenzen, indem jedes Modul des einen oder der mehreren Module in einer Reihenfolge verarbeitet wird, die durch die Abfolge angegeben wird;wobei das Verfahren durch eine oder mehrere Rechenvorrichtungen ausgeführt wird.
- 2Verfahren nach Anspruch 1, wobei jedes Modul des einen oder der mehreren Module eine Ausgabe erzeugt, die einer Ontologie entspricht, die Datenstrukturen definiert, die Genomdaten repräsentieren, wobei die Datenstrukturen zumindest Sequenzen, Proteinobjekte, Alignment-Objekte, Annotationen und Veröffentlichungen umfassen.
- 3Verfahren nach Anspruch 1 oder 2, das weiter Folgendes umfasst:Erzeugen eines Datenknotens aus der Ausgabe, wobei der Datenknoten Elemente von Genomdaten umfasst, wobei der Datenknoten mit einem letzten Modul in der Abfolge verbunden ist;Empfangen einer vierten Eingabe, die ein Element der Genomdaten zu dem Datenknoten hinzufügt oder von ihm entfernt, über die graphische Benutzeroberfläche;Empfangen einer fünften Eingabe, die ein bestimmtes Modul auswählt, um den Datenknoten zu verarbeiten, über die graphische Benutzeroberfläche;Hinzufügen des bestimmten Moduls zu dem Ende der Abfolge;Erzeugen einer zweiten Ausgabe für den Arbeitsablauf, gestützt auf die eine oder mehreren Nukleotidsequenzen, indem jedes Modul der Abfolge, einschließlich des bestimmten Moduls, in der Reihenfolge verarbeitet wird, die durch die Abfolge angezeigt ist.
- 4Verfahren nach einem der vorangegangenen Ansprüche, wobei das eine oder die mehreren Module eine Mehrzahl von Modulen umfassen, wobei das Erzeugen der Ausgabe für den Arbeitsablauf das Verwenden der Ausgabe von der Quelle als Eingabe für ein erstes Modul und das Verwenden der Ausgabe von dem ersten Modul als Eingabe für ein zweites Modul umfasst.
- 5Verfahren nach einem der vorangegangenen Ansprüche, wobei das mindestens eine Modul so konfiguriert ist, dass es die eine oder mehreren Nukleotidsequenzen verarbeitet, indem es mit einem externen Webserver und/oder einem externen Datenbankserver kommuniziert.
- 6Verfahren nach einem der vorangegangenen Ansprüche, das weiter Folgendes umfasst:Speichern von Arbeitsablauf-Daten, die die Abfolge beschreiben;Veranlassen, dass die Arbeitsablauf-Daten mit mehreren Nutzern geteilt werden;nachfolgendes Rekonstruieren der Abfolge in einer zweiten graphischen Benutzeroberfläche, gestützt auf die Arbeitsablauf-Daten;Empfangen einer vierten Eingabe über die graphische Benutzeroberfläche, die die Abfolge so modifiziert, dass sie ein oder mehrere zusätzliche Module umfasst;Erzeugen einer zweiten Ausgabe, gestützt auf die eine oder mehreren Nukleotidsequenzen, indem jedes Modul in der Abfolge, einschließlich des einen oder der mehreren zusätzlichen Module, in einer Reihenfolge verarbeitet wird, die durch die Abfolge angegeben ist.
- 7Verfahren nach einem der vorangegangenen Ansprüche, wobei das eine oder die mehreren Module ein erstes Modul umfassen, das eine erste Ausgabe gestützt auf die Quelle erzeugt, und ein zweites Modul, das die erste Ausgabe mit einer zweiten Ausgabe von einem dritten Modul mischt, das nicht in der Abfolge ist, wobei die Quelle, das erste Modul, das zweite Modul und das dritte Modul alle Knoten in einem Arbeitsablauf sind.
- 8Verfahren nach einem der vorangegangenen Ansprüche, das weiter das Darstellen von Bedienelementen zum Auswählen des einen oder der mehreren Module umfasst, wobei die Bedienelemente zumindest Folgendes umfassen:ein erstes Bedienelement zum Auswählen eines ersten Moduls, das nach Veröffentlichungen in einer Online-Datenbank sucht, gestützt auf Genomdaten, ein zweites Bedienelement zum Auswählen eines zweiten Moduls, das ein Sequenz-Alignment für mehrere Sequenzen ausgibt, und ein drittes Bedienelement zum Auswählen eines dritten Moduls, das Proteinfamilien für eine Nukleotidsequenz identifiziert.
- 9Verfahren nach einem der vorangegangenen Ansprüche, wobei das Empfangen der einen oder mehreren dritten Eingaben das Darstellen von visuellem Feedback umfasst, während ein erster Knoten ausgewählt wird, der anzeigt, dass Genomdaten, die von dem ersten Knoten ausgegeben wurden, als Eingabe für einen zweiten Knoten verknüpft werden können.
- 10Verfahren nach einem der vorangegangenen Ansprüche, wobei das eine oder die mehreren Module mindestens zwei Module umfassen, wobei das Verarbeiten jedes der Module des einen oder der mehreren Module in einer Reihenfolge, die durch die Abfolge angezeigt ist, das automatische Verarbeiten jedes der Module umfasst, ohne menschliches Eingreifen zwischen dem Anfang des Verarbeitens eines ersten Moduls in der Abfolge und dem Erzeugen der Ausgabe, indem das Verarbeiten eines letzten Moduls in der Abfolge abgeschlossen wird.
- 11Ein oder mehrere computerlesbare Medien, die Befehle speichern, die, wenn sie durch eine oder mehrere Rechenvorrichtungen ausgeführt werden, folgende Auswirkungen haben:Empfangen einer ersten Eingabe, die eine Quelle angibt, von der eine oder mehrere Nukleotidsequenzen erhalten werden sollen;Empfangen einer oder mehrerer zweiter Eingaben, die ein oder mehrere Module zum Verarbeiten von Daten auswählen, einschließlich mindestens ein Modul, um die eine oder mehreren Nukleotidsequenzen zu verarbeiten;Darstellen graphischer Komponenten, die die Quelle und das eine oder die mehreren Module als Knoten in einem Arbeitsbereich repräsentieren, in einer graphischen Benutzeroberfläche;Empfangen einer oder mehrerer dritter Eingaben über die graphische Benutzeroberfläche, die die Quelle und das eine oder die mehreren Module als ein Arbeitsablauf anordnen, der eine Abfolge von Knoten aufweist, wobei die Abfolge für jedes bestimmte Modul der ausgewählten Module anzeigt, dass die Ausgabe von entweder der Quelle oder einem anderen bestimmten Modul in das bestimmte Modul eingegeben werden soll;Erzeugen einer Ausgabe für den Arbeitsablauf, gestützt auf die eine oder mehreren Nukleotidsequenzen, indem jedes Modul des einen oder der mehreren Module in einer Reihenfolge verarbeitet wird, die durch die Abfolge angegeben wird.
- 12Das eine oder die mehreren computerlesbaren Medien nach Anspruch 11, wobei jedes Modul des einen oder der mehreren Module eine Ausgabe erzeugt, die einer Ontologie entspricht, die Datenstrukturen definiert, die Genomdaten repräsentieren, wobei die Datenstrukturen zumindest Sequenzen, Proteinobjekte, Alignment-Objekte, Annotationen und Veröffentlichungen umfassen.
- 13Das eine oder die mehreren computerlesbaren Medien nach Anspruch 11 oder 12, wobei die Befehle, wenn sie von der einen oder den mehreren Rechenvorrichtungen ausgeführt werden, weiter folgende Auswirkungen haben:Erzeugen eines Datenknotens aus der Ausgabe, wobei der Datenknoten Elemente von Genomdaten umfasst, wobei der Datenknoten mit einem letzten Modul in der Abfolge verbunden ist;Empfangen einer vierten Eingabe, die ein Element der Genomdaten zu dem Datenknoten hinzufügt oder von ihm entfernt, über die graphische Benutzeroberfläche;Empfangen einer fünften Eingabe, die ein bestimmtes Modul auswählt, um den Datenknoten zu verarbeiten, über die graphische Benutzeroberfläche;Hinzufügen des bestimmten Moduls zu dem Ende der Abfolge;Erzeugen einer zweiten Ausgabe für den Arbeitsablauf, gestützt auf die eine oder mehreren Nukleotidsequenzen, indem jedes Modul der Abfolge, einschließlich des bestimmten Moduls, in der Reihenfolge verarbeitet wird, die durch die Abfolge angezeigt ist.
- 14Das eine oder die mehreren computerlesbaren Medien nach einem der Ansprüche 11 bis 13, wobei das eine oder die mehreren Module eine Mehrzahl von Modulen umfassen, wobei das Erzeugen der Ausgabe für den Arbeitsablauf das Verwenden der Ausgabe von der Quelle als Eingabe für ein erstes Modul und das Verwenden der Ausgabe von dem ersten Modul als Eingabe für ein zweites Modul umfasst.
- 15Das eine oder die mehreren computerlesbaren Medien nach einem der Ansprüche 11 bis 14, wobei das mindestens eine Modul so konfiguriert ist, dass es die eine oder mehreren Nukleotidsequenzen verarbeitet, indem es mit einem externen Webserver und/oder einem externen Datenbankserver kommuniziert.
- 16Das eine oder die mehreren computerlesbaren Medien nach einem der Ansprüche 11 bis 15, wobei die Befehle, wenn sie durch die eine oder mehreren Rechenvorrichtungen ausgeführt werden, weiter folgende Auswirkungen haben:Speichern von Arbeitsablauf-Daten, die die Abfolge beschreiben;Veranlassen, dass die Arbeitsablauf-Daten mit mehreren Nutzern geteilt werden;nachfolgendes Rekonstruieren der Abfolge in einer zweiten graphischen Benutzeroberfläche, gestützt auf die Arbeitsablauf-Daten;Empfangen einer vierten Eingabe über die graphische Benutzeroberfläche, die die Abfolge so modifiziert, dass sie ein oder mehrere zusätzliche Module umfasst;Erzeugen einer zweiten Ausgabe, gestützt auf die eine oder mehreren Nukleotidsequenzen, indem jedes Modul in der Abfolge, einschließlich des einen oder der mehreren zusätzlichen Module, in einer Reihenfolge verarbeitet wird, die durch die Abfolge angegeben ist.
- 17Das eine oder die mehreren computerlesbaren Medien nach einem der Ansprüche 11 bis 16, wobei das eine oder die mehreren Module ein erstes Modul umfassen, das eine erste Ausgabe gestützt auf die Quelle erzeugt, und ein zweites Modul, das die erste Ausgabe mit einer zweiten Ausgabe von einem dritten Modul mischt, das nicht in der Abfolge ist, wobei die Quelle, das erste Modul, das zweite Modul und das dritte Modul alle Knoten in einem Arbeitsablauf sind.
- 18Das eine oder die mehreren computerlesbaren Medien nach einem der Ansprüche 11 bis 17, die weiter das Darstellen von Bedienelementen zum Auswählen des einen oder der mehreren Module umfasst, wobei die Bedienelemente zumindest Folgendes umfassen:ein erstes Bedienelement zum Auswählen eines ersten Moduls, das, gestützt auf Genomdaten, nach Veröffentlichungen in einer Online-Datenbank sucht, ein zweites Bedienelement zum Auswählen eines zweiten Moduls, das ein Sequenz-Alignment für mehrere Sequenzen ausgibt, und ein drittes Bedienelement zum Auswählen eines dritten Moduls, das Proteinfamilien für eine Nukleotidsequenz identifiziert.
- 19Das eine oder die mehreren computerlesbaren Medien nach einem der Ansprüche 11 bis 18, wobei das Empfangen der einen oder mehreren dritten Eingaben das Darstellen von visuellem Feedback umfasst, während ein erster Knoten ausgewählt wird, der anzeigt, dass Genomdaten, die von dem ersten Knoten ausgegeben werden, als Eingabe für einen zweiten Knoten verknüpft werden können.
- 20Das eine oder die mehreren computerlesbaren Medien nach einem der Ansprüche 11 bis 19, wobei das eine oder die mehreren Module mindestens zwei Module umfassen, wobei das Verarbeiten jedes der Module des einen oder der mehreren Module in einer Reihenfolge, die durch die Abfolge angezeigt ist, das automatische Verarbeiten jedes der Module umfasst, ohne menschlichen Eingriff zwischen dem Anfang des Verarbeitens eines ersten Moduls in der Abfolge und dem Erzeugen der Ausgabe, indem das Verarbeiten eines letzten Moduls in der Abfolge abgeschlossen wird.
Independent claims20
158 paragraphs in 22 sections, as filed
FIELD OF INVENTION
0001The present invention relates to data processing techniques for genome data about data that describes the nucleotide sequences.
BACKGROUND
0002The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been devised or followed earlier. Therefore, should unless otherwise indicated, are not believed that any of the approaches described in this section, only due to the inclusion in this section reflect the state of the art.
0003There exist a wide variety of genomic data, including, without limitation, data structures like DNA sequences and protein sequences annotations to these structures and publications. Genome data can be obtained from a wide variety of sources. Sequence data are, for example, one type of genome data. Typical sources of sequence data include network-based databases, such as GenBank, which is provided by the United States National Institute of Health, the European Nucleotide Archive ( "ENA") and the Protein Data Bank, operated by the Research Collaboratory for Structural Bioinformatics. These resources allow the user to access the sequence data in a variety of formats, such as plain-text files or files in FASTA format. In general, the sequence data includes a header with a Sequenzbezeichner and other metadata, and a body which comprises a sequence. In the sequence data can be accessed in a variety of ways, including to pages in a website, files that can be downloaded via HTTP and / or FTP protocols, or via a REST-based API.
0004Another type of genome data consists of annotations. Annotations can include, for example research results that relate to specific points of a sequence about an observation that a site is a binding site for a particular protein or a variant of a particular disease. The UC Santa Cruz (UCSC) Genome Browser is a popular web-based interface that allows access to various sources of annotation. Each Sequenzbezeichner may be linked to one or more Annotationseinträgen and each entry can be linked to one or more specific locations in a sequence.
0005There are also a wide variety of tools for processing of genome data. A common class of tools aliniert example sequences and comparing these sequences. Some of these tools are described in "Computer Graphical User Interface Supporting Aligning Genomic Sequences" with the Attorney Docket Number 60152-0017, filed the same day as this application, the content of which is incorporated herein by reference for all purposes as if they were in their total indicated. Another example tool is BLAST, a network-based tool for identifying similarities between an unknown protein and known proteins. A number of example algorithms for processing genomic data are in "Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids" described by Richard Durbin, Cambridge University Press 1998, the entire contents is incorporated herein by reference for all purposes as if it were indicated in its entirety. These and other tools to identify in general to be processed genomic data based on inputs, such as inputs, specify the sequences, or inputs, based on the sequences can be obtained or derived. The tools then perform one or more processing algorithms with respect to the genome data about statistical analysis, comparisons, searches, filters operations, machining, etc. The tools then create a report of all the results of the processing.
0006The analysis of genome data has become an increasingly important task. Unfortunately, such analyzes are often complex, since they are based on large amounts of separate data sources and with interconnected tools. One researcher, for example, be interested to determine how variations in a specific gene sequence affect a particular disease. The researchers can begin the analysis by retrieving a sequence from a database. The researcher can then encode the sequence by means of a first tool as a protein charge variants of the protein by means of a second tool and perform a similarity search in a big way on other databases to find species that have similar proteins. The researchers can then access other tools and databases to search for sequences in these species which code for the protein, and finally execute an algorithm to image search, to identify other proteins that bind to the protein. As a consequence of the complexity of this task, the work of the researcher disorganized and difficult to reproduce or be extended to other sequences.
0007While this application often refers to genomic data, many of the techniques that are described herein, in fact applicable to any type of data. Other uses of the techniques described herein may include, without limitation, data analysis in the field of natural language processing, the social sciences, financial data, historical and comparative linguistics, and market research.
BRIEF DESCRIPTION OF THE DRAWINGS
0008In the figures:
0009<figref>1</figref> shows an exemplary flow diagram for use of a workflow;
0010<figref>2</figref> is a block diagram of an exemplary system in which the techniques described herein may be used;
0011<figref>3</figref> is a screenshot which shows an exemplary interface to apply the techniques described herein;
0012<figref>4</figref> is a screenshot showing the representation of data nodes in the exemplary interface;
0013<figref>5</figref> is a screenshot showing the controls for importing data in the exemplary interface;
0014<figref>6</figref> is a screenshot showing the addition of a data node to the workspace of the exemplary interface;
0015<figref>7</figref> is a screenshot showing the addition of a process node to the workspace of the exemplary interface;
0016<figref>8th</figref> is a screenshot showing controls for connecting nodes in the workspace of the exemplary interface;
0017<figref>9</figref> is a screenshot showing nodes connected to the working area of the exemplary interface;
0018<figref>10</figref> is a screenshot showing the execution of a part of the workflow via the exemplary interface;
0019<figref>11</figref> is a screenshot showing the interaction with an output from a process node in the workflow via the exemplary interface;
0020<figref>12</figref> is a screenshot showing the work area with different types of knots of the workflow;
0021<figref>13</figref> is a screenshot showing an automated chain of nodes to retrieve items from a database by means of the user interface;
0022<figref>14</figref> is a pair of screenshots that show the splitting of data from one node to create a new data node in the workspace of the user interface; and
0023<figref>15</figref> is a block diagram showing a computer system upon which an embodiment of the invention may be implemented.
DETAILED DESCRIPTION
0024are in the following description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the present invention. It is clear, however, that the present invention without these specific details may be executed. In other instances, known structures and devices are shown in block diagram form to avoid obscuring the present invention unnecessarily.
1.0. GENERAL OVERVIEW
0025Methods and computer devices disclosed to process genomic data in at least partially automated workflows modules. According to one embodiment a method comprising: receiving a first input indicative of a source can be obtained from the one or more nucleotide sequences. The method further comprises receiving one or more second inputs, selecting the one or more modules for processing the data, including at least a module for processing the one or more nucleotide sequences. The method further includes displaying graphical components that represent the source and the one or more modules as nodes in a workspace, in a graphical user interface. The method further includes receiving one or more third input via the graphical user interface, the source and the one or arranging the plurality of modules as a workflow that includes a sequence of nodes. The sequence shows that for each particular module of the selected modules, the output of either one of the sources or another specific module to be inputted to the particular module. The method further comprises generating an output for theWorkflow, based on the one or more nucleotide sequences, by each module of the one or more modules is processed in an order which is to specify the sequence. The method is processed by one or more computing devices.
0026In one embodiment, each module of the one or more modules produces an output corresponding to an ontology, the data structures defined, representing the genome data. The data structures include at least sequences, protein properties, alignment objects, annotations and publications.
0027In one embodiment, the method further comprises generating a data node from the output. The data node includes elements of genome data. The data node is connected to the last module in the sequence. The method further comprises receiving a fourth input, which adds an element of the genome data to the data node or removed from, via the graphical user interface. The method further comprises receiving a fifth input, which selects a specific module to process the data node via the graphical user interface. The method further includes adding the particular module to the end of the sequence. The method further comprises generating a second output for the workflow, based on the one or more nucleotide sequences, by each module is processed in the sequence, including the particular module, in the order indicated by the sequence.
0028In one embodiment, comprise one or more modules a plurality of modules, wherein generating the output for the workflow using the output from the source as an input for the first module, and using the output of the first module as an input to a second module includes. In one embodiment, the at least one module configured so that it processes one or more nucleotide sequences, by communicating with at least one external web server and / or an external database server.
0029In one embodiment, the method further includes storing of workflow data that describes the sequence. The method further comprises causing the workflow data to be shared with multiple users. The method further includes the following restore sequence in a second graphical user interface, based on the workflow data. The method further comprises receiving a fourth input through the second graphic user interface, which modified the sequence in that it comprises one or more additional modules. The method further comprises generating a second output, based on the one or more nucleotide sequences, by each module is processed in the sequence, including the one or more additional modules in an order indicated by the sequence.
0030In one embodiment, include one or more modules, a first module that generates a first output, based on the source, and a second module, which mixes the first issue of a second edition of a third module that is not in the sequence wherein the source, the first module, the second module and the third module all nodes are in a workflow. In one embodiment, the method further comprises displaying of controls for selecting the one or more modules, the controls include at least a first control element for selecting a first module that searches in an online database for items, based on genome data, a second operating element for select them a second module that outputs a sequence alignment of several sequences, and a third control element for selecting a third module, the protein families for a nucleotide sequence identified.
0031In one embodiment, receiving the one or more third input comprises displaying visual feedback, while a first node is selected, indicating that genome data outputs can be linked from the first node as an input to a second node. In one embodiment, the one or the plurality of modules at least two modules, and processing each of the modules of one or more modules in an order indicated by the sequence comprises automatically processing each of the modules without human intervention between the start the processing of the first module in the sequence and generate an output by the processing of the last module will be completed in the sequence.
0032In other aspects, the invention encompasses a computer apparatus and a computer readable medium configured to carry out the foregoing steps.
2.0. FUNCTIONAL OVERVIEW
0033In one embodiment, the processing and analysis of genomic data by using a structure that is referred to herein as a "workflow", simplistic. Instead of each step of a research or data processing task to perform manually or instead of writing a proprietary script to these stepsexecute, a researcher can use the techniques described herein to produce a reusable and easily modifiable workflow that links these separate steps in a connected structure and some or performs all the steps of a task in an automated manner, with minimal or no user intervention.
2.1. WORKFLOWS
0034As used herein, a workflow is a set of connected nodes represent the amounts of data and transactions to be carried out on these amounts of data. In general, the nodes connected to form one or more parent sequences of nodes. Certain nodes in a sequence representing an operation, while other nodes represent data that has been issued by a process which is represented by a previous node in the sequence, and / or inputs to a process, which represents a subsequent node in the sequence becomes. The first node of an operation, for example, a data mining operation representing that retrieves data from a source, the second node of the workflow, the data representing the output from this source, the third node of the workflow may represent a process that to be performed on these data, the fourth node, an amount of data representing the results from this process etc.
0035A workflow can include any number of nodes. The usefulness of the workflow model but is generally best realized in a series of nodes that includes two or more process nodes. Furthermore, a workflow have branches. Some of these branches may be merged. Several processes can produce, for example, a single data set, or more amounts of data can be entered in a single operation. Other branches split up. an amount of data can be entered, for example, in two separate transactions, or a process can create multiple similar or dissimilar datasets.
0036An exemplary implementation of an operation is in "Document-Based Workflows", US Patent Application 2010/0070464, published March 18, 2010 described. "Document-Based Workflows" describes workflows, in which a single node type, which is referred to as a document, can act as a process node and / or a data node, with the meanings of which are described herein. Therefore, many of the techniques that are described therein, applicable to the operations described herein. The content of "Document-Based Workflows" is hereby incorporated by reference for all purposes as if it were given in its entirety here.
DATA KNOT
0037Workflow nodes that represent data are referred to herein as a data node. Data nodes may include data sets that have been imported from a data source, such as a sequence database or a library of publications, search results, manually entered data from a user and / or output data from one process node. The amounts of data that are represented by a data node, one or more elements of a similar type. A dataset may be, for example, an array of elements. Elements may comprise any type of data structure. In the context of genome data sample elements include, for example, without limitation, sequences, publications, annotations, Gendatenstrukturen, protein data structures motif data structures, disease data structures, patient data structures, etc. A data node, the amount of data that it represents include directly or a data node may include the amount of data indirectly by referring to or the location (s) on which the data set is located. A data node may further include metadata that describes the amount of data about a data type, which correspond to the elements in the data set, the rights data, field notes and / or a reference to the original source of the data set, such as one or more data and / or process node ,
0038In one embodiment, a workflow interface enabling a user to interact with each data node in a workflow for the purpose of observation. Therefore, a user can generate data nodes in certain positions of an operation in which the user wants to view data that is processed by the workflow. The workflow interface, for example, represent the data node as a group of named elements, where an interactive control panel for each item belongs. A user can select the operating element for an item to access different interfaces for viewing sequences, metadata, analysis and other information pertaining to the selected item. Data nodes can further facilitate other interactions, as described in other sections.
OPERATION KNOT
0039Process nodes are nodes that represent processes running on a dataset to be. A process node can be any type of operation represent, which is supported by a workflow application. Examples of operations that can be represented by process node and relating to genome data, described in subsequent sections.
0040In one embodiment, each process node comprises a reference to a particular module, which is responsible to perform the operation of the node and one or more optional configuration parameters for that module. A module is a reusable execution unit that performs an operation. The module may include, for example, specific commands to perform an operation, based on a given set of data. Or, the module may include instructions to send the specified amount of data to an external tool, such as an external runtime library or an external Web server, and then get a possible result. In one embodiment, a workflow application supports an extensible programming interface, which users can define a variety of different types of modules, each executing different operations.
0041An operation node may include metadata that connect the nodes with one or more input nodes. The term input node refers to any data nodes or other process nodes, which produces data on which a certain operation node performs an operation. Certain process nodes are not necessarily connected to any input node. This process node can still have an implied or user-configurable amount of data to which the operations are performed. An operation node, for example, run a query operation on a database, in which case the database constitutes an implied amount of data on which the query is performed. An operation node may also include metadata that connect the nodes with one or more output nodes. The term output node refers to any other node, including both data nodes and process nodes to the data that is produced by a process which is carried out at a particular node, are sent. Some process nodes are not necessarily connected to any output node. This may be for operation node of the case, which perform a final operation, about storage of results, or for process nodes which have not been executed.
0042For simplicity, describes various workflows in the form of process nodes. However, these operations may also have intermediate supporting data nodes representing data with which the user can interact.
typed data
0043An obstacle to the compatibility between different genomic data tools lies in the wide variety of formats that use the different tools in order to structure their results. In one embodiment workflows simplify this obstacle by converting the outputs of different tools in defined data types. Procedures For example, use a lot of data types that are defined by a particular schema or a specific ontology. The scheme or the ontology can define universal structures to represent common units of genomic data, such sequences, proteins, annotations, publications, etc.
0044Instead of working with ambiguous formatted plain-text files that correspond to the inputs and outputs of each of the workflow node standardized, predictable data types. This is because process modules must accept entries that correspond to a specified data type, and must generate outputs corresponding to a specified data type. In one embodiment, certain process modules include metadata or associated with them that define one or more input or output data types that can handle the process module. A workflow application allows a user only to connect a specific process node with other process nodes whose operation module inputs or outputs handle that or this one or this correspond more specified input or output data types, or data node, the data with this one these more specified input or output data types include.
0045To simplify the problem of using typed data in a workflow, a workflow application can provide various data conversion components. A module can transmit outputs from an external tool to a suitable conversion component together with information that will help the conversion component in understanding the data, such as the name of the tool from which the data has been received. The conversion component parses the data, and generates converted data structures based on it. Similarly, the workflow application can provide conversion components that convert the converted data in conventional input formats that are expected of different tools.Still other modules can perform such conversions by means of its own custom code.
2.2. PROCESS FLOW WORKFLOW
0046<figref>1</figref> shows an exemplary flow diagram <figref>100</figref> for use of an operation, according to one embodiment. The flow chart<figref>100</figref> is only an exemplary flow diagram for use of a workflow. Other flowcharts may include fewer or additional elements in potentially different arrangements.
0047block <figref>110</figref> comprises receiving a first input indicative of a source of one or more nucleotide sequences to be received. The first input defines essentially a data node of an operation. A data node represents an amount of data that one or more nucleotides in this case, which are received from the source. Example sources include, without limitation, one or more files in a local file system, a website, a web-based search, one or more records, an existing workflow, the contents of the clipboard, a library of previously stored sequences or process node. An input indicative of a source can be received via any suitable user interface technology. An exemplary interface for indicating sources is described in other sections. An input indicating a source, may instead be received via a text input, such as an XML file for a previously saved workflow, or a command line.
0048The data node, which is defined by the first entry is not necessarily the only data nodes in the workflow still the only data node in block <figref>110</figref> is defined. block<figref>110</figref> For example, further comprises receiving one or more other entries that specify one or more sources for other nucleotide sequences or other types of genomic data.
0049block <figref>120</figref> comprises receiving one or more second inputs, selecting the one or more modules to process data. The one or more second inputs can select, for example, one or more process modules, as described here. Therefore every second entry defines a process node for the workflow. The one or more second inputs may further specify one or more configuration parameters for the one or more modules, if necessary. The process nodes, which are defined by the one or more second entries are not necessarily all of the operation node in a workflow and block<figref>120</figref> may further comprise receiving one or more other entries that define different operation nodes.
0050As the first entry, each of the second input via any suitable user interface technology are received, including those that are described in other sections, or by text entries. In one embodiment, the modules are selected from a set of predefined modules. In one embodiment, the pre-defined modules, both modules can include, which are offered by a provider of a workflow application, as well as user-generated modules. In one embodiment, selects a second input from a module that has not been defined in advance, but is instead generated by the second input. A user can provide, for example, code or other commands for non-reusable module while defines the workflow.
0051block <figref>130</figref> comprises displaying graphical components in a graphical user interface, which represent the source and the selected modules. The display can for example comprise separate icons or other graphical representations for the source and for each of the modules. The complexity of graphical components may vary from embodiment to embodiment. Some embodiments may, for example, a source with a simple icon, while other embodiments may be a source by, identifier for some or all of the data items lists in the graphical component that is part of the source associated with the source. Examples of suitable graphical components are described herein.
0052In one embodiment, a workflow application block <figref>130</figref> run. The workflow application identifies the source and the one or more modules that are given by the first and second inputs. The workflow application then generates a visual representation of the source and or the modules in an application workspace. The application sarbeitsbereich represents a workflow that belong to the source and the one or more selected modules.
0053In one embodiment, the block happened <figref>110</figref>-<figref>130</figref> simultaneously. A user can provide, for example, the first input and the workflow application can respond immediately, by displaying a graphical component for the source. The user can subsequently provide each of the second input and the workflowApplication creates a new graphical component in response to each second input.
0054block <figref>140</figref> comprises receiving one or more third input, which arrange the source and the one or more modules in sequence. The one or more third inputs may include, for example, a fourth input, which specifies the source as the first node in a sequence, and additional inputs, which connects the one or more modules in a sequence according to the source. Similar to the first input of each third inputs may be received via any suitable user interface technology, including those that are described in other sections, or on a text input. In one embodiment, the one or more third input via the graphical user interface are received. A third input may comprise, for example, dragging a cursor from an output connection point, which is one of a graphical representation of the source, to an input junction, which is one of a graphical representation of a module.
0055The sequence indicates that expenditure of either the source or another specific module to be entered in the particular module for each particular module of the selected engines. The sequence can in reality include more nodes than just the source and the one or more modules. The source and the one or more modules may be, for example, arranged so that they follow an existing sequence of nodes and / or other nodes may be arranged so that they follow the source and the one or more modules. In addition, the workflow in reality comprise several sequences of nodes. A workflow can include, for example, two sequences that are completely separate from one another, or the workflow may comprise a sequence which is branched into another module in a different sequence or branches of this.
0056In one embodiment, the one or more third input via an interface to receive that, conditions are imposed on the types of nodes that can be connected. If any input, for example, tried to arrange a module according to a source or another module which outputs a data type that is not supported by the module, the interface will refuse to place the module in the manner which appears from the input becomes.
0057block <figref>150</figref> includes updating the graphical user interface to indicate the sequence that was arranged by the one or more third input. The source and the one or more modules can be rearranged in accordance with the sequence, for example, in the graphical user interface. Or the source and or the modules can be connected by lines or other suitable connecting elements in an order with each other, which is indicated by the sequence. In one embodiment, updated in response to each of the third entry, the graphical user interface is to display a new arrangement, instead of waiting until all the received one or more third inputs.
0058In one embodiment, the block can <figref>140</figref> and <figref>150</figref> simultaneously with the block <figref>110</figref>-<figref>130</figref> be executed. The user can add, for example, a source and a first module to a workspace and then connect the source to the first module. The user can then add a second module to the workspace, and then connect the second module to the original module. The graphical user interface can be updated continuously, while the user provides these inputs.
0059block <figref>160</figref> comprises processing each module of the selected modules in an order indicated by the sequence. The processing of the modules in a workflow is described in subsequent sections.
0060block <figref>170</figref> includes generating an output based on this processing. Since the one or more nucleotide sequences were used as input to at least one module, the output is based at least on the one or more nucleotide sequences. Of course, the output can be further based on other data entries in the workflow if it is defined. The output generated by the processing of the last or next to last module in the sequence. Thus comprises block<figref>160</figref> basically block <figref>170</figref>,
0061block <figref>180</figref> includes the optional save the output. The output may for example be stored in a local database or a local file system. Or the output can be uploaded to a web-based database. Or, the output can be sent to another user. block<figref>180</figref> can be implemented as part of the processing of the last module in a sequence, but need not be. The last node in a workflow can be for example a process node executing the save operation. Or the last node in the workflow, a data node to the output of block<figref>170</figref> be. In the second case block<figref>180</figref> run outside the processing workflow. The user can, for example, to import the amount of data that is represented by the last node in the workflow, manually into a database.Or the user can copy and paste the amount of data in a report, which is then stored in a file.
2.3. PROCESSING A WORKFLOW
0062In one embodiment, the processing of an operation includes processing a sequence of nodes. A first operation node in the sequence of nodes corresponding to a first module. The processing workflow includes executing the first module, based on data input from a data node representing a source. An output is generated based on the execution of the first module. In one embodiment, the process further comprises the execution of a second module, which is represented by a second operation node in the sequence. Based on the output of the first module, a second output is generated, based on the execution of the second module. In one embodiment, the process further comprises the step executing each module, which is represented by one of the following process node, by the output of an immediately preceding process node as input until all process nodes have been processed in the sequence.
0063In one embodiment, the processing of an operation includes the "processing" of a data node. The processing of a data node includes documents of the data node with an amount of data that has been output from a previous process nodes. The processing of the data node may further comprise receiving interactive processing of the data set by the user, as described below, but need not be. The amount of data is transmitted as an input to each successive process node.
0064In one embodiment, the processing of an operation includes processing a plurality of sequences of nodes. Any given sequence or at any given node may depend on outputs from each other given sequence or any other given node in the workflow. However, once a node or a sequence, which depends from another node or another sequence was processed, the other node or the other sequence can be processed, regardless of how the timing for any other node, or any other sequence is. Multiple sequences can be performed, for example, parallel or at any other time with respect to each other.
0065In one embodiment, the Execute ( "processing") of a module comprises the execution of commands that are defined for the module. The commands are executed optional, based on one or more configuration parameters that are defined in a second entry. In one embodiment, the send commands an inquiry to an external component, such as a network-based server or an external application. includes the request or referenced data input to the module during processing, which may have been reformatted in accordance with the commands of the module, but it does not have. In response, the module receives data from the external component. The module can optionally format the returned data newly or other processing before it returns as output.
AUTOMATED WORKFLOW
0066In one embodiment, some or all of the operations are processed automatically in a non-interactive manner. Once such workflow has been defined, the processing of the workflow requires no further user input between the timing at which the first node has been processed, and the time at which the last module has been processed.
INTERACTIVE WORKFLOW
0067While some work processes that are described here are designed so that they produce outputs without human intervention, other workflows are designed to support a user in identifying processes and investigations, rather than simply to produce an output. For such operations, can provide in various stages of designing and utilizing workflow, the user various inputs to interact with the data flow and / or edit it. A user can, for example, to execute part of the workflow. Based on outputs from the execution of this part of the user may choose to perform other parts of the workflow, and / or to define the workflow so new that it comprises additional nodes or sequences of nodes.
0068In one embodiment, a user can edit the data within each data node. Therefore, a data node has a position in the workflow representing where the user can make an informed decision with respect to how the workflow should proceed. In contrast, the required workflow for data processing, in which the user does not want to intervene, no data node. Thus can follow without intervening data nodes to each other several process nodes, indicating that the processing of data is done fully automated at these positions. However, since data node can also serve the purpose of observation, requires the presence ofData node is not that the user needs to edit the data in the data nodes.
0069A user can for example edit a data node, by adding or removing elements, thereby allowing the user the amount of data for all subsequent workflow processes to filter, for which the data node provides inputs. The user may have a first part of an operation "run" to generate the data nodes. The user can then edit the data node before proceeding with the second part of the workflow or even created him first. Similarly, the user can create new data node, by transmitting or copying elements from another data node. These new data nodes can be connected to process node in the workflow.
2.4. REUSE OF WORK PROCESSES
0070In one embodiment, the user can save a workflow for subsequent reuse. A workflow data structure, such as an XML file or another data object, for example, can describe the workflow. A user can save the workflow data structure in a file system or a database. The user can subsequently access the workflow data structure to run the workflow again. The user can for example load the workflow data structure into a workflow application. The workflow application can display graphical representations of the workflow, which is described by the workflow data structure. The user can then run the workflow as it was created at the time when the workflow has been saved, or the user can modify the workflow to potentially handle different data sources in different ways.
0071In one embodiment, certain stored Procedures can be used as templates from which the user can quickly create new workflows. In one embodiment, a workflow can be shared with other users. A user can, for example, send a workflow to another user via email or link to the workflow. If the other user has access to the same data sources and the same modules - for example by means of a central server resources - to perform and / or change the workflow of other users. If the other user does not have access to the same data sources and the same modules, various techniques for finding replacement sources and modules can be applied. Or a shared workflow, in order to avoid problems with resource dependencies, modules and data sources to embed on another user probably does not have access.
0072In one embodiment, the user can configure a stored, non-interactive workflow so that it is executed in response to trigger or periodically. A data source can be changed, for example, periodically or in response to specific events. A user can create and save to produce, for example, updated report data when the data source has changed, or data in a database to reimport, having regard to the changes an automated workflow. In one embodiment, a process to monitor the output of an automated workflow and a user send an update message as soon as the output is changed. A user can, for example, configure a workflow so that it runs every morning. The workflow can usually provide the same results. The user can request to receive an automated email when changes the output of the workflow. The user can then examine the new data.
3.0. STRUCTURAL OVERVIEW
0073<figref>2</figref> is a block diagram of an example system <figref>200</figref>In which the techniques described herein may be applied, according to an embodiment. The various components of the system<figref>200</figref> For example, the flow chart <figref>100</figref> implement, such as it is described above.
0074The system <figref>200</figref> includes a workflow system <figref>210</figref>, The workflow system<figref>210</figref> includes one or more computing devices that a number of components <figref>220</figref>-<figref>260</figref> implement, provide different functionality with respect to workflows. The workflow system<figref>210</figref> For example, a client computing device and a server computing device comprise. As another example, the workflow system<figref>210</figref> a single computing device include. The components<figref>220</figref>-<figref>260</figref> may be on one or more computing devices and software that is executed by this hardware, any combination of hardware. In one embodiment, the components are<figref>220</figref>-<figref>260</figref> herein collectively referred to as "workflow application".
0075A workflow-generating component <figref>230</figref> provides a workflow interface component <figref>240</figref> a user <figref>205</figref> ready. The workflow-generating component<figref>230</figref> generates workflows in response to various inputs of User in the workflow interface component <figref>240</figref>, The various user inputs, for example, the workflow generating component<figref>230</figref> instruct add data nodes and process nodes to a workflow to edit these nodes and establish links between specific nodes. The workflow-generating component<figref>230</figref> updates the workflow interface component <figref>240</figref>To display representations of the nodes and / or the compounds in a workflow, while the user inputs are received.
0076The workflow-generating component <figref>230</figref> generated process node, the process modules <figref>250</figref> represent to process workflow data. process modules<figref>250</figref> are execution units, enter the data and / or output, as described in other sections. The workflow-generating component<figref>230</figref> learns about the availability of this process modules <figref>250</figref> and configuration options, and limitations of the modules <figref>250</figref>By metadata of the module <figref>255</figref> accesses. If a workflow application, for example, called for the first time, the workflow application can a folder or other metadata<figref>255</figref> for the modules <figref>250</figref> scan and then all modules found <figref>250</figref> make available for use in the process node of a workflow. The workflow-generating component<figref>230</figref> then generates interface controls in the workflow interface component <figref>240</figref>That allows the user <figref>205</figref> allow you to create a new process node and the process node with one of the modules <figref>250</figref> link. The workflow-generating component<figref>230</figref> can generate a data node, based on inputs from the user <figref>205</figref>Indicating an amount of data.
0077The workflow-generating component <figref>230</figref> can continue to generate data node, based on outputs from the processing of a process node, as from an immediately preceding process node in the current sequence, or by a process node at the end of another workflow.
0078The workflow-generating component <figref>230</figref> further generates data node, based on data from a user of a converted data storage <figref>290</figref> and / or data sources <figref>280</figref> were selected. The workflow-generating component<figref>230</figref> retrieves data quantities from the converted data storage <figref>290</figref> and / or the data sources <figref>280</figref> from. While the amount of data to be accessed, the data sources are<figref>280</figref> from the data conversion component <figref>260</figref> converted what uniform, typed data structures provides. Amounts of data from the converted data storage<figref>290</figref> are organized on the other side even as a uniform, typed data structures. The workflow-generating component<figref>230</figref> presents these amounts of data to the user <figref>205</figref> in the workflow interface component <figref>240</figref> and receives in response to inputs, select the specific data elements of the data volumes to be included in the data nodes.
0079Once a workflow has been created, save the workflow generating component <figref>230</figref> a workflow data structure that represents the workflow in a workflow memory <figref>235</figref>, The workflow memory<figref>235</figref> for example, can be a temporary location in memory, in a directory on a local file system or in a database.
0080The workflow-interface component <figref>240</figref> further comprises control elements, by which the user <figref>205</figref> The workflow-processing component <figref>220</figref> may instruct at least a portion of a currently loaded workflow or a workflow in the workflow storage <figref>235</figref> to process. The workflow-interface component<figref>240</figref> For example, a "workflow run" - and / or "running nodes" representing button, the workflow-processing component <figref>220</figref> caused at least to process a part of the workflow, which is currently in the workflow interface component <figref>240</figref> is displayed. The workflow-processing component<figref>220</figref> can also or instead carry out operations that are indicated by other inputs, such as command-line input or input from a task scheduler.
0081The workflow-processing component <figref>220</figref> performs operations that through workflow data structures in the workflow storage <figref>235</figref> describes using workflow-processing techniques, as described in other sections. In the course of processing a workflow, calls the workflow-processing component<figref>220</figref> process modules <figref>250</figref> on that are referenced by process node in the workflow. The workflow-processing component<figref>220</figref> can a typed dataset that was issued by a previous node, if there is one, to a called process module <figref>250</figref> to hand over. Most process modules<figref>250</figref> will process one or more commands with respect to the input data set and then an output consisting of typed data, to the workflow processing component <figref>220</figref> hand back. In one embodiment, the workflow-processing component<figref>220</figref> generate data node in the presently processed workflow and / or update to include the amount of data of the process modules <figref>250</figref> were issued. Inone embodiment, when the workflow-processing component <figref>220</figref> runs in an interactive mode, the workflow-processing component <figref>220</figref> the workflow interface component <figref>240</figref> Next update so that it displays representations of added or updated data node. In one embodiment, the workflow-processing component<figref>220</figref> be configured to the final output of a workflow in a variety of places, which differs from the workflow interface component <figref>240</figref> differ, stores, prints or displays.
0082In one embodiment, one or more process modules <figref>250</figref> self-contained instructions for processing a dataset. Code for relatively frequent and / or simple operations, such as mixing or filtering a dataset, for example, directly into a module<figref>250</figref> be provided. In one embodiment, processes a process module<figref>250</figref> an amount of data only by means of self-contained commands, without any external tools <figref>270</figref> call. In one embodiment interacts a module<figref>250</figref> with one or more external tools <figref>270</figref>To process a data set. The one or more external tools<figref>270</figref> Various algorithms for processing genomic data implement. The process modules<figref>250</figref> Send part or all of the amount of data entered or processed data that are based on it, to the external tool <figref>270</figref> for processing. The process modules<figref>250</figref> receive it then output. The process modules<figref>250</figref> can optionally process an output before the output as typed data to the workflow-processing component <figref>220</figref> hand back. In one embodiment, users can create their own process modules<figref>250</figref> provide an API that the workflow system <figref>210</figref> also can be used. The external tools<figref>270</figref> For example, local runtime libraries <figref>270a</figref> comprise about redistributable libraries of Java or Python code directly on procedure calls in a process module <figref>250</figref> can be called. The external tools<figref>270</figref> may also include client-side libraries in the workflow interface component <figref>240</figref> running on the computing device of the user. A module<figref>250</figref> for example, can be implemented using client-side JavaScript tools. Such tools can ask for user input, the output of the module<figref>250</figref> influence, but it need not.
0083The external tools <figref>270</figref> can also local application server <figref>270c</figref> and network-based application server <figref>270d</figref> include, by which a process module <figref>250</figref> can communicate via one or more networks via any suitable protocol, including HTTP, FTP, REST-based protocols, JSON, etc. In one embodiment, all process modules <figref>250</figref> coded objects that extend a common class. The common class implements logic for communicating with each of these four types of external tools<figref>270</figref>, In one embodiment, the external tools can<figref>270</figref> Other tools include that are not shown. In one embodiment, some external tools can<figref>270</figref> continue with other external tools <figref>270</figref> communicate, to generate an output. In one embodiment, some external tools can<figref>270</figref> produce outputs based on the requesting data from data sources <figref>280</figref>,
0084Data sources <figref>280</figref> can any source of data includes, on the system of <figref>210</figref> may be accessed, including local files <figref>280a</figref> in each of a plurality of formats and interrogated local databases <figref>280b</figref>, Data sources<figref>280</figref> can continue to network-based storage <figref>280c</figref> comprise, on which can be accessed through various web-based interfaces, including SOAP or REST-based interfaces. In one embodiment, in order to operate the system<figref>210</figref> speed, network-based storage <figref>280c</figref> locally as local files <figref>280a</figref> and / or local databases <figref>280b</figref> cached. The workflow system<figref>210</figref> For example, periodic database dumps from the network-based storage <figref>280c</figref> Download. Data sources<figref>280</figref> can continue Websites <figref>280d</figref> include. The process modules<figref>250</figref> and or the data conversion component <figref>260</figref> For example, "screen scraping" elements have to publications or other data from web pages <figref>280d</figref> to extract from certain sites.
0085In one embodiment, data must be used by the external tools <figref>270</figref> and the data sources <figref>280</figref> be issued in converted data <figref>290</figref> be converted before it by the workflow system <figref>210</figref> are processed. The workflow system<figref>210</figref> , a data conversion component <figref>260</figref> provide to data from the external tools <figref>270</figref> and data sources <figref>280</figref> reformat in typed data structures represented by an ontology <figref>291</figref> are defined. The ontology<figref>291</figref> may be created for each type of data. In one embodiment, an ontology<figref>291</figref> for genome data, the following core data types: sequences, DNA sequences, mRNA sequences, RNA sequences, protein sequences, protein objects, paper objects, alignment objects and Genobjekte.
0086The converted data <figref>290</figref> are then stored, at least temporarily, in order to process the work flow, and / or persistent, so user <figref>205</figref> can access below in other workflows and projects it. In one embodiment, the data conversion component<figref>260</figref> further convert data back to a shape of the external tools <figref>270</figref> and data sources <figref>280</figref> is expected so that it as an input for external tools <figref>270</figref> or for storage in the data sources <figref>280</figref> can be used. In one embodiment, the process modules may<figref>250</figref> also or instead be responsible for converting some of the data sets directly, which to an associated tool <figref>270</figref> sent or received by him. In one embodiment, some process modules<figref>250</figref> directly to converted data <figref>290</figref> supported, which are permanently stored in a database, in contrast to data from the data sources <figref>280</figref>,
0087The system <figref>200</figref> is only one example of a system in which the techniques described herein may be employed. Other systems may include additional or fewer elements comprise, in potentially different arrangements. In one system, for example, absent any number of types of external tools<figref>270</figref> or data sources <figref>280</figref>, In another system, the converted data is missing<figref>290</figref> and the data conversion component <figref>260</figref>, In yet another system lacks a graphical workflow interface component<figref>240</figref>, Many other variants are also possible.
4.0. EXEMPLARY INTERFACE AND WORKFLOW
0088<figref>3</figref> is a screenshot <figref>300</figref>, An exemplary interface <figref>305</figref> for performing techniques shows that are described herein, according to an embodiment. the interface<figref>305</figref> for example, can make it easier to receive inputs from a user to define a workflow and / or interact with the processing of a workflow. the interface<figref>305</figref> is an example of an operation interface component <figref>240</figref>,
0089the interface <figref>305</figref> may include various graphical representations of elements such as nodes, data elements, modules, files, etc. In order to simplify the disclosure, this application describes characteristics of the graphical user interface sometimes represented in the form of elements itself, rather than graphical representations of these elements. One skilled in the art will understand that, as it is customary when graphic user interfaces are described, literal descriptions of a graphical user interface that includes non-graphical user interface components, shall be understood as a description of the graphical user interface in that it comprises graphical representations of these components , The description may for example describe a step of "selecting a node from a workspace", if a skilled person will understand, in fact, that what is selected, is an illustration of a node in the workspace.
0090the interface <figref>305</figref> a work space section <figref>310</figref>In which a workflow <figref>320</figref> is displayed. The various components of the workflow<figref>320</figref> are described with reference to subsequent figures. The workspace<figref>310</figref> further comprises Zoom Controls <figref>312</figref>To the visible area of the workspace <figref>310</figref> To zoom in or out. In one embodiment, the visible section of the workspace<figref>310</figref> movable over different combinations of cursor inputs and / or selection of scrolling controls.
0091the interface <figref>305</figref> further includes a header area <figref>390</figref>, The header area<figref>390</figref> includes controls <figref>391</figref>-<figref>397</figref> for general workflow operations. Storing control element<figref>391</figref> facilitates inputs for storing workflows <figref>320</figref>, The opening-control element<figref>392</figref> facilitates entry to open a previously saved workflow in the work area <figref>310</figref>, Running control element<figref>393</figref> facilitates inputs for processing the entire workflow <figref>320</figref>, Or, if one or more specific nodes of a workflow<figref>320</figref> currently selected, facilitates the execution-control element <figref>393</figref> Entry for processing a portion of the workflow <figref>320</figref>Which belongs to the one or more particular nodes. The controls<figref>394</figref>-<figref>396</figref> facilitate inputs to generate various kinds of presentations, based on outputs from the workflow <figref>320</figref>, The operating element<figref>397</figref> facilitating inputs to import selected outputs of the workflow <figref>320</figref>Including amounts of data that are stored in the intermediate data nodes in a data store.
0092the interface <figref>305</figref> further includes a sidebar area <figref>370</figref>Which is generally provided for controls that creating new nodes in a workflow <figref>320</figref> facilitate. The sidebar area<figref>370</figref> includes four panes <figref>371</figref>-<figref>374</figref>, The currently displayed pane, looking Pane<figref>371</figref>, Is a database search operating element <figref>375</figref>, The database search operating element<figref>375</figref> allows the user to perform a conceptual-based search various databases of genomic data. The user can search results, in part or as a whole, to the workspace<figref>320</figref> drag to create one or more new data nodes. The sidebar<figref>370</figref> further includes an Import Pane <figref>372</figref>. a task pane <figref>373</figref> and a library pane <figref>374</figref>,
0093the interface <figref>305</figref> further includes an overview of the range <figref>380</figref>, The survey area<figref>380</figref> generally provides a context-sensitive detailed view of information about a currently selected object in the work area <figref>310</figref> . As shown, provides the overview area <figref>380</figref> For example, a "publication view" of a particular publication element is that in a data node of the workflow <figref>320</figref> was selected. Depending on the data type and / or the node type of the currently selected element in the workspace<figref>320</figref> , the overview area <figref>380</figref> represent different organized views of different information fields. Some views may comprise a single information field, while other views can include many information fields. In one embodiment, the information contained in the survey area<figref>380</figref> are displayed, the user definable. The survey area<figref>380</figref> can be scrollable, depending on which view is displayed.
0094<figref>4</figref> is a screenshot <figref>400</figref>, The representation of the data nodes in the exemplary interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>400</figref> shows a part of the workspace <figref>310</figref>, Including graphical representations of three different data nodes <figref>430</figref>-<figref>450</figref>, And the overview area <figref>380</figref>, Each of the node data<figref>430</figref>-<figref>450</figref> comprises an amount of data elements that are shown in the respective graphs. The data node<figref>430</figref> for example, includes at least the data elements <figref>431a</figref>-g, which are another object protein respectively. A user can by means of the scrolling control element<figref>435</figref> make additional data items visible.
0095The user can select a particular data item, such as the element <figref>431a</figref>, By clicking on it or not. By any other suitable selection technique In response, the survey area<figref>380</figref> with a view <figref>485</figref> Updating of information on the selected item <figref>431a</figref> are linked. The information in the view<figref>485</figref> may change in response to a user the window areas <figref>481</figref>-<figref>484</figref> clicks. Each of the panes<figref>481</figref>-<figref>484</figref> brings a different amount of information about the protein <figref>431a</figref> in the view <figref>485</figref>Including summary information (in the pane <figref>481</figref>), References (in the pane <figref>482</figref>), Sequence data (in the pane <figref>483</figref>) And PDB data (in the pane <figref>484</figref>).
0096A user can also an entire data nodes <figref>430</figref>-<figref>450</figref> Select, by clicking on it or by any other suitable selection technique. Clicking on a data node can cause the overview area<figref>380</figref> a different view of other information than the information shows that in <figref>4</figref> are shown. The view for the node<figref>430</figref> for example, can include overview information for a whole amount of data, such as statistical analysis of a histogram showing how similar the protein elements <figref>431</figref> are to each other.
0097In contrast, when an operation node is selected, the overview area <figref>380</figref> include metadata describing the module, which is part of the process nodes, information on the latest version of the module and / or fields for entering values for configurable parameters of the module.
0098<figref>5</figref> is a screenshot <figref>500</figref>, The controls for importing data in the exemplary interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>500</figref> shows the workspace <figref>310</figref>That Sidebar <figref>370</figref> and the header <figref>390</figref>, The Import pane<figref>372</figref> was in the Sidebar <figref>370</figref> selected. Therefore shows the Sidebar<figref>370</figref> an import-control element <figref>575</figref>To receive inputs, select a file. The screenshot<figref>500</figref> further shows a file manager window <figref>560</figref>From which the user a representation of a file <figref>561</figref> can select. The user can then move the cursor<figref>565</figref> use to the display of the file <figref>561</figref> "Take" to and through the import control element <figref>575</figref> To "pull". A feedback-graphic<figref>562</figref> can be displayed to show the user that the user is actually the file <figref>561</figref> by means of the cursor <figref>565</figref> draws. Once the cursor<figref>565</figref> on the import control element <figref>575</figref> is, the user can view the file <figref>561</figref> in the interface area <figref>575</figref> "Drop" to the user interface <figref>305</figref> to instruct, try the file format of the file <figref>561</figref> to recognize the file <figref>561</figref> automatically into one or convert a plurality of data elements that can be used in an operation, and these data elements in the interface <figref>305</figref> to import.
0099<figref>6</figref> is a screenshot <figref>6:00 am</figref>, Of adding a data node <figref>630</figref> to the workspace <figref>310</figref> the exemplary interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>6:00 am</figref> showing parts of the workspace <figref>310</figref> and Sidebar <figref>370</figref>, The sidebar<figref>370</figref> still shows the Import Pane <figref>372</figref> with the import-control element <figref>575</figref>, Additionally includes the Sidebar<figref>370</figref> Representations of data elements, <figref>661</figref> and <figref>662</figref>, The data elements<figref>661</figref> and <figref>662</figref> are sequences derived from the file <figref>561</figref> were imported. Adjacent to the representations of the data elements<figref>661</figref> and <figref>662</figref> are operating <figref>665</figref> and <figref>666</figref> To add the data elements <figref>661</figref> or. <figref>662</figref> a data node, which is currently in the workspace <figref>310</figref> selected is. If there is no currently selected data nodes in the workspace<figref>310</figref> are, a new data node is created when one of the operating elements <figref>665</figref> or <figref>666</figref> was selected.
0100As in <figref>6</figref> shown, includes the workspace <figref>310</figref> a representation of the data node <figref>630</figref>That was generated in response to a user, the operating element <figref>665</figref> clicked. Thus, as in the content display area<figref>631</figref> the knot <figref>630</figref> As shown, the node <figref>630</figref> the imported data element <figref>661</figref>, The representation of the node<figref>630</figref> further comprises removing control element <figref>639</figref>That to remove the node <figref>630</figref> of the workspace <figref>310</figref> leads, and a node-title <figref>632</figref>, Of the art describes as the default, by the node <figref>630</figref> was created (ie the fact that he "imported" was).
0101<figref>7</figref> is a screenshot <figref>7:00</figref>, Of adding a process node <figref>740</figref> to the workspace <figref>310</figref> the exemplary interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>7:00</figref> showing parts of the workspace <figref>310</figref> and Sidebar <figref>370</figref>, The process Pane<figref>373</figref> was in the Sidebar <figref>370</figref> selected, which causes the sidebar <figref>370</figref> the operating element groups <figref>770</figref>. <figref>780</figref> and <figref>790</figref> includes. The operating element groups<figref>770</figref>. <figref>780</figref> and <figref>790</figref> each include controls <figref>771</figref>-<figref>775</figref>. <figref>781</figref>-<figref>782</figref> and <figref>791</figref>-<figref>792</figref> To add operation node to the workspace <figref>310</figref>, The workspace<figref>310</figref> includes a presentation of the recently added process node <figref>740</figref>, The process node<figref>740</figref> may have been created and added to its associated presentation to the workspace in response to the control element <figref>775</figref> of the workspace <figref>370</figref> was selected. As shown, the process node<figref>740</figref> a different shade than the data nodes <figref>630</figref>, In one embodiment, all data nodes are differently shaded or colored as process nodes.
0102The controls <figref>771</figref>-<figref>775</figref>. <figref>781</figref>-<figref>782</figref> and <figref>791</figref>-<figref>792</figref> for example, have been produced by one or more plug-in directories have been scanned, in which the workflow application expected to find modules. The operating element groups<figref>770</figref>. <figref>780</figref> and <figref>790</figref> may have been created, based on module metadata that categorizes each of the module plug-ins. The operating element group<figref>770</figref> belongs to a "sequence" category of modules. Selecting one of its controls<figref>771</figref>-<figref>775</figref> creates a process node executing an operation that is implemented by a "MSA" module, a "BLAST" module, a "transcription" module, a "translation" module and a "ScanProsite" module. The operating element group<figref>780</figref> belongs to a "simple" category of modules and includes a mixed-operating element <figref>781</figref> and a filter-control element <figref>782</figref>, Selecting one of the operating elements<figref>781</figref>-<figref>782</figref> generates a process node that is performing an operation that is implemented by a "mixed" module and a "filter" module. The operating element group<figref>790</figref> belongs to a "search" category of modules. Selecting one of its controls<figref>791</figref>-<figref>792</figref> creates a process node that performs an operation that is implemented by a "PubMed" module and a "UniProtKB" module. Because users can create their own modules easily and because the workflow application automatically generates controls for all modules, which generates a user, the operating element groups<figref>770</figref>. <figref>780</figref> and <figref>790</figref> and the controls <figref>771</figref>-<figref>775</figref>. <figref>781</figref>-<figref>782</figref> and <figref>791</figref>-<figref>792</figref> only a small selection of the operating element groups and controls, in the sidebar <figref>370</figref> may appear.
0103<figref>8th</figref> is a screenshot <figref>800</figref>, The controls for connecting nodes in the workspace <figref>310</figref> the exemplary interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>800</figref> shows a part of the workspace <figref>310</figref>Including representations of the Node <figref>630</figref> and <figref>740</figref>, The representation of the node<figref>740</figref> was closer to the representation of the node <figref>630</figref> in response moved on user inputs, such as input from the user, the node to <figref>740</figref> Pull in the currently displayed location and place. The knot<figref>630</figref> includes an input connection point <figref>634</figref> and an output junction <figref>635</figref>, Similarly includes the node<figref>740</figref> an input junction <figref>744</figref> and an output junction <figref>745</figref>,
0104In one embodiment, a user can each node to every other node to connect by pulling its output junction to the input junction of the other node or its input junction to the output junction of the other node. The node whose output connection point has been connected to the input junction of the other node, provides input for the other node ready and it is therefore assumed that it is placed before the other node in the sequence.
0105As in <figref>8th</figref> shown, the output junction was <figref>635</figref> the knot <figref>630</figref> to the input junction <figref>744</figref> the knot <figref>740</figref> drawn. The junction<figref>635</figref> has changed its color, and the cursor <figref>865</figref> appears as a connection symbol to indicate that the user is currently the joint <figref>635</figref> draws. The junctions labels<figref>861</figref> and <figref>862</figref> also appear while the user is the junction <figref>635</figref> draws, which provides information about the selected junction and other junctions when it is appropriate. In one embodiment, a junction with onlyanother junction are connected when the two connection points belong to the same data type. To assist a user in identifying which connection points belong to the same data type, the user interface<figref>305</figref> further change the appearance of all contact points which are compatible with the connection point, which is currently selected. Two joints are compatible if they have opposite types of connection at least support (input vs. output), a common data type, are not in the same node and not both in data nodes. Thus, it was because the input junction<figref>744</figref> with the output junction <figref>635</figref> is compatible, the input junction <figref>744</figref> shaded with a consistent color without border, as opposed to the input node <figref>634</figref>Which is transparent and has an edge. The junction<figref>745</figref> also has an edge, which indicates that they joint the <figref>635</figref> can not receive. The knot<figref>745</figref> but is currently shaded, since it represents an output that no other node is currently provided. A variety of other techniques to change the appearance of compatible nodes may also or instead.
0106<figref>9</figref> is a screenshot <figref>900</figref>, The nodes connected in the workspace <figref>310</figref> the exemplary interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>900</figref> shows a part of the workspace <figref>310</figref>In which the node <figref>630</figref> and <figref>740</figref> were joined by the drag-and-drop operation, which is described above. The workspace<figref>310</figref> includes a representation of the compound <figref>961</figref> between nodes <figref>630</figref> and <figref>740</figref>, nodes<figref>630</figref> and <figref>740</figref> now form a sequence, thus forming a functional workflow <figref>320</figref>,
0107<figref>10</figref> is a screenshot <figref>1000</figref>, The running of a portion of the workflow <figref>320</figref> by means of exemplary interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>1000</figref> showing the header area <figref>290</figref> and a part of the workspace <figref>310</figref>Including nodes <figref>630</figref> and <figref>740</figref>, After the node<figref>630</figref> and <figref>740</figref> has connected, a user can choose the workflow <figref>320</figref> perform. Thus, the user can execute control element<figref>393</figref> . click In response, the workflow application workflow<figref>320</figref> run by the sequence represented by the node <figref>630</figref> is represented, enters into the ScanProsite module, which by the node <figref>740</figref> is represented, and performs the ScanProsite module. The ScanProsite module interacts in turn with a Web server that implements an algorithm to identify motifs in the sequence. The ScanProsite module receives a response from the Web server, interprets this response as a data set of designs and makes this data set of workflow application ready. The workflow application generates a data node<figref>1030</figref>Adds the identified motifs to the data nodes <figref>1030</figref> as data elements <figref>1031a</figref>-<figref>1031c</figref> added, adds the data nodes <figref>1030</figref> to the workflow <figref>320</figref> added by the data node <figref>1030</figref> to the node <figref>740</figref> with a new connection <figref>1061</figref> connects, and then adds appropriate representations of the new data to the workspace <figref>310</figref> added. These representations are in the work area<figref>310</figref> of the <figref>10</figref> shown.
0108<figref>11</figref> is a screenshot <figref>1100</figref>, Of interacting with outputs from one process node in the workflow <figref>320</figref> by means of exemplary interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>1100</figref> showing parts of the survey area <figref>380</figref> and the workspace <figref>310</figref>Including workflow <figref>320</figref>, As in <figref>10</figref> was created. A particular data element<figref>1031a</figref> was from the node <figref>1030</figref> selected. In response, the survey area<figref>380</figref> updated with information associated with the element <figref>1031a</figref> are linked, including labeling <figref>1181</figref>, metadata <figref>1182</figref> and a sequence <figref>1183</figref>,
0109<figref>12</figref> is a screenshot <figref>1200</figref>, Of the work area <figref>310</figref> with different types of knots of the workflow <figref>320</figref> shows, according to an embodiment. The screenshot<figref>1200</figref> showing parts of the Sidebar <figref>370</figref> and the workspace <figref>310</figref>, While the node<figref>630</figref> in the workspace <figref>310</figref> has been scrolled out of view, including the workspace <figref>310</figref> Now a number of additional nodes to the workflow <figref>320</figref> were added. Specifically, the node<figref>1030</figref> now as an input to a process node <figref>1240</figref> connected, in turn, as an input to the operation node <figref>1250</figref> connected is. Another data nodes<figref>1230</figref> was also to the workspace <figref>310</figref> added. The data node<figref>1230</figref> is as an input to the operation node <figref>1250</figref> connected.
0110The library pane <figref>374</figref> is in the sidebar <figref>370</figref> selected. Therefore includes the Sidebar<figref>370</figref> three controls <figref>1281</figref>-<figref>1283</figref> to add items from a library. In one embodiment, a library is a local memory in which users can store data items of interest to the user. Thus, the library pane<figref>370</figref> many more controls include, depending on which elements were added by a user. In one embodiment, library elements are shared with a group of users. Each of the control elements<figref>1281</figref>-<figref>1283</figref> corresponds to a different library item. Selecting one of the operating elements<figref>1281</figref>-<figref>1283</figref> leading to the addition of the associated library member to the data currently selected node or creating a new data node when no compatible Data node is selected. The data node<figref>1230</figref> was generated, for example, as the user, the operating element <figref>1282</figref> clicked.
0111The process node <figref>1240</figref> belongs to a filter module. The process node<figref>1240</figref> For example, to the workspace <figref>310</figref> have been added in response to a user, the operating element <figref>782</figref> clicked. By default, the filter module is configured to the data nodes<figref>1030</figref> so filtered that it only the first element <figref>1031a</figref> includes, but the user can filter behavior to the process node <figref>1240</figref> heard reconfigure by the node <figref>1240</figref> selects and parameter values change, which in the overview area <figref>380</figref> are shown in response to the selecting.
0112The process node <figref>1250</figref> is part of a mixing module. The process node<figref>1250</figref> For example, to the workspace <figref>310</figref> have been added in response to a user, the operating element <figref>781</figref> clicked. The process node<figref>1250</figref> includes a plurality of input connection points <figref>1253</figref> and <figref>1254</figref>To allow the node <figref>1250</figref> to enable to receive multiple inputs. The mixing module is configured so that it generates a data set from the plurality of inputs. The knot<figref>1250</figref> mixed, for example, as shown, the output of the node <figref>1240</figref> with the data in the node <figref>1230</figref>,
0113<figref>13</figref> is a screenshot <figref>1300</figref>, The automated chain of nodes to retrieve items from a database using the user interface <figref>305</figref> shows, according to an embodiment. The screenshot<figref>1300</figref> shows a part of the workspace <figref>310</figref>, Including most of the workflow <figref>320</figref>, The workflow<figref>320</figref> now includes a process node <figref>1340</figref> and a data node <figref>1330</figref>, The process node<figref>1340</figref> receives the output of mixed-node <figref>1250</figref> and sends the output to a module for searching a PubMed database for articles. The process node<figref>1340</figref> was generated in response to a user, the operating element <figref>791</figref> has selected.
0114After the process node <figref>1340</figref> has been added, the user workflow <figref>320</figref> executed expenditure for the operation node <figref>1340</figref> to create. These expenses, which include a lot of publications, were in the data nodes<figref>1330</figref> than at least the data elements <figref>1331a</figref>-<figref>1331j</figref> stored. A user can other data elements by means of the scrolling control element<figref>1235</figref> consider that in the node <figref>1330</figref> are.
0115<figref>14</figref> is a pair of screenshots <figref>1400</figref> and <figref>1450</figref>, The splitting of data from said data node <figref>1330</figref> show to a new data node <figref>1430</figref> in the workspace <figref>310</figref> UI <figref>305</figref> to generate, in accordance with an embodiment. The screenshots<figref>1400</figref> and <figref>1450</figref> show a part of the workspace <figref>310</figref>, While the user performs the splitting, or the same part of the workspace <figref>310</figref>After the user has performed the splitting. In the screenshot<figref>1400</figref> the user has three elements of the data nodes <figref>1330</figref> selected: the elements <figref>1331c</figref>. <figref>1331e</figref> and <figref>1331g</figref>, The user pulls these elements of the nodes<figref>1330</figref> to an empty area in the workspace <figref>310</figref>, The cursor icon<figref>1465</figref> indicates the current position of the cursor in the working area as well as the number of elements, which pulls the cursor.
0116In the screenshot <figref>1450</figref> the user has selected the elements at the location of the data node <figref>1430</figref> "Stored". Thus, the node was<figref>1430</figref> and generates a representation of the node, including the data elements <figref>1331c</figref>. <figref>1331e</figref> and <figref>1331g</figref>, Was added to the workspace <figref>310</figref> added. Meanwhile, the data elements<figref>1331c</figref>. <figref>1331e</figref> and <figref>1331g</figref> from the data node <figref>1330</figref> removed as a result of the operation, with the result that the elements <figref>1331k</figref>-m in the node <figref>1330</figref> be visible. but removed the dividing elements of a node In some embodiments, not necessarily members of the original node, but instead copied the elements in a new node.
0117The knot <figref>1435</figref> can now turn to the workflow <figref>320</figref> to be added. He can, for example, back to the node<figref>1340</figref> are connected, thereby requires in effect that the node <figref>1340</figref> its input into two separate data node splits. Or the node<figref>1430</figref> may be the first node in a different and independent sequence of nodes in the workflow <figref>320</figref> be used.
0118the interface <figref>305</figref> is only one example of an interface for carrying out the techniques described herein. Other interfaces may include fewer or additional elements in potentially different arrangements.
5.0. EXEMPLARY OPERATION KNOT WORKFLOW
0119Examples of process nodes, which can be useful for processing genomic data are described below. There may be in fact a lot more types of modules than those listed here. Procedures for other types of data may include some of these process node, but may also or instead include other process nodes that reflect the algorithms for processing of the other kinds of data.
0120In one embodiment, standard modules include a mixing module for mixing volumes of data from multiple nodes and a filter module for filtering a dataset, based on configurable criteria.
0121In one embodiment, a type of process node corresponds to a "DNA into protein Translate 'module. The module accepts input in the form of a sequence. The module uses a locally implemented algorithm to translate the DNA sequence. The module provides outputs in the form of a protein data structure. Exemplary parameters for the configurable module may include, without limitation, include a frame parameter and a complement parameters.
0122In one embodiment, corresponds to a different type of process nodes a "Multiple sequence alignment" module. The module accepts input in the form of a data volume of several sequences, for example, in a multi-FASTA formatted file. The module provides outputs in the form of alignment data, such as in an MSA-alignment file. An overview of the range of an operation may display the alignment of data in a detailed view area using techniques as described in the earlier referenced application, "Computer Graphical User Interface Supporting Aligning Genomic Sequences".
0123In one embodiment, a different type of process node corresponds to a "scan of the protein family (Pfam)" - module. The module accesses a network based application server that performs a hidden Markov module via a protein sequence and the one or more protein families likely calculated based on motifs in the protein sequence. The module accepts input in the form of a protein data structure. The module provides outputs in the form of one or more protein families data structures.
0124In one embodiment, a different kind of operation corresponds to a node "mica" module. The module accesses a network based application server to find genes by an interpolated Markov model. Configurable parameters for the module include, without limitation, a gene coding type, a topology type, a number of input sequences and an output data type, which can be an annotation or sequence.
0125In one embodiment, corresponds to a different type of process nodes a "BLAST" module. The module accesses a network-based application server that searches a sequence query using various libraries of genomic data for genes. The module provides information about matching results back.
0126In one embodiment, a different kind of operation corresponds to a node "FASTA sequence reading device" module that converts a FASTA data structure in a protein or a DNA sequence by means of a local application.
0127In one embodiment, a different kind of operation corresponds to a node "UniProt" module. The engine queries online UniProt database to obtain information on input protein sequences or objects. Configurable parameters for the module include, without limitation, an organism parameters, a gene ontology (GO) parameter, a verification parameter and a Prosite parameters.
0128In one embodiment, a different kind of operation corresponds to a node "PubMed" module. The engine queries online PubMed database for all items that match the input data. include configurable parameters for the module, without limitation, an ID parameter and a verification parameter.
0129In one embodiment, various workflow node types that are supported by the system described herein, without limitation, include nodes that represent one or more of the following: source functions, the data from one or more sources using various query and / or scraping techniques recall; Aggregate, about summary and average results of data; Filter functions, select the sub-sets of sets of biological objects, based on various match criteria and / or thresholds; Sequence Distribution functions that select a portion of a sequence or an alignment and use this section as a new data object; Comparison functions to compare two or more biological objects or sets of objects, based on one or more specified metrics, and determine whether the differences are statistically significant; Comparison functions compare the patient cohorts; Conversion and modification functions; Sequence alignment generating functions; Sequence alignment analysis functions; Forecast functions Make a sequence predict that are for annotations or other features of interest; Annotation to generate annotations automatically; Annotationssuchfunktionen; Functions for the processing of natural languages Publications; Search tools to identify diseases that are associated with specific Genobjekten; and memory functions that store various annotations or other issues in different locations and make available to other users spending.
0130In one embodiment, another exemplary modules that can be connected to different workflow node types, without limitation, modules that implement the following types of analysis: alleles tests, genotype frequency tests, Hardy-Weinberg equilibrium tests, rates of missing genotypes, inbreeding tests, Identity-by-state and Identity-by-Descent statistics for individuals and pairs of individuals, non-Mendelian transmission in family data, complete linkage Hierarchical clustering, multidimensional scaling analysis to visualize substructures significance test of whether two individuals of the same population belong limited cluster solutions to phenotype, cluster size and / or external match criteria, subsequent association analyzes that depend on cluster solutions, standard alleles tests, exact tests according to Fisher, Cochran Armitage trend test, shell-Haenszel and Breslow-Day test for stratified sampling, dominant / rezessiv- and General model tests, model comparison tests (eg. B. general vs. multiplicative model), family-based association tests, about transmission disequilibrium tests or SIBSHIP tests, quantitative characteristics, associations and interactions, association tests that depend on one or more single nucleotide polymorphisms ( "SNPs"), asymptotic and empirical P values , flexible clustered Permutationsschemata, analysis of genotype-probability data and factional alleles counts (after imputation), related haplotype tests, fallen / remote control offers transmission disequilibrium test the association on the probabilistic haplotype phase proxy association process for the investigation of individual SNP associations in their local haplotypischen context Imputationsheuristiken not to test typed SNPs when a Referenzpaneel is given, Joint-SNP and copy number variation ( "CNV") - tests for copy number variants, filtering and summary procedure for segmental (rare) CNV data, case / control comparison tests for global CNV properties, permutationsbasierte association method for identifying specific loci, gene-based association tests, screening for epistasis, gene-environment interaction with uniform and bisected environments and / or Fixed Effects and random-effects models.
6.0. EXEMPLARY USE CASES
0131The following examples show how a user can use a workflow to simplify various objectives associated with genomic data. The examples are for explanation and are not limiting the types of destinations to which the workflow can be applied. There are of course many other types of workflows that are not described below, including, without limitation, Procedures for epigenetics, impact of copy number variations, evolutionary biology and non-coded RNA analysis.
6.1. Gen ANNOTATION
0132One use for the operations described here, in the solution of problems in gene annotation. Strains of Lactobacillus acidophilus that are often in probiotics and possible vaccination vectors of interest, sometimes have, for example, a protein of the surface layer to adhere to cells. A researcher can sequence a new strain of L. acidophilus. When the researchers aliniert the new strain with the reference sequences, finds the researchers found that the new strain, the SLPA gene is missing, however, he has an unknown insertion. The researcher decides that the new strain can be interesting and wants to find out if the new strain is a gene, the likely function of the protein, which encodes the new strain, the biological context of the new strain and how does the master to proteins, which are already known.
0133An example workflow, to help achieve these goals may be as follows. A first set of one or more process node loads the correct DNA sequence and metadata, including the source of the sequencing process, the date and quality. A second set of one or more process nodes delivers a MICA tool to predict one or more genes, based on the loaded data. A third set of one or more process nodes executes a multiple sequence alignment, and a multi-sequence comparison, to compare the sequence with a corresponding region of the L. acidophilus reference sequence, whereby a genome object with annotated genes is generated. A fourth set of one or more process nodes translated the genome object in a protein sequence. A fifth set of one or more process nodes says Pfams and GO-termini advance. A sixth set of one or more process node performs BLAST on the protein sequence in order to answer the questions which Pfams and GO terms among the top results are most common, as these terms overlap with those of the predicted protein and how closely related bacteria the top results with L. acidophilus are. A seventh set of one or more process node looks for known routes, which connected the top results, means MetaCyc and EC numbers. An eighth set of one or more process nodes retrieves PubMed data from the top BLAST results. A ninth amount of one orseveral process nodes found genes and characteristics that are upstream and downstream of the insertion, and determines what are the functions of these bodies. A tenth amount of one or more process nodes delivers a Merkmals- / Annotationsmodul to compare the gene with the PubMed annotations BLAST results to identify unique characteristics of the gene. An eleventh set of one or more process node adds appropriate annotations that relate to the unique characteristics, as annotations for the gene and linked the annotations back to the genome.
0134In one embodiment, different processes may require human intervention to identify key data points before proceeding to the next node. In one embodiment, the workflow is fully automated, with no human intervention. In one embodiment, such a workflow for re-use are stored. Next time, when the researchers discovered a new strain could run the same workflow researchers with respect to the new strain, easily modified by the entry of the original workflow.
6.2. SEQUENCE STRUCTURE FUNCTION DISEASE
0135Another use of the work processes that are described here, is to solve problems with sequence-structure-function disease with respect to a gene. One researcher, for example, create a workflow that answers questions like what implications it in a gene are of the polymorphisms of how and when the gene is expressed, what implications there of the polymorphisms in the protein that it encodes to with interact its cofactor, and how these are connected with the implications of the role of the gene in the disease.
0136An example workflow in order to contribute to achieving these goals can be as follows: A first set of one or more process node performs a multi-sequence alignment of the variants of the gene. A second set of one or more process nodes used recorded observation of microarray expression data under different experimental conditions in order to identify patterns of activity changed. The identification method may include multiple hypotheses's t-test and / or other algorithms, to identify mutations statistically correlated with expression changes under one or more conditions. A third set of one or more process node adds a history of annotated features, such regulatory sequences that are to be compared with the experimental expression data. A fourth set of one or more process node performs a multi-sequence alignment of the variants of the protein. A fifth set of one or more process node performs a recorded observation of the activity level, such as the binding affinity. A sixth set of one or more process node examines individual assays in a tabular view module to determine the sequence of the variants according to binding affinity.
0137A seventh set of one or more process node uses a TreeView module to determine how the binding affinity can be influenced by amino acid mutations, including the prediction of important interactions (H bonds, pi-pi interactions, steric interactions). An eighth set of one or more process nodes examined in PubChem to additional relevant assays. A ninth set of one or more process node looks for the ways in which is located the protein. A tenth set of one or more process node searches in PubMed for elements that are connected to both the gene and a disease. An eleventh set of one or more process node imports the other components of the biological pathway, which is believed that it connects gene and disease, Uniprot and / or other databases, and records the biological pathway on.
6.3. PROTEIN DESIGN
0138Another use for the workflow described here is in Resolving Problems Protein Design. A researcher may want to design, for example, a set of candidate proteins to perform a specific chemical function as tyrosine decarboxylase. The researchers will make these proteins produce and then try in a bacterium which lacks this activity.
0139An example workflow, which helps to achieve these objectives can be as follows. A first set of one or more process nodes examined in Uniprot for proteins with Pfam PF00282 (pyridoxal-dependent decarboxylase). A second set of one or more process nodes examined in Pfam PF00282 by and combines the results with the Uniprot results. A third set of one or more process nodes filters BLAST on this amount of protein sequences against itself to generate all possible pairwise BLAST comparisons. A fourth set of one or more process nodes are clustered the results, based on theBLAST scores. A fifth set of one or more process nodes considered the annotations of each cluster and determines whether there is more than one cluster, which is annotated with tyrosine decarboxylase activity. A sixth set of one or more process nodes aliniert the sequences of each of the clusters, annotated Y decarb activity. The alignment achieved two objectives. First, the alignment allows a comparison of the conserved regions within and between alignments. Second, the alignment grouped subsets of alinierten sequences visually or algorithmically, based on similarity. A seventh set of one or more process node generates a set of candidate proteins that far are representative of the alignments that the candidate proteins have a common sequence in the conserved regions. An eighth set of one or more process node considers the bacteria producing the BLAST result proteins and determine which bacteria is the test bacterium of the researcher most similar, based on phylogenetic information. This may include that the Y-decarb is named candidatel by a bacterium that is added to the list of candidate candidatel proteins and the non-conserved regions of the other candidates having the sequence of candidatel be filled. A ninth set of one or more process nodes examined in PubMed or other databases by comparing the niche and metabolism 'of the two species. A tenth set of one or more process nodes examined in PubMed or other databases to find out what is known about candidatel in the other bacterium beyond. An eleventh set of one or more process nodes analyzes the annotated features of candidatel to investigate whether any areas to be disturbed in the other candidates. A twelfth set of one or more process nodes analyzes the structure of candidate 1 and the alignment of each of the other candidates with candidate 1 to determine where the changes in the sequence can influence the structure. This may involve concatenating the candidate sequence in the structure of candidatel or both concatenating to the closest existing PDB-structure as well as the detection of differences comprise. A thirteenth set of one or more process node exports the sequences of the candidate proteins in a database for subsequent referencing.
6.4. ASSOCIATION STUDY THROUGHOUT THE GENOME
0140Another use for workflows which are described here, is in association studies in the whole genome. One researcher, for example, be curious which SNPs (if any) are connected in an amount of genomes with the presence of disease. The researcher has a lot of individuals at different stages of the disease, which is measured by a biomarker concentration. For all individuals of the genotype by a SNP chip with 1 million SNPs is created. A quality control has been carried out.
0141An example workflow to help to realize the association studies throughout the genome looks like this: A first set of one or more process nodes used Plink, specially Trat and / or R modules to calculate summary statistics, including Allele- frequency and SNP frequency. A second set of one or more process node uses these modules to adjust the statistics to include population stratification (ie distortion due to ancestors / relatives in the case or the control group), where identity-by-state (IBS) - or multidimensional scaling is used. A third set of one or more process node uses these modules to perform a plurality of association tests to determine how the genotype is associated with diseases. The tests include the Fisher exact test, chi-square, correlation and regression. Specifically, the following analyzes can be performed: Basic alleles (as each allele is linked), gentypische tests (such as each pair is linked alleles), additive model (was the presence of two copies of one allele, over no, the dual effect, as the presence of a copy, while there was not), the dominant model (at least one of the child alleles vs. no), recessive model (two minor alleles compared with one or no). A fourth set of one or more process node uses a Manhattan-Plot module, possibly with a ggplot2-R-packet to produce a plot of the P-values of all SNPs along the genomic axis. A fifth set of one or more process node extracts SNPs according to the multi-hypotheses correction, with a P value less than a threshold (for example, p = 10 ^ -. 8). A sixth set of one or more process node generates multiple sequence alignments of the portion of each of the SNPs. The multi-sequence alignments are ordered by case / control group. A seventh set of one or more process nodes examined in dbSNP or other databases to the SNPs to determine whether they are linked to anything else. An eighth set of one or more process node searches for other annotations in these areas to form functional hypotheses.
7.0. HARDWARE OVERVIEW
0142In one embodiment, the techniques described herein may be implemented by one or more special-computing devices. The special computing devices may be hard-wired to perform the techniques, or may be digital electronic devices include, as one or more application specific integrated circuits (ASICs) or field programmable logic arrays (FPGAs) that are permanently programmed to perform the techniques, or may comprise one or more general-purpose hardware processors programmed to perform the techniques in response to program instructions in firmware, in memory, in other storage devices, or a combination thereof. Such special-computing devices may also combine custom hard-wired logic, ASICs or FPGAs with custom programming to perform the techniques. The specialty computing devices may be desktop computer systems, portable computer systems, handheld devices, network devices or any other device, hardwired and / or includes program logic to implement the techniques.
0143<figref>15</figref> For example, a block diagram a computer system <figref>1500</figref> shows, upon which an embodiment of the invention may be implemented. The computer system<figref>1500</figref> includes a bus <figref>1502</figref> or other communication mechanism for transferring information, and a hardware processor <figref>1504</figref>Coupled to the bus <figref>1502</figref> is connected to process information. The hardware processor<figref>1504</figref> can be for example a general purpose microprocessor.
0144The computer system <figref>1500</figref> also includes a main memory <figref>1506</figref>, Such as memory (RAM) or other dynamic storage device, coupled to the bus <figref>1502</figref> is connected to store information and commands from the processor <figref>1504</figref> to be executed. The main memory<figref>1506</figref> can also be used to store temporary variables or other intermediate information during execution of instructions by the processor <figref>1504</figref> to be executed. Such instructions, when stored in non-volatile storage media to which the processor<figref>1504</figref> can access, making the computer system <figref>1500</figref> to a special device that is adapted to perform the operations specified in the instructions.
0145The computer system <figref>1500</figref> further comprising a read only memory (ROM) <figref>1508</figref> or other static storage device coupled to bus <figref>1502</figref> is connected to static information and instructions for processor <figref>1504</figref> save. A memory device<figref>1510</figref>Such as a magnetic disk or optical disk, is provided and the bus <figref>1502</figref> connected to storing information and instructions.
0146The computer system <figref>1500</figref> can via the bus <figref>1502</figref> with a display <figref>1512</figref> be connected, such as a picture tube (CRT), for displaying information to a user. An input device<figref>1514</figref>Comprising alphanumeric and other keys, is coupled to the bus <figref>1502</figref> connected to information and select commands to the processor <figref>1504</figref> transferred to.
0147Another type of user input device is cursor control <figref>1516</figref>, Such as a mouse, a trackball, or cursor direction keys for communicating direction information and select commands to the processor <figref>1504</figref> and to transmit the movement of the cursor on the display <figref>1512</figref> control. The input device typically has two degrees of freedom in two axes (x z. B.) to provide a first axis and a second axis (z. B. y) that allow the device positions in a plane.
0148The computer system <figref>1500</figref> , the techniques described herein using customized hard-wired logic, one or more ASICs implement or FPGAs, firmware and / or program logic, which together with the computer system, the computer system <figref>1500</figref> cause or program so that it becomes a specific device. According to one embodiment, the techniques described herein by the computer system<figref>1500</figref> performed in response to the processor <figref>1504</figref> one or more sequences of one or more instructions executing that in main memory <figref>1506</figref> are included. Such commands can be in the main memory<figref>1506</figref> be read from another storage medium, such as the storage device <figref>1510</figref>, Executing the sequence of instructions in the main memory<figref>1506</figref> are included, causes the processor <figref>1504</figref>To perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.
0149The term "storage medium" as used herein, refers to any non-volatile medium, the data and / or storing instructions that cause a device to operate in a specific way. Such storage media may comprise non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device<figref>1510</figref>, Volatile memories include dynamic memory, such as the main memory<figref>1506</figref>, Common forms of storage media include, for example, floppy disks, floppy disks, hard drives, solid state drives, magnetic tape, any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, RAM, PROM and EPROM, Flash -EPROM, NVRAM, any other memory chip or any other cassette. The terms computer-readable medium and storage medium, as used herein, include non-volatile media, and transmission media.
0150Storage media differ from transmission media, but can be used with them. Transmission media help to transfer information between storage media. Transmission media include, for example, coaxial cables, copper wire and fiber, including cable, the bus to<figref>1502</figref> form. Transmission media can also take the form of acoustic or light waves, such as those generated during radio and infra-red data communications.
0151Various types of media may be involved in carrying one or more sequences of one or more instructions to processor <figref>1504</figref> to transmit for execution. The instructions may be stored on a magnetic disk or a solid-state drive of a remote computer, for example, initially. The remote computer can load the instructions into its dynamic memory and transmit the instructions over a telephone line using a modem. A modem in the vicinity of the computer system<figref>1500</figref> is, can receive the data on the telephone line and use an infrared transmitter to convert the data to an infrared signal. An infrared receiver can receive the data transmitted in the infrared signal and appropriate circuitry can use the data on the bus<figref>1502</figref> lay. The bus<figref>1502</figref> transfers the data to the main memory <figref>1506</figref>From which the processor <figref>1504</figref> retrieves and executes the instructions. The commands from the main memory<figref>1506</figref> be received, can optionally in the storage device <figref>1510</figref> be stored either before or after it by the processor <figref>1504</figref> be executed.
0152The computer system <figref>1500</figref> also includes a communication interface <figref>1518</figref>That the bus <figref>1502</figref> connected is. The communication interface<figref>1518</figref> provides a two-way data communication link with a network connection <figref>1520</figref> prepared with a local area network <figref>1522</figref> connected is. The communication interface<figref>1518</figref> For example, an ISDN card, a cable modem, a satellite modem, or a modem be to prepare a data communication connection to a corresponding type of telephone line. As another example, the communication interface<figref>1518</figref> be a LAN card to establish a data communication connection to a compatible LAN. Wireless links may also be implemented. In such an implementation, transmits and receives the communications interface<figref>1518</figref> electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.
0153The network connection <figref>1520</figref> typically provides data communication through one or more networks to other data devices ready. The network connection<figref>1520</figref> For example, a connection over a local network <figref>1522</figref> to a host computer <figref>1524</figref> manufacture or to data equipment operated by an Internet Service Provider (ISP) <figref>1526</figref> operate. The ISP<figref>1526</figref> in turn provides data communication services through the world wide packet data communication network, now commonly referred to as "Internet" <figref>1528</figref> referred to as. The local network<figref>1522</figref> and the Internet <figref>1528</figref> both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on the network link<figref>1520</figref> and the communication interface <figref>1518</figref>That the digital data to and from computer system <figref>1500</figref> transmitted, are example forms of transmission media.
0154The computer system <figref>1500</figref> can via the / the network (s), the network connection <figref>1520</figref> and the communication interface <figref>1518</figref> send messages and receive data, including program code. In the Internet example, a server<figref>1530</figref> a requested code for an application program through Internet <figref>1518</figref>, The ISP <figref>1526</figref>, The local area network <figref>1522</figref> and the communication interface <figref>1518</figref> transfer.
0155The received code may by the processor <figref>1504</figref> be carried out as it was received, and / or in the storage device <figref>1510</figref> or other non-volatile memory is stored, to be executed later.
0156In the foregoing specification embodiments of the invention with respect to many specific details have been described, which may vary from implementation to implementation. The specification and drawings are therefore intended to be understood in a descriptive and not a limiting sense. The sole and exclusive definition of the scope of the invention and which is regarded by the parties as the scope of the invention, is the literal and equivalent scope of the amount of claims that are specified in this application, in the specific form in which such claims are given , including any subsequent correction.
Contents22
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10534595B1 | Cited by | United States of America | Applicant |
| US10871887B2 | Cited by | United States of America | Applicant |
| US11119630B1 | Cited by | United States of America | Applicant |
| US10452673B1 | Cited by | United States of America | Applicant |
| US10838697B2 | Cited by | United States of America | Applicant |
| US10402054B2 | Cited by | United States of America | Applicant |
| US9785773B2 | Cited by | United States of America | Applicant |
| US10706220B2 | Cited by | United States of America | Applicant |
| US11048706B2 | Cited by | United States of America | Applicant |
| US10403011B1 | Cited by | United States of America | Applicant |
| US11004244B2 | Cited by | United States of America | Applicant |
| US10540333B2 | Cited by | United States of America | Applicant |
| US12204845B2 | Cited by | United States of America | Applicant |
| US10157200B2 | Cited by | United States of America | Applicant |
| US10191926B2 | Cited by | United States of America | Applicant |
| US10706434B1 | Cited by | United States of America | Applicant |
| US10877638B2 | Cited by | United States of America | Applicant |
| US9953445B2 | Cited by | United States of America | Applicant |
| US9965534B2 | Cited by | United States of America | Applicant |
| US10545985B2 | Cited by | United States of America | Applicant |
| US9740369B2 | Cited by | United States of America | Applicant |
| US11379407B2 | Cited by | United States of America | Applicant |
| US10318398B2 | Cited by | United States of America | Applicant |
| US10860299B2 | Cited by | United States of America | Applicant |
| US10762291B2 | Cited by | United States of America | Applicant |
| US9621676B2 | Cited by | United States of America | Applicant |
| US9652510B1 | Cited by | United States of America | Applicant |
| US9779525B2 | Cited by | United States of America | Applicant |
| US12554781B2 | Cited by | United States of America | Applicant |
| US11252248B2 | Cited by | United States of America | Applicant |
| US10572496B1 | Cited by | United States of America | Applicant |
| US10795723B2 | Cited by | United States of America | Applicant |
| US10248294B2 | Cited by | United States of America | Applicant |
| US10042524B2 | Cited by | United States of America | Applicant |
| US10853338B2 | Cited by | United States of America | Applicant |
| US9921734B2 | Cited by | United States of America | Applicant |
| US11157951B1 | Cited by | United States of America | Applicant |
| US11138279B1 | Cited by | United States of America | Applicant |
| US10346410B2 | Cited by | United States of America | Applicant |
| US11645250B2 | Cited by | United States of America | Applicant |
| US10444940B2 | Cited by | United States of America | Applicant |
| US9880987B2 | Cited by | United States of America | Applicant |
| US9857958B2 | Cited by | United States of America | Applicant |
| US10824604B1 | Cited by | United States of America | Applicant |
| US11500827B2 | Cited by | United States of America | Applicant |
| US10180929B1 | Cited by | United States of America | Applicant |
| US11934847B2 | Cited by | United States of America | Applicant |
| US10873603B2 | Cited by | United States of America | Applicant |
| US10698938B2 | Cited by | United States of America | Applicant |
| US10296617B1 | Cited by | United States of America | Applicant |
| US12105719B2 | Cited by | United States of America | Applicant |
| US11182204B2 | Cited by | United States of America | Applicant |
| US9886467B2 | Cited by | United States of America | Applicant |
| US10866685B2 | Cited by | United States of America | Applicant |
| US10152306B2 | Cited by | United States of America | Applicant |
| US10976892B2 | Cited by | United States of America | Applicant |
| US10360252B1 | Cited by | United States of America | Applicant |
| US11625529B2 | Cited by | United States of America | Applicant |
| US10554516B1 | Cited by | United States of America | Applicant |
| US10509844B1 | Cited by | United States of America | Applicant |
| US9767172B2 | Cited by | United States of America | Applicant |
| US10360702B2 | Cited by | United States of America | Applicant |
| US9734217B2 | Cited by | United States of America | Applicant |
| US10120545B2 | Cited by | United States of America | Applicant |
| US10356032B2 | Cited by | United States of America | Applicant |
| US9880696B2 | Cited by | United States of America | Applicant |
| US10275778B1 | Cited by | United States of America | Applicant |
| USRE47594E | Cited by | United States of America | Applicant |
| US10102369B2 | Cited by | United States of America | Applicant |
| US12147647B2 | Cited by | United States of America | Applicant |
| US10732803B2 | Cited by | United States of America | Applicant |
| US10180977B2 | Cited by | United States of America | Applicant |
| US10545655B2 | Cited by | United States of America | Applicant |
| US9646396B2 | Cited by | United States of America | Applicant |
| US12079456B2 | Cited by | United States of America | Applicant |
| US10929436B2 | Cited by | United States of America | Applicant |
| US11138180B2 | Cited by | United States of America | Applicant |
| US10223748B2 | Cited by | United States of America | Applicant |
| US11444854B2 | Cited by | United States of America | Applicant |
| US10853352B1 | Cited by | United States of America | Applicant |
| US12353678B2 | Cited by | United States of America | Applicant |
| US11907513B2 | Cited by | United States of America | Applicant |
| US10025834B2 | Cited by | United States of America | Applicant |
| US11392759B1 | Cited by | United States of America | Applicant |
| US9898528B2 | Cited by | United States of America | Applicant |
| US11886382B2 | Cited by | United States of America | Applicant |
| US9619557B2 | Cited by | United States of America | Applicant |
| US10956406B2 | Cited by | United States of America | Applicant |
| US10460602B1 | Cited by | United States of America | Applicant |
| US10127021B1 | Cited by | United States of America | Applicant |
| US10728277B2 | Cited by | United States of America | Applicant |
| US10540061B2 | Cited by | United States of America | Applicant |
| US10795909B1 | Cited by | United States of America | Applicant |
| US10650086B1 | Cited by | United States of America | Applicant |
| US10699071B2 | Cited by | United States of America | Applicant |
| US10261763B2 | Cited by | United States of America | Applicant |
| US11977863B2 | Cited by | United States of America | Applicant |
| US10805321B2 | Cited by | United States of America | Applicant |
| US11244102B2 | Cited by | United States of America | Applicant |
| US9965937B2 | Cited by | United States of America | Applicant |
17 members in 6 offices
Members17
| Document | Office | Kind | |
|---|---|---|---|
| GB201404574D0 | United Kingdom | D0 | |
| CA2845606A1 | Canada | A1 | |
| NL2012437A | Netherlands (Kingdom of the) | A | |
| DE102014103482A1This record | Germany | A1 | |
| US2014282177A1 | United States of America | A1 | |
| AU2014200965A1 | Australia | A1 | |
| GB2517529A | United Kingdom | A | |
| NL2012437B1 | Netherlands (Kingdom of the) | B1 | |
| US9501202B2 | United States of America | B2 | |
| US2017046481A1 | United States of America | A1 | |
| AU2014200965B2 | Australia | B2 | |
| US10431327B2 | United States of America | B2 | |
| US2019371435A1 | United States of America | A1 | |
| CA2845606C | Canada | C | |
| US11074993B2 | United States of America | B2 | |
| US2021383896A1 | United States of America | A1 | |
| US11676686B2 | United States of America | B2 |
10 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Grant decision by examination section/examining divisionR018 | R018 | |
| Response to examination communicationR016 | R016 | |
| Change of applicant/patenteeR081 | R081 | |
| Change of applicant/patenteeR081 | R081 | |
| Change of applicant/patenteeR081 | R081 | |
| Change of representativeR082 | R082 | |
| Amendment of ipc main classPREVIOUS MAIN CLASS: G06F0019260000R079 | R079 | |
| Request for examination validly filedR012 | R012 | |
| Request for examination validly filedR012 | R012 | |
| Amendment of/additions to inventor(s)R083 | R083 |
Numbers
- Publication
- 102014103482
- Application
- 10103482
Titles2
- German
- Graphische Benutzeroberfläche eines Computers mit genomischem Arbeitsablauf
- English
- The graphical user interface of a computer with genomic workflow
Classification
- CPC, 10
- G16B45/00
- G16B20/00
- G06Q10/10
- G06F16/951
- G16B50/00
- G16B50/10
- G06F16/953
- G06F3/0481
- G06F3/0482
- G06F3/04847
- IPC, 3
- G06F19 26
- G16B45 00
- G16B50 10