CA2845606C

Computer graphical user interface with genomic workflow

Abstract

Methods and computer apparatuses are disclosed for processing genomic data in at least partially automated workflows of modules. A method comprises: specifying a source from which nucleic acid sequence(s) are to be obtained; selecting module(s) for processing data, including at least one module for processing the one or more nucleic acid sequences; presenting, in a graphical user interface, graphical components representing the source and the module(s) as nodes within a workspace; receiving, via the graphical user interface, inputs arranging the source and the module(s) as a workflow comprising a series of nodes, the series indicating, for each particular module, that output from one of the source or another particular module is to be input into the particular module; generating an output for the workflow based upon the nucleic acid sequence(s) by processing each module in an order indicated by the series.

CA2845606C, drawing sheet 1
Sheet 1 of 15

Term

Projected expiry 11 March 2034.

  1. Priority and filed
  2. Granted
  3. Today
  4. Projected expiry

40 claims: 4 independent, 36 dependent

  1. 1
    A method implemented as a set of stored instructions when executed by a computer processor for:receiving a first input specifying a source from which one or more nucleic acid sequences are to be obtained, the one or more nucleic acid sequences being converted by a data conversion component into converted data in a data structure defined by an ontology associated with a workflow;receiving one or more second inputs selecting one or more modules for processing data, including at least one module for processing the converted data, and when the at least one module processes the converted data by sending at least a portion of the converted data to one or more external tools, the one or more external tools processing the portion of the converted data and returning processed data, the processed data being converted by the data conversion component to the data structure defined by the ontology;presenting, in a graphical user interface, graphical components representing the source and the one or more modules as nodes within a workspace;re ceiving, via the graphical user interface, one or more third inputs arranging the source and the one or more modules as the workflow comprising a series of nodes, the series indicating, for each particular module of the selected modules, that output from one of the source or another particular module is to be input into the particular module;generating an output for the workflow, wherein the output comprises a set of one or more items of genomic data that are based upon the one or more nucleic acid sequences that are processed by each module of the one or more modules in an order indicated by the series;generating a first data node from the output, the first data node comprising the set of one or more items of genomic data, the first data node linked to a last module in the series;re ceiving, via the graphical user interface, fourth input that selects a subset of one or more items of genomic data from the set of one or more items of genomic data in the first data node;re ceiving, via the graphical user interface, fifth input that moves the subset of one or more items of genomic data to a location on the graphical user interface not associated with the -38Date Reçue/Date Received 2020-04-21 first data node;generating a second data node comprising the subset of one or more items of genomic data, wherein the output for the workflow is reconfigured to generate multiple data nodes, one corresponding to the first data node comprising the set of one or more items of genomic data other than the subset of one or more items of genomic data, and another corresponding to the second data node comprising, the subset of one or more items of genomic data;wherein the method is performed by one or more computing devices.
  2. 11
    One or more non-transitory computer-readable media having stored instructions that, when executed by one or more computing devices, cause:-40Date Reçue/Date Received 2020-04-21 receiving a first input specifying a source from which one or more nucleic acid sequences are to be obtained, the one or more nucleic acid sequences being converted by a data conversion component into converted data in a data structure defined by an ontology associated with a workflow;receiving one or more second inputs selecting one or more modules for processing data, including at least one module for processing the converted data, and when the at least one module processes the converted data by sending at least a portion of the converted data to one or more external tools, the one or more external tools processing the portion of the converted data and returning processed data, the processed data being converted by the data conversion component to the data structure defined by the ontology;presenting, in a graphical user interface, graphical components representing the source and the one or more modules as nodes within a workspace;receiving, via the graphical user interface, one or more third inputs arranging the source and the one or more modules as the workflow comprising a series of nodes, the series indicating, for each particular module of the selected modules, that output from one of the source or another particular module is to be input into the particular module;generating an output for the workflow, wherein the output comprises a set of one or more items of genomic data that are based upon the one or more nucleic acid sequences that are processed by each module of the one or more modules in an order indicated by the series;generating a first data node from the output, the first data node comprising the set of one or more items of genomic data, the first data node linked to a last module in the series;receiving, via the graphical user interface, fourth input that selects a subset of one or more items of genomic data from the set of one or more items of genomic data in the first data node;receiving, via the graphical user interface, fifth input that moves the subset of one or more items of genomic data to a location on the graphical user interface not associated with the first data node;generating a second data node comprising the subset of one or more items of genomic data, wherein the output for the workflow is reconfigured to generate multiple data nodes, one corresponding to the first data node comprising the set of one or more items of genomic data -41Date Reçue/Date Received 2020-04-21 other than the subset of one or more items of genomic data, and another corresponding to the second data node comprising, the subset of one or more items of genomic data.
  3. 21
    A method implemented as a set of stored instructions when executed by a computer processor for:presenting, in a graphical user interface, graphical components representing a source from which one or more nucleic acid sequences are to be obtained and one or more sets of -43Date Reçue/Date Received 2020-04-21 instructions for processing data, including at least one set of instructions for processing the one or more nucleic acid sequences, wherein the source and the one or more sets of instructions are represented as nodes within a workspace;wherein the source and the one or more sets of instructions are arranged as a workflow comprising a series of nodes, the series of nodes indicating, for each particular set of instructions of the one or more sets of instructions, that output from one of the source or another particular set of instructions is to be input into the particular set of instructions;generating an output for the workflow, wherein the output comprises a set of one or more items of genomic data that are based upon the one or more nucleic acid sequences that are processed by each set of instructions of the one or more sets of instructions in an order indicated by the series of nodes;generating a first data node from the output, the first data node comprising the set of one or more items of genomic data, the first data node linked to a last set of instructions in the series;receiving, via the graphical user interface, a first input that selects a subset of one or more items of genomic data from the set of one or more items of genomic data in the first data node;receiving, via the graphical user interface, a second input that moves the subset of one or more items of genomic data to a location on the graphical user interface not associated with the first data node;generating a second data node comprising the subset of one or more items of genomic data, wherein the output for the workflow is reconfigured to generate multiple data nodes;and wherein the method is performed by one or more computing devices.
  4. 22
    The method of Claim 21, wherein each set of instructions of the one or more sets of instructions generates output that conforms to an ontology defining data structures that represent genomic data, the data structures representing at least all of:sequences, protein objects, alignment objects, annotations, and publications.
  5. 23
    The method of Claim 21, further comprising:receiving, via the graphical user interface, third input selecting a particular set of instructions to process the first data node;-44Date Reçue/Date Received 2020-04-21 adding the particular set of instructions to the end of the series;and generating third output for the workflow based upon the one or more nucleic acid sequences by processing each set of instructions in the series, including the particular set of instructions, in the order indicated by the series.
  6. 24
    The method of Claim 21, wherein the one or more sets of instructions comprises at least two sets of instructions, wherein generating the output for the workflow comprises using output from the source as input to a first set of instructions, and using output from the first set of instructions as input to a second set of instructions.
  7. 25
    The method of Claim 21, wherein the at least one set of instructions is configured to process the one or more nucleic acid sequences by communicating with at least one of an external web server or an external database server.
  8. 26
    The method of Claim 21, further comprising:saving workflow data describing the series;causing the workflow data to be shared with multiple users;subsequently reconstructing the series in a second graphical user interface based on the workflow data;receiving sixth input, via the second graphical user interface, modifying the series to include one or more additional sets of instructions;and generating second output based upon the one or more nucleic acid sequences by processing each set of instructions in the series, including the one or more additional sets of instructions, in an order indicated by the series.
  9. 27
    The method of Claim 21, wherein the one or more sets of instructions include a first set of instructions that generates first output based upon the source, and a second set of instructions that merges the first output with second output from a third set of instructions that is not in the series, wherein the source, first set of instructions, second set of instructions, and third set of instructions are all nodes within a workflow.
  10. 28
    The method of Claim 21, further comprising presenting controls for selecting the one or more sets of instructions, wherein the controls include at least:a first control for selecting a -45Date Reçue/Date Received 2020-04-21 first set of instructions that searches for publications in an online database based on genomic data, a second control for selecting a second set of instructions that outputs a sequence alignment for multiple sequences, and a third control for selecting a third set of instructions that identifies protein families for a nucleic acid sequence.
  11. 29
    The method of Claim 21, further comprising presenting visual feedback while a first node is selected that indicates that genomic data output from the first node can be linked as input to a second node.
  12. 30
    The method of Claim 21, wherein the one or more sets of instructions comprises at least two sets of instructions, wherein processing each set of instructions of the one or more sets of instructions in an order indicated by the series comprises automatically processing each set of instructions, without human intervention between beginning processing of a first set of instructions in the series and generating the output by concluding processing of a last set of instructions in the series.
  13. 31
    One or more non-transitory computer-readable media having stored instructions that, when executed by one or more computing devices, cause:presenting, in a graphical user interface, graphical components representing a source from which one or more nucleic acid sequences are to be obtained and one or more sets of instructions for processing data, including at least one set of instructions for processing the one or more nucleic acid sequences, wherein the source and the one or more sets of instructions are represented as nodes within a workspace;wherein the source and the one or more sets of instructions are arranged as a workflow comprising a series of nodes, the series of nodes indicating, for each particular set of instructions of the one or more sets of instructions, that output from one of the source or another particular set of instructions is to be input into the particular set of instructions;generating an output for the workflow, wherein the output comprises a set of one or more items of genomic data that are based upon the one or more nucleic acid sequences that are processed by each set of instructions of the one or more sets of instructions in an order indicated by the series of nodes;generating a first data node from the output, the first data node comprising the set of -46Date Reçue/Date Received 2020-04-21 one or more items of genomic data, the first data node linked to a last set of instructions in the series;receiving, via the graphical user interface, a first input that selects a subset of one or more items of genomic data from the set of one or more items of genomic data in the first data node;receiving, via the graphical user interface, a second input that moves the subset of one or more items of genomic data to a location on the graphical user interface not associated with the first data node;generating a second data node comprising the subset of one or more items of genomic data, wherein the output for the workflow is reconfigured to generate multiple data nodes.
  14. 32
    The one or more non-transitory computer-readable media of Claim 31, wherein each set of instructions of the one or more sets of instructions generates output that conforms to an ontology defining data structures that represent genomic data, the data structures representing at least all of:sequences, protein objects, alignment objects, annotations, and publications.
  15. 33
    The one or more non-transitory computer-readable media of Claim 31, wherein the instructions, when executed by the one or more computing devices, further cause:receiving, via the graphical user interface, third input selecting a particular set of instructions to process the first data node;adding the particular set of instructions to the end of the series;and generating third output for the workflow based upon the one or more nucleic acid sequences by processing each set of instructions in the series, including the particular set of instructions, in the order indicated by the series.
  16. 34
    The one or more non-transitory computer-readable media of Claim 31, wherein the one or more sets of instructions comprises at least two sets of instructions, wherein generating the output for the workflow comprises using output from the source as input to a first set of instructions, and using output from the first set of instructions as input to a second set of instructions. -47Date Reçue/Date Received 2020-04-21
  17. 35
    The one or more non-transitory computer-readable media of Claim 31, wherein the at least one set of instructions is configured to process the one or more nucleic acid sequences by communicating with at least one of an external web server or an external database server.
  18. 36
    The one or more non-transitory computer-readable media of Claim 31, wherein the instructions, when executed by the one or more computing devices, further cause:saving workflow data describing the series;causing the workflow data to be shared with multiple users;subsequently reconstructing the series in a second graphical user interface based on the workflow data;receiving sixth input, via the second graphical user interface, modifying the series to include one or more additional sets of instructions;and generating second output based upon the one or more nucleic acid sequences by processing each set of instructions in the series, including the one or more additional sets of instructions, in an order indicated by the series.
  19. 37
    The one or more non-transitoiy computer-readable media of Claim 31, wherein the one or more sets of instructions include a first set of instructions that generates first output based upon the source, and a second set of instructions that merges the first output with second output from a third set of instructions that is not in the series, wherein the source, first set of instructions, second set of instructions, and third set of instructions are all nodes within a workflow.
  20. 38
    The one or more non-transitory computer-readable media of Claim 31, wherein the instructions, when executed by the one or more computing devices, further cause presenting controls for selecting the one or more sets of instructions, wherein the controls include at least:a first control for selecting a first set of instructions that searches for publications in an online database based on genomic data, a second control for selecting a second set of instructions that outputs a sequence alignment for multiple sequences, and a third control for selecting a third set of instructions that identifies protein families for a nucleic acid sequence.
  21. 39
    The one or more non-transitory computer-readable media of Claim 31, wherein the instructions, when executed by the one or more computing devices, further cause presenting -48Date Reçue/Date Received 2020-04-21 visual feedback while a first node is selected that indicates that genomic data output from the first node can be linked as input to a second node.
  22. 40
    The one or more non-transitory computer-readable media of Claim 31, wherein the one or more sets of instructions comprises at least two sets of instructions, wherein processing each set of instructions of the one or more sets of instructions in an order indicated by the series comprises automatically processing each set of instructions, without human intervention between beginning processing of a first set of instructions in the series and generating the output by concluding processing of a last set of instructions in the series.
Independent claims22