EP3673370B1

System and method for distributed resource requirement and allocation

Abstract

This record has no abstract on file.

EP3673370B1, drawing sheet 1
Sheet 1 of 20

Term

Projected expiry 19 September 2038.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Projected expiry

15 claims: 3 independent, 12 dependent

  1. 1
    A computer-implemented method for allocating resources of a distributed compute cluster for execution of a distributed compute job in the distributed compute cluster, the computer-implemented method comprising:obtaining (202) duration information from historical runs indicative of an amount of time taken by each of one or more tasks of a distributed compute phase of the distributed compute job to execute;sorting (204) the one or more tasks into one or more groups based on the duration information and determining (206) a resource requirement for each of the one or more groups;and determining (208) , based on the resource requirement for each of the one or more groups, a time-varying allocation of the resources of the distributed compute cluster for the phase.
  2. 8
    The computer-implemented method of any one of claims 1 to 7, wherein the one or more tasks are to be performed in one or more compute slots and further wherein sorting the one or more tasks into one or more groups comprises determining a partition of the one or more groups that meets a desired number for the one or more groups and minimizes a total area of the one or more compute slots.
  3. 9
    A computing device (1400) for allocating resources of a distributed compute cluster for execution of a distributed compute job in the distributed compute cluster, the computing device (1400) comprising:at least one processing unit (1402) ;and a non-transitory memory ( 1404 ) communicatively coupled to the at least one processing unit (1402) and storing computer-readable program instructions (1406) executable by the at least one processing unit (1402) for: obtaining duration information from historical runs indicative of an amount of time taken by each of one or more tasks of a distributed compute phase of the distributed compute job to execute, sorting the one or more tasks into one or more groups based on the duration information and determining a resource requirement for each of the one or more groups, and determining, based on the resource requirement for each of the one or more groups, a time-varying allocation of the resources of the distributed compute cluster for the phase.