US11727935B2

Natural language processing for optimized extractive summarization

Summary by NHIP

Optimized Extractive Summarization

The method determines per-party utterance subsets from multi-party transcripts and generates summaries using an unsupervised machine learning model. It selects an optimal summary from eligible candidates based on utility measures derived from similarity constraints linking specific utterances to per-party subsets.

Claim Score by NHIP

Read claim 1, the broadest

Abstract

There is a need for more effective and efficient predictive natural language summarization. This need can be addressed by, for example, solutions for performing predictive natural language summarization using a constrained optimization model. In one example, a method includes identifying one or more per-party utterance subsets in a multi-party call transcript; generating a plurality of eligible extractive summaries that comply with one or more optimization constraints; for each eligible extractive summary of the plurality of eligible extractive summaries, determining an overall summary utility measure; generating the optimal extractive summary based at least in part on each overall summary utility measure for an eligible extractive summary of the plurality of eligible extractive summaries; and performing one or more summary-based actions based at least in part on the optimal extractive summary.

US11727935B2, drawing sheet 1
Sheet 1 of 45

Term

14.8 yearsleft in the term

Expires 24 July 2041, including 221 days of term adjustment.

  1. Priority and filed
  2. Granted
  3. Today
  4. Expires

18 claims: 3 independent, 15 dependent

  1. 1
    Broadest claimClaim Score 16, narrow(NHIP)A computer-implemented method comprising:determining, by one or more processors, one or more per-party utterance subsets from multi-party interaction transcript data, wherein: (1) the multi-party interaction transcript data comprises a plurality of interaction utterances associated with a plurality of interaction parties, and (2) each of the one or more per-party utterance subsets comprises one or more of the plurality of interaction utterances that are associated with one of the plurality of interaction parties;generating, by the one or more processors and using an unsupervised machine learning model, an optimal extractive summary, wherein the optimal extractive summary comprises one of a plurality of eligible extractive summaries that is selected based at least in part on respective overall summary utility measures for each of the plurality of eligible extractive summaries, wherein:(i) each of the plurality of eligible extractive summaries comprises a covered subset of the plurality of interaction utterances that complies with one or more optimization constraints, and the one or more optimization constraints comprise a similarity-based optimization constraint based at least in part on a covered subset for a particular one of the plurality of eligible extractive summaries comprising (a) a particular one of the plurality of interaction utterances that is in a particular one of the one or more per-party utterance subsets, and (b) additional ones of the plurality of interaction utterances, each of which is in any per-party utterance subset of the one or more per-party utterance subsets other than the particular one of the one or more per-party utterance subsets having a threshold-satisfying utterance similarity measure with respect to the particular one of the plurality of interaction utterances, and (ii) each of the respective overall summary utility measures is based at least in part on an information quality measure and a linguistic quality measure for each of the plurality of interaction utterances in the covered subset for the particular one of the plurality of eligible extractive summaries;andinitiating, by the one or more processors, the performance of one or more summary-based actions based at least in part on the optimal extractive summary.
  2. 9
    An apparatus comprising one or more processors and memory including program code, the memory and the program code configured to, with the one or more processors, cause the apparatus to at least:determine one or more per-party utterance subsets from multi-party interaction transcript data, wherein: (1) the multi-party interaction transcript data comprises a plurality of interaction utterances associated with a plurality of interaction parties, and (2) each of the one or more per-party utterance subsets comprises one or more of the plurality of interaction utterances that are associated with one of the plurality of interaction parties;generate, using an unsupervised machine learning model, an optimal extractive summary, wherein the optimal extractive summary comprises one of a plurality of eligible extractive summaries that is selected based at least in part on respective overall summary utility measures for each of the plurality of eligible extractive summaries, wherein: (i) each of the plurality of eligible extractive summaries comprises a covered subset of the plurality of interaction utterances that complies with one or more optimization constraints, and the one or more optimization constraints comprise a similarity-based optimization constraint based at least in part on a covered subset for a particular one of the plurality of eligible extractive summaries comprising (a) a particular one of the plurality of interaction utterances that is in a particular one of the one or more per-party utterance subsets, and (b) additional ones of the plurality of interaction utterances, each of which is in any per-party utterance subset of the one or more per-party utterance subsets other than the particular one of the one or more per-party utterance subsets having a threshold-satisfying utterance similarity measure with respect to the particular one of the plurality of interaction utterances, and(ii) each of the respective overall summary utility measures is based at least in part on an information quality measure and a linguistic quality measure for each of the plurality of interaction utterances in the covered subset for the particular one of the plurality of eligible extractive summaries;andinitiate the performance of one or more summary-based actions based at least in part on the optimal extractive summary.
  3. 17
    A computer program product comprising at least one non-transitory computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions configured to:determine one or more per-party utterance subsets from multi-party interaction transcript data, wherein: (1) the multi-party interaction transcript data comprises a plurality of interaction utterances associated with a plurality of interaction parties, and (2) each of the one or more per-party utterance subsets comprises one or more of the plurality of interaction utterances that are associated with one of the plurality of interaction parties;generate, using an unsupervised machine learning model, an optimal extractive summary, wherein the optimal extractive summary comprises one of a plurality of eligible extractive summaries that is selected based at least in part on respective overall summary utility measures for each of the plurality of eligible extractive summaries, wherein: (i) each of the plurality of eligible extractive summaries comprises a covered subset of the plurality of interaction utterances that complies with one or more optimization constraints, and the one or more optimization constraints comprise a similarity-based optimization constraint based at least in part on a covered subset for a particular one of the p lurality of eligible extractive summaries comprising (a) a particular one of the plurality of interaction utterances that is in a particular one of the one or more per-party utterance subsets, and (b) additional ones of the plurality of interaction utterances, each of which is in any per-party utterance subset of the one or more per-party utterance subsets other than the particular one of the one or more per-party utterance subsets having a threshold-satisfying utterance similarity measure with respect to the particular one of the plurality of interaction utterances, and(ii) each of the respective overall summary utility measures is based at least in part on an information quality measure and a linguistic quality measure for each of the plurality of interaction utterances in the covered subset for the particular one of the plurality of eligible extractive summaries;and initiate the performance of one or more summary-based actions based at least in part on the optimal extractive summary.