US20260133940A1

Updates and Deletes in Retrieval-Access Generation Ingestion Versioning

Claim Score by NHIP

Read claim 15, the broadest

Abstract

A system can, based on a retrieval-augmented generation (RAG) process ingesting data, query a search system to identify at least one first portion of the data that has at least one third generation identifier that is greater than second generation identifiers in a checkpoint; create second chunks from the at least one first portion of the data; in response to second chunks being duplicates of first chunks, based on hash values, store first identifiers of the second chunks; remove the stored chunk identifiers from second identifiers of chunks that correspond to chunks that existed in a RAG system prior to the ingesting, to produce stale chunk identifiers; remove, from the RAG system, third chunks that are identified by the stale chunk identifiers, and, in response to the second chunks being determined to be unique relative to the first chunks, store the second chunks in the RAG system.

US20260133940A1, drawing sheet 1
Sheet 1 of 12

Term

18.1 yearsto projected expiry

Projected expiry 13 November 2044, counted from filing; an application has no term until it is granted.

  1. Priority and filed
  2. Published
  3. Today
  4. Projected expiry

20 claims: 3 independent, 17 dependent

  1. 1
    A system, comprising:at least one processor;and at least one memory that stores executable instructions that, when executed by the at least one processor, facilitate performance of operations, comprising: storing a checkpoint that comprises pairs and first hash values, wherein the checkpoint identifies data that has been ingested into a retrieval-augmented generation system, wherein each pair of the pairs comprises an identification of at least some of first data stored in a storage system and a second generation identifier that corresponds to the at least some of the first data, and wherein each first hash value of the first hash values comprises a hash of a first chunk of the at least some of the first data, wherein a group of first chunks comprises the first chunk;and based on executing a retrieval-augmented generation process comprising performance of an iteration of ingesting more data from the storage system and to send the data to be ingested by the retrieval-augmented generation system, wherein the retrieval-augmented generation process is configured to ingest the data via a communications protocol that omits tracking of previously-ingested data, querying a search system to identify at least one first portion of the data that has at least one respective third generation identifier that is greater than the respective second generation identifiers in the checkpoint, wherein the search system stores respective metadata of respective first data from the storage system, and wherein the respective metadata comprises respective first generation identifiers that indicate respective updates to the respective first data, creating second chunks from the at least one first portion of the data, in response to respective second chunks of the second chunks being determined to be duplicates of respective first chunks of the group of first chunks, based on the respective first hash values and respective second hash values of the respective second chunks, storing first identifiers of the respective second chunks, to produce stored chunk identifiers, removing the stored chunk identifiers from second identifiers of chunks that correspond to chunks that existed in the retrieval-augmented generation system prior to the performance of the iteration of the ingesting of the data, to produce stale chunk identifiers, removing, from the retrieval-augmented generation system, third chunks that are identified by the stale chunk identifiers, and in response to the respective second chunks being determined to be unique relative to the respective first chunks, storing the respective second chunks in the retrieval-augmented generation system, wherein the second chunks correspond to portions of the data that are updated relative to what is stored in the retrieval-augmented generation system.
  2. 9
    A method, comprising:based on performance of an iteration of ingesting data from a storage system and to send the data to be ingested by a retrieval-augmented generation system, querying, by a system comprising at least one processor, a search system to identify at least one first portion of the data that has at least one respective third generation identifier that is greater than respective second generation identifiers in a checkpoint that comprises pairs and first hash values, wherein each pair of the pairs comprises an identification of at least some of first data stored in a storage system and a second generation identifier that corresponds to the at least some of the first data, wherein each first hash value of the first hash values comprises a hash of a first chunk of the at least some of the first data, wherein a group of first chunks comprises the first chunk, wherein the search system stores respective metadata of respective first data from the storage system, and wherein the respective metadata comprises respective first generation identifiers that indicate respective updates to the respective first data;creating, by the system, second chunks from the at least one first portion of the data;responsive to respective second chunks of the second chunks being determined to be duplicates of respective first chunks of the group of first chunks, based on the respective first hash values and respective second hash values of the respective second chunks, storing, by the system, first identifiers of the respective second chunks, to produce stored chunk identifiers;removing, by the system, the stored chunk identifiers from second identifiers of chunks that correspond to chunks that existed in the retrieval-augmented generation system prior to the performance of the iteration of the ingesting of the data, to produce stale chunk identifiers;removing, by the system and from the retrieval-augmented generation system, third chunks that are identified by the stale chunk identifiers;and responsive to the respective second chunks being determined to be unique relative to the respective first chunks, storing, by the system, the respective second chunks in the retrieval-augmented generation system.
  3. 15
    Broadest claimClaim Score 22, narrow(NHIP)A non-transitory computer-readable medium comprising instructions that, in response to execution, cause a system comprising at least one processor to perform operations, comprising:based on ingesting data from a storage system and to a retrieval-augmented generation system, querying a search system to identify at least one first portion of the data that has at least one respective third generation identifier that is greater than respective second generation identifiers in a state file that comprises pairs and first hash values, wherein each pair of the pairs comprises an identification of at least some of first data stored in a storage system and a second generation identifier that corresponds to the at least some of the first data, wherein each first hash value of the first hash values comprises a hash of a first chunk of the at least some of the first data, wherein a group of first chunks comprises the first chunk, and wherein the search system stores respective metadata of respective first data from the storage system that comprises respective first generation identifiers that indicate respective updates to the respective first data;creating second chunks from the at least one first portion of the data;where the respective first hash values and respective second hash values of respective second chunks of the second chunks indicate that at least some of the respective second chunks of the second chunks are duplicates of respective first chunks of the first chunks, storing first identifiers of the at least some of the respective second chunks, to produce stored chunk identifiers;removing the stored chunk identifiers from second identifiers of chunks that correspond to chunks that existed in the retrieval-augmented generation system prior to the ingesting of the data, to produce stale chunk identifiers;and removing, from the retrieval-augmented generation system, third chunks that are identified by the stale chunk identifiers.