Backward compatibility testing of software in a mode that disrupts timing
Abstract
The apparatus may operate in a timing test mode in which the apparatus is configured to disrupt processing occurring on the one or more processors while the application is running with the one or more processors. timing. The application may test for errors while the device is running in the timing test mode.

Term
10.1 yearsleft in the term
Expires 31 October 2036.
- Priority
- Filed
- Granted
- Today
- Expires
134 claims: 4 independent, 130 dependent
- 1一种装置,其包括: 一个或多个处理器; 存储器,其耦合到所述一个或多个处理器;以及 操作系统(OS),其存储在所述存储器中,被配置来在所述一个或多个处理器中的至少 一个子集上运行,其中所述操作系统被配置来选择性地以正常模式或时序测试模式运行, 其中在所述时序测试模式下,所述操作系统被配置来在用所述一个或多个处理器运行应用 程序时,扰乱所述一个或多个处理器上发生的处理的时序,并且在所述装置正在所述时序 测试模式下运行时测试所述应用程序以发现硬件组件和/或软件组件同步中的错误, 其中所述一个或多个处理器包括一个或多个中央处理单元(CPU)核心,其中在所述时 序测试模式下,所述一个或多个CPU核心的至少一个子集被配置来以比所述一个或多个CPU 核心在正常模式下的标准操作频率高的一个或多个频率操作。
- 2如权利要求1所述的装置,其中在所述时序测试模式下,所述操作系统被配置来在所 述一个或多个处理器上运行应用程序时,通过实时修改所述装置的一个或多个硬件设置来 扰乱所述一个或多个处理器上发生的处理的时序。
- 3如权利要求1所述的装置,其中在所述时序测试模式下,所述操作系统被配置来在所 述一个或多个处理器上运行应用程序时,通过以扰乱时序的方式向所述装置的一个或多个 硬件组件发送命令来扰乱所述一个或多个处理器上发生的处理的时序。
- 4如权利要求1所述的装置,其中,在所述时序测试模式下,所述OS修改所述一个或多 个CPU核心的所述至少一个子集中的特定CPU核心的操作频率。
- 5如权利要求1所述的装置,其中在所述时序测试模式下,所述一个或多个CPU核心的 所述至少一个子集被配置来以相同频率操作。
- 6如权利要求1所述的装置,其中在所述时序测试模式下,所述一个或多个CPU核心的 所述至少一个子集被配置来以不同频率操作。
- 7如权利要求1所述的装置,其中在所述时序测试模式下,所述一个或多个CPU核心的 所述至少一个子集中的一个或多个给定CPU核心被配置来以一个频率操作,而所述一个或 多个CPU核心的所述至少一个子集中的一个或多个其他CPU核心被配置来以另一频率操作。
- 8如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器的一个或多个高 速缓存,其中在所述时序测试模式下,所述一个或多个高速缓存的至少一个子集被配置来 以与所述一个或多个高速缓存在正常模式下的标准操作频率不同的一个或多个频率操作。
- 9如权利要求8所述的装置,其中所述一个或多个高速缓存的所述至少一个子集被配 置一次来以与所述一个或多个高速缓存在正常模式下的标准操作频率不同的一个或多个 频率操作。
- 10如权利要求8所述的装置,其中,在所述时序测试模式下,所述OS修改所述一个或多 个高速缓存的所述至少一个子集中的高速缓存的操作频率。
- 11如权利要求8所述的装置,其中在所述时序测试模式下,所述一个或多个高速缓存 的所述至少一个子集被配置来以相同频率操作。
- 12如权利要求8所述的装置,其中在所述时序测试模式下,所述一个或多个高速缓存 的所述至少一个子集被配置来以不同频率操作。
- 13如权利要求8所述的装置,其中在所述时序测试模式下,所述一个或多个高速缓存 的所述至少一个子集中的至少一个高速缓存被配置来以一个频率操作,而所述一个或多个 高速缓存的所述至少一个子集中的一个或多个其他高速缓存被配置来以另一频率操作。
- 14如权利要求8所述的装置,其中所述一个或多个高速缓存的所述至少一个子集中的 所述高速缓存中的一个或多个被配置来以比所述一个或多个高速缓存在正常模式下的标 准操作频率高的一个或多个频率操作。
- 15如权利要求8所述的装置,其中所述一个或多个高速缓存的所述至少一个子集中的 所述高速缓存中的一个或多个被配置来以比所述一个或多个高速缓存在正常模式下的标 准操作频率低的一个或多个频率操作。
- 16如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器的一个或多个高 速缓存,其中在所述时序测试模式下,所述一个或多个高速缓存的至少一个子集被配置来 以与所述一个或多个高速缓存在正常模式下的可用方式计数不同的可用方式计数操作。
- 17如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器的一个或多个高 速缓存,其中在所述时序测试模式下,所述一个或多个高速缓存的至少一个子集被配置来 以与所述一个或多个高速缓存在正常模式下的可用组计数不同的可用组计数操作。
- 18如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器的一个或多个高 速缓存,其中被配置来在正常模式下作为内含式高速缓存操作的所述一个或多个高速缓存 的至少一个子集被重新配置来在所述时序测试模式下作为独占式高速缓存操作。
- 19如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器的一个或多个高 速缓存,其中被配置来在正常模式下作为独占式高速缓存操作的所述一个或多个高速缓存 的至少一个子集被重新配置来在所述时序测试模式下作为内含式高速缓存操作。
- 20如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器的一个或多个高 速缓存,其中所述一个或多个高速缓存的至少一个子集的在正常模式下基于虚拟地址的高 速缓存和高速缓存查找行为被改变成在所述时序测试模式下基于虚拟地址的高速缓存和 高速缓存查找行为。
- 21如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器的一个或多个指 令高速缓存,其中所述一个或多个指令高速缓存的至少一个子集的在正常模式下被启用的 预提取功能在所述时序测试模式下被禁用。
- 22如权利要求1所述的装置,其中在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过运行干扰所述应用程序的一个或多个程序来扰 乱所述一个或多个处理器上发生的处理的时序。
- 23如权利要求22所述的装置,其中所述一个或多个程序从所述应用程序夺取资源。
- 24如权利要求22所述的装置,其中所述一个或多个程序暂停所述应用程序。
- 25如权利要求22所述的装置,其中所述一个或多个程序与所述应用程序竞争资源。
- 26如权利要求1所述的装置,其中在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过以扰乱时序的方式更改所述操作系统的功能来 扰乱所述一个或多个处理器上发生的处理的时序。
- 27如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),其 中,在所述时序测试模式下,所述装置被配置来在所述一个或多个处理器上运行应用程序 时,通过减少可用于运行所述应用程序的所述CPU的资源来扰乱所述一个或多个处理器上 发生的处理的时序。
- 28如权利要求27所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小队列的大小。
- 29如权利要求28所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小存储队列的大小。
- 30如权利要求28所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小加载队列的大小。
- 31如权利要求28所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小调度队列的大小。
- 32如权利要求28所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小停用队列的大小。
- 33如权利要求27所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小高速缓存的大小。
- 34如权利要求33所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小1级指令高速缓存的大小。
- 35如权利要求33所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小1级数据高速缓存的大小。
- 36如权利要求33所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小更高级别高速缓存的大小。
- 37如权利要求27所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小表后备缓冲器(TLB)的大小。
- 38如权利要求37所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小指令转换后备缓冲器(ITLB)的大小。
- 39如权利要求37所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小数据转换后备缓冲器(DTLB)的大小。
- 40如权利要求27所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小一个或多个指令管的执行速率。
- 41如权利要求27所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小一个或多个特定指令的执行速率。
- 42如权利要求27所述的装置,其中减少可用于运行所述应用程序的所述CPU的资源包 括减小由所述CPU执行的所有指令的执行速率。
- 43如权利要求27所述的装置,其中所述CPU包括算术逻辑单元(ALU),其中减少可用于 运行所述应用程序的所述CPU的资源包括减小由所述ALU执行的一个或多个特定指令的执 行速率。
- 44如权利要求27所述的装置,其中所述CPU包括算术逻辑单元(ALU),其中减少可用于 运行所述应用程序的所述CPU的资源包括减小由所述ALU执行的所有指令的执行速率。
- 45如权利要求27所述的装置,其中所述CPU包括地址生成单元(AGU),其中减少可用于 运行所述应用程序的所述CPU的资源包括减小由所述AGU执行的一个或多个特定指令的执 行速率。
- 46如权利要求27所述的装置,其中所述CPU包括地址生成单元(AGU),其中减少可用于 运行所述应用程序的所述CPU的资源包括减小由所述AGU执行的所有指令的执行速率。
- 47如权利要求27所述的装置,其中所述CPU包括单指令多数据(SIMD)单元,其中减少 可用于运行所述应用程序的所述CPU的资源包括减小由所述SIMD单元执行的一个或多个特 定指令的执行速率。
- 48如权利要求27所述的装置,其中所述CPU包括单指令多数据(SIMD)单元,其中减少 可用于运行所述应用程序的所述CPU的资源包括减小由所述SIMD单元执行的所有指令的执 行速率。
- 49如权利要求27所述的装置,其中所述CPU包括一个或多个处理器核心,其中减小可 用于运行所述应用程序的所述CPU的资源包括先占所述一个或多个处理器核心的一个或多 个个别处理器核心。
- 50如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过暂停一个或多个应用程序线程来扰乱所述一个 或多个处理器上发生的处理的时序。
- 51如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过暂停多个应用程序线程来扰乱所述一个或多个 处理器上发生的处理的时序。
- 52如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过暂停所有应用程序线程来扰乱所述一个或多个 处理器上发生的处理的时序。
- 53如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过随机地对一个或多个应用程序线程的暂停进行 时序的选择来扰乱所述一个或多个处理器上发生的处理的时序。
- 54如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过有条理地对一个或多个应用程序线程的暂停进 行时序的选择来扰乱所述一个或多个处理器上发生的处理的时序。
- 55如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过延迟在使用所述OS的功能时引入的时序来扰乱 所述一个或多个处理器上发生的处理的时序。
- 56如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过加速在使用所述OS的功能时引入的时序来扰乱 所述一个或多个处理器上发生的处理的时序。
- 57如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所述 一个或多个处理器上运行应用程序时,通过改变在使用所述OS的功能时引入的时序来扰乱 所述一个或多个处理器上发生的处理的时序。
- 58如权利要求1所述的装置,其中,在时序测试模式下,当所述OS执行所述应用程序所 请求的处理时,由所述OS所花费的时间和所述OS用于执行所述处理的所述一个或多个处理 器中的特定处理器与在正常模式下的操作期间由所述OS所花费的时间和用于执行所述处 理的所述特定处理器不同。
- 59如权利要求1所述的装置,其中,在时序测试模式下,当所述OS与来自所述应用程序 的请求无关地执行处理时,由所述OS所花费的时间和所述OS用于执行所述处理的所述一个 或多个处理器中的特定处理器与在正常模式下由所述OS所花费的时间和用于执行所述处 理的所述特定处理器不同。
- 60如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写或使所述CPU的一个或多个高速缓存和/或转换后备缓冲器无效来扰乱所述 一个或多个处理器上发生的处理的时序。
- 61如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写所述CPU的一个或多个高速缓存来扰乱所述一个或多个处理器上发生的处 理的时序。
- 62如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过使所述CPU的一个或多个高速缓存无效来扰乱所述一个或多个处理器上发生的 处理的时序。
- 63如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写所述CPU的一个或多个转换后备缓冲器来扰乱所述一个或多个处理器上发 生的处理的时序。
- 64如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过使所述CPU的一个或多个转换后备缓冲器无效来扰乱所述一个或多个处理器上 发生的处理的时序。
- 65如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写所述CPU的一个或多个指令转换后备缓冲器(ITLB)来扰乱所述一个或多个 处理器上发生的处理的时序。
- 66如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过使所述CPU的一个或多个指令转换后备缓冲器(ITLB)无效来扰乱所述一个或多 个处理器上发生的处理的时序。
- 67如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写所述CPU的一个或多个数据转换后备缓冲器(DTLB)来扰乱所述一个或多个 处理器上发生的处理的时序。
- 68如权利要求1所述的装置,其中所述一个或多个处理器包括中央处理单元(CPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过使所述CPU的一个或多个数据转换后备缓冲器(DTLB)无效来扰乱所述一个或多 个处理器上发生的处理的时序。
- 69如权利要求1所述的装置,其中所述一个或多个处理器包括多个处理器核心,其中 在所述时序测试模式下,所述装置被配置来在所述多个处理器核心中的除了由所述应用程 序指定的处理器核心以外的一个处理器核心上执行所述应用程序的线程。
- 70如权利要求1所述的装置,其中所述一个或多个处理器包括具有两个或更多个集群 的中央处理单元(CPU),所述集群共享更高级别高速缓存,每个集群包含耦合到集群级别高 速缓存的一个或多个核心,其中,在所述时序测试模式下,所述装置被配置来将所述应用程 序的线程重新指派给与正常模式下所述线程被指派的集群不同的集群中的核心。
- 71如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器和所述存储器的 一个或多个总线,其中在所述时序测试模式下,所述一个或多个总线的至少一个子集被配 置来以与所述一个或多个总线在正常模式下的标准操作频率不同的一个或多个频率操作。
- 72如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器和所述存储器的 一个或多个总线,其中在所述时序测试模式下,所述一个或多个总线的至少一个子集被配 置来以比所述一个或多个总线在正常模式下的标准操作频率高的一个或多个频率操作。
- 73如权利要求1所述的装置,其还包括耦合到所述一个或多个处理器和所述存储器的 一个或多个总线,其中在所述时序测试模式下,所述一个或多个总线的至少一个子集被配 置来以比所述一个或多个总线在正常模式下的标准操作频率低的一个或多个频率操作。
- 74如权利要求1所述的装置,其中所述一个或多个处理器包括具有一个或多个GPU核 心的图形处理器单元,其中在所述时序测试模式下,所述一个或多个GPU核心的至少一个子 集被配置来以与所述一个或多个GPU核心在正常模式下的标准操作频率不同的一个或多个 频率操作。
- 75如权利要求74所述的装置,其中在所述时序测试模式下,所述一个或多个GPU核心 的所述至少一个子集被配置来以相同频率操作。
- 76如权利要求74所述的装置,其中在所述时序测试模式下,所述一个或多个GPU核心 的所述至少一个子集被配置来以不同频率操作。
- 77如权利要求74所述的装置,其中在所述时序测试模式下,所述一个或多个GPU核心 的所述至少一个子集中的至少一个GPU核心被配置来以一个频率操作,而所述一个或多个 GPU核心的所述至少一个子集中的一个或多个其他GPU核心被配置来以另一频率操作。
- 78如权利要求74所述的装置,其中所述一个或多个GPU核心的所述至少一个子集中的 所述GPU核心中的一个或多个被配置来以比所述一个或多个GPU核心在正常模式下的标准 操作频率高的一个或多个频率操作。
- 79如权利要求74所述的装置,其中所述一个或多个GPU核心的所述至少一个子集中的 所述GPU核心中的一个或多个被配置来以比所述一个或多个GPU核心在正常模式下的标准 操作频率低的一个或多个频率操作。
- 80如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),其 中,在所述时序测试模式下,所述装置被配置来在所述一个或多个处理器上运行应用程序 时,通过更改所述GPU的功能来扰乱所述一个或多个处理器上发生的处理的时序。
- 81如权利要求80所述的装置,其中更改所述GPU的功能包括用具有与所述装置的正常 操作不同的时序的固件替换GPU固件。
- 82如权利要求80所述的装置,其中更改所述GPU的功能包括扰乱对象处理。
- 83如权利要求80所述的装置,其中在所述时序测试模式下,所述操作系统请求所述 GPU执行扰乱所述GPU的时序的处理。
- 84如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),其 中,在所述时序测试模式下,所述装置被配置来在所述一个或多个处理器上运行应用程序 时,通过减少GPU资源来扰乱所述一个或多个处理器上发生的处理的时序。
- 85如权利要求84所述的装置,其中减少GPU资源包括减小GPU高速缓存的大小。
- 86如权利要求84所述的装置,其中减少GPU资源包括减小所述GPU的转换后备缓冲器 的大小。
- 87如权利要求84所述的装置,其中减少GPU资源包括减小所述GPU的指令转换后备缓 冲器(ITLB)的大小。
- 88如权利要求84所述的装置,其中减少GPU资源包括减小所述GPU的数据转换后备缓 冲器(DTLB)的大小。
- 89如权利要求84所述的装置,其中减少GPU资源包括减小一个或多个指令的执行速 率。
- 90如权利要求84所述的装置,其中减少GPU资源包括减小指令管的执行速率。
- 91如权利要求84所述的装置,其中减少GPU资源包括减小一个或多个特定GPU指令的 执行速率。
- 92如权利要求84所述的装置,其中减少GPU资源包括减小由所述GPU执行的所有指令 的执行速率。
- 93如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写或使所述GPU的一个或多个高速缓存和/或转换后备缓冲器无效来扰乱所述 一个或多个处理器上发生的处理的时序。
- 94如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写所述GPU的一个或多个高速缓存来扰乱所述一个或多个处理器上发生的处 理的时序。
- 95如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过使所述GPU的一个或多个高速缓存无效来扰乱所述一个或多个处理器上发生的 处理的时序。
- 96如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写所述GPU的一个或多个指令转换后备缓冲器(ITLB)来扰乱所述一个或多个 处理器上发生的处理的时序。
- 97如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过使所述GPU的一个或多个指令转换后备缓冲器(ITLB)无效来扰乱所述一个或多 个处理器上发生的处理的时序。
- 98如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过回写所述GPU的一个或多个数据转换后备缓冲器(DTLB)来扰乱所述一个或多个 处理器上发生的处理的时序。
- 99如权利要求1所述的装置,其中所述一个或多个处理器包括图形处理单元(GPU),并 且其中在所述时序测试模式下,所述装置被配置来在用所述一个或多个处理器运行应用程 序时,通过使所述GPU的一个或多个数据转换后备缓冲器(DTLB)无效来扰乱所述一个或多 个处理器上发生的处理的时序。
- 100如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所 述一个或多个处理器上运行应用程序时,通过更改所述存储器的功能来扰乱所述一个或多 个处理器上发生的处理的时序。
- 101如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所 述一个或多个处理器上运行应用程序时,通过向一个或多个存储器操作添加时延来扰乱所 述一个或多个处理器上发生的处理的时序。
- 102如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所 述一个或多个处理器上运行应用程序时,通过改变一个或多个存储器操作的优先级来扰乱 所述一个或多个处理器上发生的处理的时序。
- 103如权利要求1所述的装置,其中,在所述时序测试模式下,所述装置被配置来在所 述一个或多个处理器上运行应用程序时,通过将正常模式下放在一条地址线上的信号用放 在另一地址线上的信号交换来扰乱所述一个或多个处理器上发生的处理的时序。
- 104如权利要求1所述的装置,其中在所述时序测试模式下,所述存储器被配置来以与 所述存储器在正常模式下的标准操作频率不同的一个或多个频率操作。
- 105如权利要求1所述的装置,其中在所述时序测试模式下,所述存储器被配置来以比 所述存储器在正常模式下的标准操作频率高的一个或多个频率操作。
- 106如权利要求1所述的装置,其中在所述时序测试模式下,所述存储器被配置来以比 所述存储器在正常模式下的标准操作频率低的一个或多个频率操作。
- 107如权利要求1所述的装置,其中在所述时序测试模式下,存储器时钟被配置来以与 所述存储器时钟在正常模式下的标准操作频率不同的频率运行。
- 108如权利要求1所述的装置,其中在所述时序测试模式下,存储器时钟被配置来以比 所述存储器时钟在正常模式下的标准操作频率高的频率运行。
- 109如权利要求1所述的装置,其中在所述时序测试模式下,存储器时钟被配置来以比 所述存储器时钟在正常模式下的标准操作频率低的频率运行。
- 110如权利要求1所述的装置,其中在所述时序测试模式下,总线时钟被配置来以与所 述存储器时钟在正常模式下的标准操作频率不同的频率运行。
- 111如权利要求1所述的装置,其中在所述时序测试模式下,总线时钟被配置来以比所 述存储器时钟在正常模式下的标准操作频率高的频率运行。
- 112如权利要求1所述的装置,其中在所述时序测试模式下,总线时钟被配置来以比所 述存储器时钟在正常模式下的标准操作频率低的频率运行。
- 113如权利要求1所述的装置,其中在所述时序测试模式下,所述操作系统修改存储器 时钟的频率。
- 114如权利要求1所述的装置,其中在所述时序测试模式下,所述操作系统修改内部总 线时钟的频率。
- 115如权利要求1所述的装置,其还包括耦合到所述存储器和所述一个或多个处理器 中的至少一者的存储器控制器。
- 116如权利要求115所述的装置,其中,在所述时序测试模式下,所述存储器控制器被 配置来模拟从所述存储器适当读取数据的一个或多个故障。
- 117如权利要求115所述的装置,其中,在所述时序测试模式下,所述存储器控制器被 配置来增加由所述存储器控制器执行的一个或多个读取的时延。
- 118如权利要求115所述的装置,其中,在所述时序测试模式下,所述存储器控制器被 配置来在各种类型的读取之间使用与在正常模式下运行所述装置时使用的优先级次序不 同的优先级次序。
- 119如权利要求115所述的装置,其中在所述时序测试模式下,所述存储器控制器被配 置来置换地址线。
- 120一种装置,其包括: 一个或多个处理器; 存储器,其耦合到所述一个或多个处理器;并且 其中所述装置被配置来选择性地以正常模式或时序测试模式运行,其中在时序测试模 式下,所述装置被配置来在用所述一个或多个处理器运行应用程序时,扰乱所述一个或多 个处理器上发生的处理的时序,并且在所述装置在所述时序测试模式下运行时测试所述应 用程序以发现硬件组件和/或软件组件同步中的错误, 其中所述一个或多个处理器包括一个或多个央处理器单元(CPU)核心,其中在所述时 序测试模式下,所述一个或多个CPU核心的至少一个子集被配置来以比所述一个或多个CPU 核心在正常模式下的标准操作频率高的一个或多个频率操作。
- 121如权利要求120所述的装置,其中所述一个或多个处理器包括具有一个或多个CPU 核心的中央处理器单元(CPU),其中所述一个或多个CPU核心的至少一个子集被配置来以与 所述一个或多个CPU核心在正常模式下的标准操作频率不同的一个或多个频率操作。
- 122如权利要求120所述的装置,其还包括一个或多个高速缓存,其中所述一个或多个 高速缓存的至少一个子集被配置一次来以与所述一个或多个高速缓存在正常模式下的标 准操作频率不同的一个或多个频率操作。
- 123如权利要求120所述的装置,其还包括一个或多个高速缓存,其中所述一个或多个 高速缓存的至少一个子集被配置一次来以比所述一个或多个高速缓存在正常模式下的标 准操作频率高的一个或多个频率操作。
- 124如权利要求120所述的装置,其还包括一个或多个总线,其中所述一个或多个总线 的至少一个子集被配置来以比所述一个或多个总线在正常模式下的标准操作频率高的一 个或多个频率操作。
- 125如权利要求120所述的装置,其还包括一个或多个高速缓存,其中所述一个或多个 高速缓存的至少一个子集被配置一次来以比所述一个或多个高速缓存在正常模式下的标 准操作频率低的一个或多个频率操作。
- 126如权利要求120所述的装置,其还包括一个或多个总线,其中所述一个或多个总线 的至少一个子集被配置来以与所述一个或多个总线在正常模式下的标准操作频率不同的 一个或多个频率操作。
- 127如权利要求120所述的装置,其还包括一个或多个总线,其中所述一个或多个总线 的至少一个子集被配置来以比所述一个或多个总线在正常模式下的标准操作频率低的一 个或多个频率操作。
- 128如权利要求120所述的装置,其中所述存储器被配置来以与所述存储器在正常模 式下的标准操作频率不同的一个或多个频率操作。
- 129如权利要求120所述的装置,其中所述存储器被配置来以比所述存储器在正常模 式下的标准操作频率高的一个或多个频率操作。
- 130如权利要求120所述的装置,其中所述存储器被配置来以比所述存储器在正常模 式下的标准操作频率低的一个或多个频率操作。
- 131如权利要求120所述的装置,其还包括被配置来确定由所述装置执行的每循环指 令数目的一个或多个电路。
- 132如权利要求120所述的装置,其中在所述时序测试模式下,所述装置在所述装置正 在所述时序测试模式下运行时通过监视IPC的变化来测试所述应用程序的错误。
- 133一种在具有一个或多个处理器以及耦合到所述一个或多个处理器的存储器的装 置中实现的方法,其包括: 以时序测试模式运行所述装置,其中在所述时序测试模式下,所述装置被配置来在用 所述一个或多个处理器运行应用程序时,扰乱所述一个或多个处理器上发生的处理的时 序;并且 在所述装置正在所述时序测试模式下运行时测试所述应用程序以发现硬件组件和/或 软件组件同步中的错误, 其中所述一个或多个处理器包括一个或多个中央处理单元(CPU)核心,其中在所述时 序测试模式下,所述一个或多个CPU核心的至少一个子集被配置来以比所述一个或多个CPU 核心在正常模式下的标准操作频率高的一个或多个频率操作。
- 134一种非暂时性计算机可读介质,其具有在其中体现的计算机可读可执行指令,所 述指令被配置来使具有处理器和存储器的装置在执行所述指令时实现方法,所述方法包 括: 以时序测试模式运行所述装置,其中在所述时序测试模式下,所述装置被配置来在用 一个或多个处理器运行应用程序时,扰乱所述一个或多个处理器上发生的处理的时序;并 且 在所述装置正在所述时序测试模式下运行时测试所述应用程序以发现硬件组件和/或 软件组件同步中的错误, 其中所述一个或多个处理器包括一个或多个中央处理单元(CPU)核心,其中在所述时 序测试模式下,所述一个或多个CPU核心的至少一个子集被配置来以比所述一个或多个CPU 核心在正常模式下的标准操作频率高的一个或多个频率操作。
Independent claims134
90 paragraphs in 1 section, as filed
Software backward compatibility testing in scrambled mode
[0001] This application claims the benefit of priority to commonly assigned U.S. non-provisional application No. 14/930,408, filed on November 02, 2015, the entire contents of which are incorporated herein by reference.
technical field
[0002] Aspects of the present disclosure relate to executing a computer application on a computer system. In particular, aspects of the present disclosure relate to a system or method for providing backward compatibility for applications/titles designed for older versions of computer systems.
Background technique
[0003] Modern computer systems often use several different processors for different computing tasks. For example, in addition to several central processing units (CPUs), modern computers may have graphics processing units (GPUs) dedicated to certain computational tasks in the graphics pipeline, or dedicated to digital signal processing for audio, all of which A unit is potentially part of an accelerated processing unit (APU), which may also include other units. These processors are connected to various types of memory using a bus that can be located either inside the APU or externally on the computer's motherboard.
[0004] A set of applications is typically created for a computer system such as a video game console or smartphone ("legacy device"), and when a variant or more advanced version of a computer system is released ("new device"), the legacy version is expected to be The device's application runs flawlessly on the new device without recompilation or any modification taking into account the properties of the new device. This aspect of a new device is often referred to as "backward compatibility" as contained in the new device's hardware architecture, firmware, and operating system.
[0005] Backward compatibility is often achieved through binary compatibility, where new devices are able to execute programs created for older devices. However, when the real-time behavior of such device classes is important to their operation, as in the case of video game consoles or smartphones, significant differences in the operating speed of new devices may render them incompatible with backward compatibility with older devices. Issues preventing backward compatibility may arise if the new device has lower performance than the older device; the same is true if the new device has higher performance or different performance characteristics than the older device.
It is against this background that various aspects of the present disclosure arise.
Description of drawings
[0007] The teachings of the present disclosure can be readily understood by considering the following detailed description in conjunction with the accompanying drawings, in which:
[0008] FIG. 1 is a block diagram illustrating one example of a central processing unit (CPU) core that may be configured to operate in a backward compatibility mode in accordance with aspects of the present disclosure.
[0009] FIG. 2 is a block diagram illustrating one example of a possible multi-core architecture for a CPU in accordance with aspects of the present disclosure.
[0010] FIG. 3 is a block diagram of an apparatus having a CPU configured to operate in a backward compatibility mode in accordance with aspects of the present disclosure.
[0011] FIG. 4 is a timing diagram illustrating the concept of "skew".
[0012] FIG. 5 is a flow diagram illustrating the operation of an apparatus in a sequential test mode in accordance with aspects of the present disclosure.
[0013] Introduction
[0014] Even if the CPU of the new device is binary compatible with the legacy device (ie, capable of executing programs created for the legacy device), differences in performance characteristics between the CPU of the new device and the CPU of the legacy device may still cause errors in legacy applications. error, so the new unit will not be backward compatible.
[0015] If the CPU of the new device has lower performance than the CPU of the older device, many bugs in the legacy application may arise due to the inability to meet real-time deadlines imposed by display timing, audio streaming, and the like. If the CPU of the new device has substantially higher performance than the CPU of the older device, many bugs in the legacy application may arise as a result of untested results of this high speed operation. For example, in a producer-consumer model, if a data consumer (eg, a CPU) is operating at a higher speed than initially expected, it may try to make the data available before the data producer (eg, some other component of the computer) access data. Alternatively, if a data producer (eg, a CPU) operates at a higher speed than initially expected, it may overwrite data that is still in use by a data consumer (eg, some other component of the computer).
[0016] Additionally, since the speed at which the CPU executes code depends on the characteristics of the specific code being executed, it is likely that the degree to which the performance of the CPU of a new device will increase relative to an older device will depend on the specific code being executed. This can lead to problems in the producer-consumer model described above, where both the producer and the consumer are CPUs, but execute the code of the legacy application at a relative speed not encountered on older hardware.
Detailed ways
[0017] Aspects of the present disclosure describe computer systems and methods that may allow applications written for a device to run on a second device with a higher degree of backward compatibility, the second device Binary compatible (because programs written for the first device will run on the second device) but have different timing characteristics (because programs written for the first device will execute at different rates on the second device, potentially producing in-operation mistake). The second device can potentially be a variant or more advanced version of the first device, and can potentially be configured in a "backward compatibility mode" in which the features and capabilities of the second device more closely approximate the features and capabilities of the first device ability.
[0018] In an implementation of the present disclosure, a timing test mode is created for the first device. This mode creates timing that is not (or generally not) discovered on the device, with the result that when an application is running in this mode, there is no interaction between hardware components (such as CPU, GPU, audio and video hardware) or software Errors in synchronization, such as application processing or OS processing, arise in ways that are not possible or common on the device during normal operation. Once these errors in synchronization are detected, the application software can be repaired to eliminate or mitigate these errors, thereby increasing the likelihood that the application will execute properly on a second device with different timing characteristics, ie, relative to the first device , the application will have a higher degree of backward compatibility on the second device. Since the capabilities of the second device may not be known (eg, it may be a future device that does not yet exist), it is beneficial to have a large variation in the available timing in the timing test mode.
[0019] In an implementation of the present disclosure, in the timing test mode, the operating system may configure the hardware in a particular state (eg, at a particular operating frequency not found in normal operation of the device). Additionally, in timing test mode, the operating system can change the hardware configuration while the application is running or perform various processing while the application is running (for example, competing for system resources or preempting a process that is processed by the application).
[0020] In implementations of the present disclosure, the tests may be performed on hardware other than the device. For example, using an IC selected to operate with a wider operating range than a consumer device would allow for a test mode not available on a consumer device.
[0021] FIG. 1 shows a generalized architecture of a CPU core 100. The CPU core 100 typically includes a branch prediction unit 102 that attempts to predict whether a branch will be taken, and also attempts (if the branch is taken) to predict the destination address of the branch. In cases where these predictions are correct, the efficiency of speculatively executed code will increase
Plus; therefore highly accurate branch prediction is highly desirable. Branch prediction unit 102 may include highly specialized subunits, such as a return address stack 104 that tracks return addresses from subroutines, an indirect target array 106 that tracks the destinations of indirect branches, and the past history of branches to more accurately The branch target buffer 108 and its associated prediction logic that predicts the resulting address of the branch.
The CPU core 100 typically includes an instruction fetching and decoding unit 110, which includes an instruction fetching unit 112, an instruction byte buffer 114, and an instruction decoding unit 1160. The CPU core 100 typically also includes several instruction-related high-speed Cache and Instruction Translation Lookaside Buffer (ITLB) 120 . These may include an ITLB cache hierarchy 124 that caches virtual addresses to physical address translation information such as page table entries, page directory entries, and the like. This information is used to translate the virtual address of the instruction to a physical address so that instruction fetch unit 112 can load the instruction from the cache hierarchy. By way of example and not limitation, program instructions may be cached according to a cache hierarchy including a level 1 instruction cache (L1I-cache) 122 resident in the core and in the CPU core 100. External other cache levels 176; these caches are first searched for program instructions using the physical address of the instruction. If the instruction is not found, the instruction is loaded from system memory 101 . Depending on the architecture, there may also be a micro-op cache 126 containing decoded instructions, as described below.
[0023] Once the program instructions are fetched, they are typically placed in the instruction byte buffer 114, awaiting processing by the instruction fetch and decode unit 110. Decoding can be a very complex process; it is difficult to decode multiple instructions per loop, and there may be restrictions on instruction alignment or instruction type that limit how many instructions can be decoded in a loop. Depending on the architecture, the decoded instructions may be placed in the micro-op cache 126 (if a micro-op cache exists on the new CPU) so that the decode stage can be bypassed for subsequent use of the program instructions.
[0024] The decoded instructions are typically passed to other units 130 for dispatch and scheduling. These units may use the disable queue 132 to track the status of instructions throughout the remainder of the CPU pipeline. Also, due to the limited number of general-purpose and SIMD registers available on many CPU architectures, register renaming can be performed, where when a logical (also called architectural) register is encountered in the instruction stream being executed, the assignment Physical registers 140 to represent them. The physical registers 140 may include a single instruction multiple data (SIMD) register set 142 and a general purpose (GP) register set 144, which may be much larger in size than the number of logical registers available on a particular CPU architecture, so that performance may be significantly increased. After register renaming 134 is performed, instructions are typically placed in a dispatch queue 136 from which each cycle may select several instructions (based on dependencies) for execution by execution units 150 .
Execution unit 150 generally includes: SIMD pipe 152, which performs several parallel operations on multiple data fields contained in the 128-bit or wider SIMD registers contained in SIMD register bank 142; Arithmetic and Logic Unit (ALU) 154, which performs several logical, arithmetic, and hash operations on the GPRs contained in the GP register set 144; and an address generation unit (AGU) 156, which calculates the address from which the memory should store or load. Multiple instances of each type of execution unit may exist, and the instances may have different capabilities, eg, a particular SIMD pipe 152 may be capable of performing floating-point multiply operations but not floating-point add operations.
[0026] Stores and loads are typically buffered in store queue 162 and load queue 164 so that many store operations can be performed in parallel. To aid memory operations, the CPU core 100 typically includes several data-dependent caches and a data translation lookaside buffer (DTLB) 1700. The DTLB cache hierarchy 172 caches virtual addresses to physical address translations such as page table entries, page directory entries, etc.; This information is used to translate virtual addresses of memory operations to physical addresses so that data can be stored or loaded from system memory. Data is usually cached at level 1 residing in the core
Data cache (L1D-cache) 174 and other cache levels 176 external to core 100 .
[0027] According to certain aspects of the present disclosure, a CPU may include multiple cores. By way of example and not limitation, FIG. 2 illustrates an example of a possible multi-core CPU 200 that may be used in conjunction with aspects of the present disclosure. Specifically, the architecture of the CPU 200 may include M clusters 201-1, . . . , 2015, where M is an integer greater than zero. Each cluster may have N cores 202-1, 202-2, . . . , 202-N, where N is an integer greater than one. Aspects of the present disclosure include implementations in which different clusters have different numbers of cores. Each core may include one or more corresponding dedicated local caches (eg, L1 instruction, L1 data, or L2 caches). Each of the local caches may be dedicated to a particular corresponding core, as it is not shared with any other core. Each cluster may also include cluster-level caches 203-1, . . . , 203-M that may be shared among the cores in the corresponding cluster. In some implementations, cluster level caches are not shared by cores associated with different caches. Additionally, the CPU 200 may also include one or more higher level caches 204 that may be shared among the clusters. To facilitate communication between the cores in the cluster, the clusters 201-1, . . . , 202-M may include a corresponding local bus 205-1, 202-M coupled to each of the cores and a cluster-level cache for the cluster ..., 205-M. Likewise, for To facilitate communication between the various clusters, CPU 200 may include one or more higher level buses 206 coupled to clusters 201 - 1 , . . . , 201 -M and higher level caches 204 . In some implementations, the higher-level bus 206 may also be coupled to other devices, such as a GPU, memory, or memory controller. In other implementations, the higher-level bus 206 may be connected to a device-level bus that connects to different devices within the system. In other implementations, a higher-level bus 206 may couple the clusters 201-1, . , for example, GPU, memory, or memory controller. By way of example and not limitation, implementations with such a device-level bus 208 can be produced, eg, where the higher-level cache 204 is L3 for all CPU cores, but not used by the GPU.
[0028] In the CPU 200, OS processing may occur primarily on a certain core or a certain subset of the cores. Similarly, application-level processing may occur primarily on a core or a subset of cores. Individual application threads can be designated by an application to run on a core or a subset of cores. Because of the shared cache and bus, the processing speed of a given application thread may vary depending on processing occurring in other threads (eg, application threads or OS threads) running in the same cluster as the given application thread. Depending on the specificity of the CPU 200, the core may be able to execute only one thread at a time, or may be able to execute multiple threads simultaneously (hyperthreading"). In the case of hyperthreaded CPUs, the application may also specify which threads can communicate with which Other threads execute concurrently. The performance of a thread is affected by specific processing performed by any other thread executing on the same core.
[0029] Turning now to FIG. 3, an illustrative example of an apparatus 300 configured to operate in accordance with aspects of the present disclosure is shown. According to aspects of the present disclosure, the device 300 may be an embedded system, a cell phone, a personal computer, a tablet computer, a portable gaming device, a workstation, a gaming console, or the like.
Apparatus 300 typically includes a central processing unit (CPU) 320, which may include one or more CPU cores 3230 of the type shown in FIG. 1 and described above A plurality of such cores 323 and one or more caches 325 of the configuration shown. By way of example and not limitation, CPU 320 may be part of an accelerated processing unit (APU) 310 that includes CPU 320 and a graphics processing unit (GPU) 330 on a single chip. In alternative implementations, CPU 320 and GPU 330 may be implemented as separate hardware components on separate chips. GPU 330 may also include two or more cores 332 and two or more caches 334 and (in some implementations) one or more buses for facilitating each core and Communication between caches and other components of the system. The buses may include internal bus 317 for APU 310 and external data bus 390 .
[0031] The apparatus 300 may also include a memory 340. Memory 340 may optionally include a main memory unit accessible by CPU 320 and GPU 330 . CPU 320 and GPU 330 may each include one or more processor cores, eg, a single core, two cores, four cores, eight cores, or more cores. CPU 320 and GPU 330 may be configured to access one or more memory units using external data bus 390, and in some implementations, it may be useful for apparatus 300 to include two or more different buses.
[0032] Memory 340 may include one or more memory cells in the form of integrated circuits that provide addressable memory, eg, RAM, DRAM, and the like. The memory may contain executable instructions configured to implement the method of FIG. 5 when executed to determine to operate the device 300 in a timing test mode when running an application originally created for execution on a legacy CPU. Additionally, memory 340 may include dedicated graphics memory for temporarily storing graphics resources, graphics buffers, and other graphics data for the graphics rendering pipeline.
[0033] The CPU 320 may be configured to execute CPU code, which may include an operating system (OS) 321 or applications 322 (eg, video games). The operating system may include a kernel that manages input/output (I/O) requests from software (eg, applications 322 ) and translates them into CPU 320 , GPU 330 , or other components of device 300 data processing instructions. OS 321 may also include firmware that may be stored in non-volatile memory. OS 321 may be configured to implement certain features of operating CPU 320 in sequential test mode, as described in detail below. The CPU code may include a graphics application programming interface (API) 324 for issuing draw commands or draw calls to programs implemented by the GPU 330 based on the state of the application 322. The CPU code can also implement physics simulation and other functions. Portions of the code for one or more of the OS 321 , applications 322 or API 324 may be stored in memory 340 , in a cache internal or external to the CPU, or in mass storage accessible by the CPU 320 middle.
[0034] The device 300 may include a memory controller 315. Memory controller 315 may be a digital circuit that manages the flow of data to and from memory 340 . By way of example and not limitation, the memory controller may be an integral part of APU 310, as shown in the example of Figure 3, or may be a separate hardware component.
[0035] The apparatus 300 may also include well-known support functions 350, which may communicate with other components of the system, eg, through a bus 390. Such support functions may include, but are not limited to, input/output (I/O) elements 352, one or more clocks 356 that may include separate clocks for CPU 320, GPU 330, and memory 340, respectively, as well as One or more levels of cache 358 external to GPU 330 . Device 300 may optionally include a mass storage device 360, such as a disk drive, CD-ROM drive, flash memory, tape drive, Blu-ray drive, etc., for storing programs and/or data. In one example, mass storage device 360 may receive computer-readable media 362 containing legacy application programs originally designed to run on systems with legacy CPUs. Alternatively, legacy application 362 (or a portion thereof) may be stored in memory 340 or partially in cache 358 .
[0036] The apparatus 300 may also include a display unit 380 for presenting the rendered graphics 382 prepared by the GPU 330 to the user. The apparatus 300 may also include a user interface unit 370 to facilitate interaction between the system 100 and the user. The display unit 380 may be in the form of a flat panel display, a cathode ray tube (CRT) screen, a touch screen, a head mounted display (HMD) or other device that may display text, numbers, graphic symbols or images. Display 380 may display rendered graphics 382 processed in accordance with various techniques described herein. User interface 370 may include one or more peripheral devices, such as a keyboard, mouse, joystick, light pen, game controller, touch screen, and/or other devices that may be used in conjunction with a graphical user interface (GUI). In some implementations, the state of application 322 and the underlying content of the graphics may be determined, at least in part, by user input through user interface 370, eg, where application 322 includes a video game or other graphics-intensive application.
[0037] The device 300 may also include a network interface 372 to enable the device to communicate with other devices over a network.
The network may be, for example, a local area network (LAN), a wide area network (such as the Internet), a personal area network (such as a Bluetooth network), or other types of networks. Each of the components shown and described may be implemented in hardware, software or firmware or some combination of two or more of these.
[0038] Aspects of the present disclosure overcome backward compatibility issues that arise due to timing differences when programs written for legacy systems run on newer systems that are more powerful or differently configured. By running the apparatus 300 in sequential test mode, a developer can determine how software written for an older system will perform when operating on a new system.
[0039] In accordance with aspects of the present disclosure, the apparatus 300 may be configured to operate in a sequential test mode. To understand the usefulness of this mode of operation, consider the timing diagram of Figure 4. In Figure 4, when running an application, different computing elements (eg, CPU cores) A, B, C, D can be run by the parallelogram 4<sub>1</sub>··Mountain<sub>4</sub>field<sub>1</sub>・·・8<sub>4</sub>、(2<sub>1</sub>··(<sub>4</sub>,mouth<sub>1</sub>··2<sub>4</sub>Different tasks indicated. Some tasks need to produce data for consumption by other tasks, and the other tasks cannot start work until the required data is produced. For example, suppose task A<sub>2</sub>requires the data produced by task 4, while task B<sub>2</sub>Data produced by tasks A1 and B1 is required. To ensure proper operation, typically the application will use semaphores or other synchronization strategies between tasks, such as when starting task B<sub>2</sub>Before, tasks A1 and B1 should be checked (resulting in task B<sub>2</sub>required source data) has been run to completion. Assume further that the timing shown in Figure 4 represents the timing when these tasks are implemented on legacy devices. Timings may be different on new devices (eg one with more processing power in core B), so task B1 may be in task A] already spawning task B<sub>2</sub>end before the required data. The offset in the relative timing of tasks on different processors is referred to herein as "skew." Such skew can expose software bugs in applications that will appear only on new devices or with increased frequency on new devices. For example, if on an older device, task A] is guaranteed to be in task B<sub>2</sub>before the end, then ensure that task A] is in task B<sub>2</sub>The sync code that ended before may never be tested, and if the sync code is not implemented properly, it is possible that this will only be known when running the application on the new device, e.g. Task B<sub>2</sub>May generate task 8 in task A]<sub>2</sub>Execution starts before the required data, which can lead to serious bugs in the application. Additionally, similar problems can arise when applications written to run on new devices run on older, less capable devices. To address these issues, a device such as 300 may operate in a timing test mode in which skew may be intentionally created, eg, between CPU threads, or between CPU 320 and GPU 330, or between CPU 320 and GPU 330. Between processes running on GPU 330, or between any of these and the real-time clock. Testing in this mode increases the likelihood that the application will run properly on future hardware.
[0040] According to aspects of the present disclosure, in the timing test mode, the CPU core may be configured to run at a different frequency (higher or lower) than for normal operation of the device, or the OS 321 may be continuously or Occasionally modify the frequency of the CPU cores. This can be done in a way that the CPU cores all run at the same frequency as each other, or a way that the CPU cores run at different frequencies from each other, or a way that some CPU cores run at a certain frequency and others run at another frequency Way.
[0041] By way of example and not limitation, if on a legacy device there are four cores running at 1 GHz on a consumer device in their typical mode of operation, then in sequential test mode, during consecutive ten second periods , which randomly selects a core to run at 800MHz. As a result, processes running on the selected core will run slower, exposing possible errors in the synchronization logic between that core and other cores, as other cores may try to complete the data prepared by the selected core The data is used before it is ready.
In various aspects of the present disclosure, in timing test mode, the clock rate of caches not included in the CPU core may be configured to operate at a different (higher or lower) frequency than its normal operating frequency or the same The normal operating frequency of the CPU core runs at a different frequency. If there are multiple caches that may be configured in this way, they may be configured to run at the same rate as each other, at different rates from each other, or some may run at a certain frequency and others at another frequency
rate operation.
[0043] In various aspects of the present disclosure, in the timing test mode, CPU resources may be configured to be limited in a manner that affects the execution timing of application code. Queues (eg, store and load queues, deactivation queues, and dispatch queues) may be configured to decrease in size (eg, available portions of resources may be limited). Caches (such as L1I-caches and D-caches, ITLB and DTLB cache hierarchies, and higher level caches) can be reduced in size (eg, the amount that can be stored in a fully associative cache can be reduced) The number of values, or for a cache with a limited number of ways, may reduce the available bank count or way count). The execution rate of all or specific instructions running on the ALU, AGU, or SIMD pipe may be reduced (eg, increased latency and/or decreased throughput).
[0044] In various aspects of the present disclosure, in the timing test mode, the OS may temporarily preempt (pause) the application thread. By way of example and not limitation, individual application threads may be pre-empted, or multiple threads may be pre-empted at the same time, or all threads may be pre-empted at the same time; the timing of pre-emption may be random or methodical; the number of pre-emptions and their length may be determined by Tuning to increase the likelihood that real-time deadlines (eg, for display timing or audio streaming output) can be met by the application.
In various aspects of the present disclosure, in the sequential test mode, when the OS performs processing (eg, a service (such as allocation)) requested by an application, or when the OS performs processing independently of the application request (eg, , servicing a hardware interrupt), the time spent by the OS and the processor (eg, CPU cores) used by the OS may differ from the time spent and the CPU cores used in the normal operating mode of the device. By way of example and not limitation, the time spent by the OS to perform memory allocations may be increased, or the OS may use CPU cores dedicated by applications under normal operation of the device to service hardware interrupts.
[0046] In various aspects of the present disclosure, in the timing test mode, an application thread may execute on a different CPU core than that specified by the application. By way of example and not limitation, in a system with two clusters (cluster "A" and cluster "B") each having two cores, all threads designated to execute on core 0 of cluster A may be conversely Execute on core 0 of cluster B, and all threads designated to execute on core 0 of cluster B may be executed on core 0 of cluster A conversely, resulting in different execution timing of thread processing compared to normal operation of the device , which is due to sharing the cluster high-level cache with different threads.
[0047] In various aspects of the present disclosure, in timing test mode, OS 321 may write back or invalidate CPU caches, or invalidate instruction and data TLBs randomly or methodically. By way of example and not limitation, the OS may randomly write back and invalidate the cache hierarchy of all CPU cores, resulting in delays in thread execution during invalidations and write backs and when thread requests are typically found in the cache hierarchy data time delays, resulting in timing not encountered during normal operation of the device.
In various aspects of the present disclosure, in the timing test mode, the GPU and any GPU sub-units with individually configurable frequencies may be configured to run at a different frequency than normal operation of the device, or the OS may continuously or The frequency of the GPU and any of its individually configurable subunits is occasionally modified.
[0049] Additionally, other behaviors of one or more caches (such as L1I-cache and D-cache, ITLB and DTLB cache hierarchies, and higher level caches) may disrupt the manner in which timing is in the timing test mode to modify. A non-limiting example of such a change in cache behavior modification would be to change whether a particular cache is exclusive or implicit. A cache that is inclusive in normal mode can be configured to be exclusive in timing test mode, and vice versa.
[0050] Another non-limiting example of cache behavior modification involves cache lookup behavior. in timing test mode
In normal mode, cache lookups can be implemented differently than in normal mode. Some newer processor hardware can actually slow down memory accesses compared to older processor hardware where newer hardware translates from virtual to physical addresses before a cache lookup and older hardware doesn't . For cache entries stored by physical addresses, as typically implemented for multi-core CPU caches 325, virtual addresses are always translated to physical addresses before performing a cache lookup (eg, in L1 and L2). Always translating virtual addresses to physical addresses before performing any cache lookups allows a core to write to a particular memory location to notify other cores not to write to that location. In contrast, cache lookups for cache entries stored according to virtual addresses (eg, for GPU cache 334) may be performed without translating addresses. This is faster because address translation only needs to be performed in the case of a cache miss, ie the entry is not in the cache and must be looked up in memory 340 . For example, where older GPU hardware stores cache entries through virtual addresses and newer GPU hardware stores cache entries through physical addresses, the difference between cache behavior can introduce 5 to 1000 cycles of latency in newer hardware. Delay. To test application 322 for errors caused by differences in cache lookup behavior, in timing test mode, against one or more caches Cache behavior and cache lookup behavior by (eg, GPU cache 334 ) may change from virtual address based to physical address based and vice versa.
[0051] Another non-limiting example of behavior modification would be to disable the I-cache prefetch function in timing test mode for one or more I-caches that have the I-cache prefetch function enabled in normal mode.
[0052] In various aspects of the present disclosure, in the timing test mode, the OS may replace the GPU firmware (if present) with firmware having different timing than the normal operation of the device. By way of example and not limitation, in sequential test mode, firmware may be replaced by firmware with higher overhead for each object processed, or with firmware that supports lower object counts that can be processed simultaneously, resulting in a device timing not encountered during normal operation.
[0053] In various aspects of the present disclosure, in the timing test mode, GPU resources may be configured to be limited in a manner that affects the processing timing of application requests. GPU cache 334 may be reduced in size (eg, the number of values that can be stored in a fully associative cache may be reduced, or for caches with a limited number of modes, the available set count or mode count may be reduced) . The execution rate of all or certain instructions running on GPU core 332 may be reduced (eg, increased latency and/or decreased throughput).
[0054] In various aspects of the present disclosure, in the timing test mode, OS 321 may request GPU 330 to perform processing that reduces the remaining resources available to application 322 for its processing. These requests can be random or orderly in their timing. By way of example and not limitation, OS 321 may request higher priority to render graphics objects or compute shaders, which may replace lower priority application rendering or other computations, or OS 321 may request that its processing occur at a particular on GPU cores 332 and thus unduly affect the application processing that is specified to be taking place on those GPU cores.
[0055] In aspects of the present disclosure, in timing test mode, OS 321 may randomly or methodically request GPU 330 to write back or invalidate its caches, or invalidate its instruction and data TLBs.
[0056] According to aspects of the present disclosure, APU 310 may include an internal clock 316 for an internal bus 317 that operates at a particular clock rate or set of rates referred to herein as an "internal bus clock." Internal bus 317 is connected to memory controller 315, which in turn is connected to external memory 340. Communication from the memory controller 315 to the memory 340 may occur at another particular clock rate, referred to herein as the "memory clock."
[0057] According to aspects of the present disclosure, when the device 300 is operating in a timing test mode, the memory clock and/or the internal bus clock may be configured to be different (higher or lower) than they are during normal operation of the device frequency, or the OS 321 may continuously or occasionally modify the frequency of the memory clock and/or the internal bus clock.
[0058] In various aspects of the present disclosure, in the timing test mode, the memory controller 315 may be configured to simulate the following: random failures to properly read data from external memory; increase certain functions performed by the memory controller latency of types of memory accesses; or use of a different priority order between various types of memory accesses than the priority order used during normal operation of the device. The OS 321 may modify these configurations continuously or sporadically in sequential test mode.
[0059] In accordance with aspects of the present disclosure, in the timing test mode, the memory controller 315 may be configured such that address lines are permuted, eg, a signal normally placed on one address line may be placed on another address line with signal exchange. By way of example and not limitation, if address line A is used to send column information to external memory 315 and address line B is used to send row information to external memory 340, and in timing test mode to address lines A and B The signals are exchanged, then the result will be that the timing is very different from that found during normal operation of the device.
Configuring the hardware and performing operations as described above (eg, configuring the CPU cores to run at different frequencies) can expose errors in the synchronization logic, but if the real-time behavior of the device is important, the timing test mode itself Errors in operation can result, for example, in the case of video game consoles, due to lower speed CPU cores being unable to meet real-time deadlines imposed by display timing, audio streaming output, and the like. According to aspects of the present disclosure, in the timing test mode, the apparatus 300 may operate at a speed higher than a standard operating speed. By way of non-limiting example, the higher than standard operating speed may be about 5% to about 30% higher than the standard operating speed. By way of example and not limitation, in timing test mode, the clocks of the CPU, CPU cache, GPU, internal bus and memory may be set to a frequency higher than the standard operating frequency (or standard operating frequency range) of the device. Since mass-production versions of device 300 can be constructed in a way that avoids setting the clock at higher than standard operating frequencies, it may be necessary to create specially designed hardware, such as hardware that uses memory chips at higher speeds than the corresponding mass-production devices, Either use hardware for parts of a system-on-chip (SoC) manufacturing operation that allow higher-than-average speed operation, or use higher-spec motherboard, power supply, and cooling hardware than is used on mass-production devices.
By way of example and not limitation, if the specially designed hardware allows the CPU to operate at a higher speed than that of the mass-produced device, and if on the mass-produced device there are four operating at 1 GHz in their typical mode of operation cores running, then in sequential test mode on specially designed hardware, during consecutive ten second cycles, three cores can be selected to run at 1.2GHz, while the remaining cores can run at 1GHz. Thus, processing running on selected cores will run slower than others, exposing possible bugs in synchronization logic, but not all cores running at least as fast as they would on mass-produced units As in the previous example, real-time deadlines can be met (eg, for showing timing), and the timing test mode itself is unlikely to cause errors in operation.
By way of example and not limitation, if the specially designed hardware allows the CPU to operate at a higher speed than that of the mass-produced device, and if on the mass-produced device there are four operating at 1 GHz in their typical mode of operation cores running, then in timing test mode on specially designed hardware, all cores can be selected to run at 1.2GHz and the OS321 can randomly write back and invalidate the CPU cache. If the slowdown due to cache writebacks and invalidations is less than the speedup due to higher CPU frequencies, then as above, the real-time deadline can be met, and the timing test mode itself is unlikely to cause errors in operation, in other words, the timing test mode can Skew is introduced by caching operations, and testing for synchronization errors can be performed without concern that the overall operation of the device will be slowed down and thus more prone to errors.
[0063] There are several ways in which application errors can be displayed in timing test mode. According to one implementation, specially designed hardware may include circuitry configured to determine the number of instructions per cycle (IPC) executed by apparatus 300 . The OS 321 can monitor changes in the IPC to test for errors in the application. In timing test mode, the OS can correlate significant changes in the IPC with specific modifications to device operation.
[0064] According to aspects of the present disclosure, a computer apparatus may operate in a sequential test mode. By way of example and not limitation, a computer system, such as apparatus 300, may have an operating system, such as operating system 321, configured to implement in a manner similar to method 500 shown in FIG. 5 and discussed below this test mode.
[0065] As shown at 501, the method begins. At 510, it is determined whether the system is to operate in a timing test mode. There are several ways in which this can be achieved. By way of example and not limitation, operating system 321 may prompt the user through rendered graphics 382 on display 380 to determine whether to enter timing test mode, and the user may enter appropriate instructions through user interface 370 . If it is determined that the system should not operate in the sequential test mode, then the system can operate normally, as shown at 520 . If it is determined that the system should operate in the sequential test mode, the device may be set to operate in the sequential test mode, as shown at 530 . Setting a device to operate in a timing test mode may generally involve the device's operating system (eg, OS 321 ) setting hardware states, loading firmware, and performing other operations in order to implement timing test mode-specific settings.
[0066] The apparatus 300 can be set to operate in a sequential test mode in any of a number of possible ways. By way of example and not limitation, in some implementations, a device may be configured externally, eg, through a network (eg, a local area network (LAN)). In another non-limiting example, a device may be configured internally using menus generated by the operating system and input from a user interface. In other non-limiting examples, a device may be set in a sequential test mode by physical configuration of the device hardware, eg, by manually setting the position of one or more dual in-line package (DIP) switches on the device run. Device firmware (eg, stored in ROM) may then read the settings of the DIP switches, eg, when the device is powered on. The latter implementation may be useful, for example, where the device is specially designed hardware rather than a mass-produced version of the device. In such cases, for convenience, the switch may be located outside the box or case containing the device hardware.
[0067] Once the device is set to run in sequential test mode, the device may run the application in sequential test mode, as shown at 540. There are several ways in which the operation of the system in sequential test mode may differ from normal device operation.
[0068] By way of example and not limitation, while the application 322 is running, the OS 321 may execute one or more of the following while running the application in a timing test:
[0069] . Modify hardware settings in real time, as shown in 542;
0 send commands to various hardware components of device 300 in a manner that disrupts timing, as shown at 544;
[0071] . Programs that interfere with application 322 are run, eg, by seizing resources from, suspending, or competing with applications for resources, as indicated at 546 .
[0072]. Change the functionality of OS 321 in a timing test mode in a manner that disrupts timing, as shown at 548.
[0073] Once the application 322 is running with the device 300 in the timing test mode, the application may be tested for errors, as shown at 550. Such testing may include, but is not limited to, determining whether an application stops, generates errors, or generates abnormal results (eg, significant IPC changes) that would not occur when the device is operating normally.
As an example of modifying the setting at 542, in a processor architecture of the type shown in FIG. 2, two or more CPU cores may operate at different frequencies, which may be higher than the normal operating frequency of the consumer device high frequency. Similarly, two or more caches within a device may operate at different frequencies in a timing test mode. Also, different combinations of cores and caches can run at different frequencies.
[0075] In other implementations, CPU resources may be reduced when the device operates in a sequential test mode. Examples of such CPU resource reduction include, but are not limited to, reducing the size of a store queue, load queue, or cache (eg, L1 or higher, "cache, D-cache, ITLB, or DTLB). Other examples Including but not limited to reducing the execution rate of ALU, AGU, SIMD tube or specific instructions. In addition, one or more individual cores or application threads can be preempted randomly or methodically. Additional examples include delaying or Speed up or change timing, change OS usage of cores, change virtual
To physical core assignments (eg, inter-cluster contention), exploit other asymmetries or write-backs, or invalidate caches and/or TLBs.
[0076] In other implementations, modifying the settings at 542 may include changing the functionality of the GPU 330. Examples of such modifications include running GPU cores 332 at different frequencies, running one or more of the GPU cores at different frequencies than normal for consumer devices, replacing GPU firmware with firmware having different timings than normal operation of device 300 . One or more of the GPU cores 332 may be configured to selectively operate at a frequency higher or lower than that used in the normal operating mode of the device. Other examples include perturbing GPU firmware (eg, perturbing object processing), and reducing GPU resource reductions such as cache size or execution rate.
[0077] In other implementations, GPU processing can be altered when running the device in sequential test mode, such as by changing wavefront counts via random compute threads, random preempting graphics, or by writing back or invalidating caches and/or TLBs.
[0078] An example of sending a command to a hardware component at 544 in a manner that disrupts timing includes changing the functionality of memory 340 or memory controller 315. Examples of such changes in memory or memory controller functionality include, but are not limited to, running memory clocks and/or internal bus clocks at different frequencies, inserting noise into memory operations, adding latency to memory operations, changing the priority of memory operations, and Change row and/or column channel bits to simulate different channel counts or row breaks. [0079] Aspects of the present disclosure allow a software developer to test the performance of a new application on a previous version of a device. More specifically, aspects of the present disclosure allow developers to probe the impact of timing disturbances on applications.
[0080] While the above is a complete description of the preferred embodiments of the present invention, various alternatives, modifications and equivalents may be utilized. The scope of the invention should, therefore, be determined not with reference to the above description, but should instead be determined with reference to the appended claims, along with their full scope of equivalents. Any feature described herein (whether preferred or not) may be combined with any other feature described herein (whether preferred or not). In the following claims, the indefinite articles "a" or "an" refer to the quantity of one or more of the items following the article, unless expressly stated otherwise. As used herein, in an alternative list of elements, the term "or" is used in an inclusive sense, eg, "X or Y" encompasses only X, only y, or both X and Y together, unless expressly stated otherwise. Two or more of the elements listed as an alternative may be combined. The appended claims should not be construed to include means-plus-function limitations unless such limitation is expressly recited in a given claim using the phrase "means for".
CN 108369552 B
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
29 members in 5 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 14930408 | United States of America | – | |
| 201514930408 | United States of America | A | |
| 201514930408 | United States of America | A | |
| 2016059751 | United States of America | W | |
| 2016059751 | United States of America | W | |
| 14930408 | – | – | – |
| PCTUS2016059751 | – | – | – |
| US201514930408 | – | – | – |
| WO2016US59751 | – | – | – |
Members29
| Document | Office | Kind | |
|---|---|---|---|
| US2017123961A1 | United States of America | A1 | |
| WO2017079089A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US9892024B2 | United States of America | B2 | |
| CN108369552A | China | A | |
| US2018246802A1 | United States of America | A1 | |
| EP3371704A1 | European Patent Office (EPO) | A1 | |
| EP3371704A4 | European Patent Office (EPO) | A4 | |
| JP2020502595A | Japan | A | |
| JP2020113302A | Japan | A | |
| EP3686741A1 | European Patent Office (EPO) | A1 | |
| CN111881013A | China | A | |
| US11042470B2 | United States of America | B2 | |
| EP3371704B1 | European Patent Office (EPO) | B1 | |
| JP6903187B2 | Japan | B2 | |
| JP2021152956A | Japan | A | |
| US2021311856A1 | United States of America | A1 | |
| JP6949857B2 | Japan | B2 | |
| EP3920032A1 | European Patent Office (EPO) | A1 | |
| EP3920032A4 | European Patent Office (EPO) | A4 | |
| CN108369552BThis record | China | B | |
| CN115794598A | China | A | |
| JP7269282B2 | Japan | B2 | |
| JP2023093646A | Japan | A | |
| US11907105B2 | United States of America | B2 | |
| CN111881013B | China | B | |
| US2024211380A1 | United States of America | A1 | |
| EP3686741B1 | European Patent Office (EPO) | B1 | |
| EP3920032B1 | European Patent Office (EPO) | B1 | |
| JP7759357B2 | Japan | B2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent grantGrantedGR01 | GR01 | |
| PublicationPB01 | PB01 |
Numbers
- Publication
- 108369552
- Publication, DOCDB
- 108369552
- Publication, EPODOC
- CN108369552B
- Application
- 80070684
- Application, DOCDB
- 201680070684
- Application, EPODOC
- CN201680070684
Titles2
- Chinese
- 以扰乱时序的模式进行的软件向后兼容性测试
- English
- Software backward compatibility testing in scrambled mode
Classification
- CPC, 14
- G06F11/3684
- G06F11/3668
- G06F11/3688
- G06F12/0811
- G06F12/1027
- G06F12/084
- G06F12/0875
- G06F12/1045
- G06F9/30079
- G06F9/3001
- G06F2212/452
- G06F2212/50
- G06F2212/62
- G06F9/46
- IPC, 1
- G06F11 36