Article
Duplicate In-memory Shared-intermediate Data Detection and Reuse Module in Spark Framework
US Patent US15/404100, US20180067861A1
(2017)
Abstract
A cache management system for managing a plurality of intermediate data includes a processor, and a memory having stored thereon the plurality of intermediate data and instructions that when executed by the processor, cause the processor to perform identifying a new intermediate data to be accessed, loading the intermediate data from the memory in response to identifying the new intermediate data as one of the plurality of intermediate data, and in response to not identifying the new intermediate data as one of the plurality of intermediate data identifying a reusable intermediate data having a longest duplicate generating logic chain that is at least in part the same as a generating logic chain of the new intermediate data, and generating the new intermediate data from the reusable intermediate data and a portion of the generating logic chain of the new intermediate data not in common with the reusable intermediate data.
Disciplines
Publication Date
2017
Citation Information
Zhengyu Yang, Jiayin Wang and David Evans. "Duplicate In-memory Shared-intermediate Data Detection and Reuse Module in Spark Framework" US Patent US15/404100, US20180067861A1 (2017) Available at: http://works.bepress.com/zhengyuyang/32/