官术网_书友最值得收藏!

Data caching

Many machine learning algorithms are iterative in nature and thus require multiple passes over the data. However, all data stored in Spark RDD are by default transient, since RDD just stores the transformation to be executed and not the actual data. That means each action would recompute data again and again by executing the transformation stored in RDD.

Hence, Spark provides a way to persist the data in case we need to iterate over it. Spark also publishes several StorageLevels to allow storing data with various options:

  • NONE: No caching at all
  • MEMORY_ONLY: Caches RDD data only in memory
  • DISK_ONLY: Write cached RDD data to a disk and releases from memory
  • MEMORY_AND_DISK: Caches RDD in memory, if it's not possible to offload data to a disk
  • OFF_HEAP: Use external memory storage which is not part of JVM heap

Furthermore, Spark gives users the ability to cache data in two flavors: raw (for example, MEMORY_ONLY) and serialized (for example, MEMORY_ONLY_SER). The later uses large memory buffers to store serialized content of RDD directly. Which one to use is very task and resource dependent. A good rule of thumb is if the dataset you are working with is less than 10 gigs then raw caching is preferred to serialized caching. However, once you cross over the 10 gigs soft-threshold, raw caching imposes a greater memory footprint than serialized caching.

Spark can be forced to cache by calling the cache() method on RDD or directly via calling the method persist with the desired persistent target - persist(StorageLevels.MEMORY_ONLY_SER). It is useful to know that RDD allows us to set up the storage level only once.

The decision on what to cache and how to cache is part of the Spark magic; however, the golden rule is to use caching when we need to access RDD data several times and choose a destination based on the application preference respecting speed and storage. A great blogpost which goes into far more detail than what is given here is available at:

http://sujee.net/2015/01/22/understanding-spark-caching/#.VpU1nJMrLdc

Cached RDDs can be accessed as well from the H2O Flow UI by evaluating the cell with getRDDs:

主站蜘蛛池模板: 阿城市| 上犹县| 常熟市| 永嘉县| 广平县| 泊头市| 临颍县| 老河口市| 大兴区| 深水埗区| 和田市| 万山特区| 银川市| 鹤庆县| 通河县| 县级市| 文成县| 永修县| 从化市| 尼玛县| 库尔勒市| 台山市| 宁乡县| 吉首市| 磴口县| 繁峙县| 桃江县| 乌兰县| 阜城县| 青铜峡市| 天气| 贵德县| 德安县| 怀安县| 潞西市| 绥化市| 沅江市| 平安县| 巴东县| 西乌珠穆沁旗| 兰西县|