官术网_书友最值得收藏!

Summary

In this chapter, we covered how to set up Spark locally on our own computer as well as in the cloud as a cluster running on Amazon EC2. You learned how to run Spark on top of Amazon's Elastic Map Reduce (EMR). You also learned how to use Google Compute Engine's Spark Service to create a cluster and run a simple job. We discussed the basics of Spark's programming model and API using the interactive Scala console, and we wrote the same basic Spark program in Scala, Java, R, and Python. We also compared the performance metrics of Hadoop versus Spark for different machine learning algorithms as well as SORT benchmark tests.

In the next chapter, we will consider how to go about using Spark to create a machine learning system.

主站蜘蛛池模板: 句容市| 彭泽县| 铜陵市| 盐城市| 阿鲁科尔沁旗| 工布江达县| 佛冈县| 松溪县| 普洱| 罗平县| 临清市| 古交市| 长宁区| 休宁县| 广元市| 海兴县| 资源县| 翁牛特旗| 诸暨市| 凭祥市| 平罗县| 宝山区| 安丘市| 龙山县| 昭觉县| 台北县| 海兴县| 榆中县| 临潭县| 石门县| 广东省| 吉安县| 哈尔滨市| 宁城县| 印江| 来宾市| 东平县| 邵武市| 克什克腾旗| 寻乌县| 射阳县|