财神捕鱼送注册彩金

書(shū)名： Learning Apache Spark 2
作者名： Muhammad Asif Abbasi
本章字?jǐn)?shù)： 244字
更新時(shí)間： 2021-07-09 18:46:01

How is Spark being used?

Matei Zaharia is the creator of Apache Spark project and co-founder of DataBricks, the company which was formed by the creators of Apache Spark. Matei in his keynote at the Spark summit in Europe during fall of 2015 mentioned some key metrics on how Spark is being used in various runtime environments. The numbers were a bit surprising to me, as I had thought Spark on YARN would have higher numbers than what was presented. Here are the key figures:

Spark in Standalone mode - 48%
Spark on YARN - 40%
Spark on MESOS - 11%

As we can see from the numbers, almost 90% of Apache Spark installations are in standalone mode or on YARN. When Spark is being configured on YARN, we can make an assumption that the organization has chosen Hadoop as their data operating system, and are planning to move their data onto Hadoop, which means our primary source of data ingest might be Hive, HDFS, HBase, or other No SQL systems.

When Apache Spark is installed in standalone mode, the possibility of primary sources increases, but the data on HDFS still remains a huge possibility as it is entirely likely that the customer has a Hadoop installation, but wishes to keep Spark separate as a discovery platform.

Spark can work with a variety of sources. Let's look at the most common sources that we come across:

File Formats
File Systems
Structured Data sources / Databases
Key/Value Stores

官术网_书友最值得收藏!

Learning Apache Spark 2

How is Spark being used?