官术网_书友最值得收藏!

  • Apache Spark Quick Start Guide
  • Shrey Mehrotra Akash Grade
  • 352字
  • 2021-07-02 13:39:53

What is Spark?

Apache Spark is a distributed computing framework which makes big-data processing quite easy, fast, and scalable. You must be wondering what makes Spark so popular in the industry, and how is it really different than the existing tools available for big-data processing? The reason is that it provides a unified stack for processing all different kinds of big data, be it batch, streaming, machine learning, or graph data.

Spark was developed at UC Berkeley’s AMPLab in 2009 and later came under the Apache Umbrella in 2010. The framework is mainly written in Scala and Java. 

Spark provides an interface with many different distributed and non-distributed data stores, such as Hadoop Distributed File System (HDFS), Cassandra, Openstack Swift, Amazon S3, and Kudu. It also provides a wide variety of language APIs to perform analytics on the data stored in these data stores. These APIs include Scala, Java, Python, and R.

The basic entity of Spark is Resilient Distributed Dataset (RDD), which is a read-only partitioned collection of data. RDD can be created using data stored on different data stores or using existing RDD. We shall discuss this in more detail in Chapter 3Spark RDD.

Spark needs a resource manager to distribute and execute its tasks. By default, Spark comes up with its own standalone scheduler, but it integrates easily with Apache Mesos and Yet Another Resource Negotiator (YARN) for cluster resource management and task execution.

One of the main features of Spark is to keep a large amount of data in memory for faster execution. It also has a component that generates a Directed Acyclic Graph (DAG) of operations based on the user program. We shall discuss these in more details in coming chapters.

The following diagram shows some of the popular data stores Spark can connect to:

Data stores 
Spark is a computing engine, and should not be considered as a storage system as well. Spark is also not designed for cluster management. For this purpose, frameworks such as Mesos and YARN are used. 
主站蜘蛛池模板: 淮北市| 潍坊市| 潮州市| 曲周县| 定陶县| 霍邱县| 交口县| 海淀区| 临高县| 静安区| 报价| 高邮市| 股票| 彝良县| 凤凰县| 沭阳县| 静宁县| 舞阳县| 化州市| 新晃| 朝阳区| 南雄市| 博白县| 全椒县| 锡林郭勒盟| 怀仁县| 丰城市| 乌拉特前旗| 三明市| 银川市| 忻城县| 西贡区| 綦江县| 高州市| 松滋市| 宝山区| 抚宁县| 内黄县| 喜德县| 泸溪县| 哈巴河县|