官术网_书友最值得收藏!

Chapter 1.  The Big Data Science Ecosystem

As a data scientist, you'll no doubt be very familiar with handling files and processing perhaps even large amounts of data. However, as I'm sure you will agree, doing anything more than a simple analysis over a single type of data requires a method of organizing and cataloguing data so that it can be managed effectively. Indeed, this is the cornerstone of a great data scientist. As the data volume and complexity increases, a consistent and robust approach can be the difference between generalized success and over-fitted failure!

This chapter is an introduction to an approach and ecosystem for achieving success with data at scale. It focuses on the data science tools and technologies. It introduces the environment, and how to configure it appropriately, but also explains some of the nonfunctional considerations relevant to the overall data architecture. While there is little actual data science at this stage, it provides the essential platform to pave the way for success in the rest of the book.

In this chapter, we will cover the following topics:

  • Data management responsibilities
  • Data architecture
  • Companion tools
主站蜘蛛池模板: 基隆市| 灵寿县| 延边| 资源县| 呼玛县| 桂阳县| 宝丰县| 栾川县| 瑞昌市| 桐城市| 旅游| 平乐县| 东至县| 永靖县| 左权县| 麻栗坡县| 景德镇市| 鄂托克前旗| 喀喇| 横峰县| 策勒县| 肇州县| 齐河县| 贵德县| 葫芦岛市| 吴旗县| 正蓝旗| 泰兴市| 宜丰县| 宝清县| 舞阳县| 崇州市| 兴业县| 府谷县| 德昌县| 大竹县| 梓潼县| 利川市| 新建县| 镇赉县| 苗栗县|