官术网_书友最值得收藏!

What this book covers

Chapter 1, The Need for Data Lake, helps you understand what Data Lake is, its architecture and key components, and the business contexts where Data Lake can be successfully deployed. You will also learn the limitations of the traditional data architectures and how Data Lake addresses some of these inadequacies and provides significant benefits.

Chapter 2, Data Intake, helps you understand the Intake Tier in detail where we will explore the process of obtaining huge volumes of data into Data Lake. You will learn the technology perspective of the various External Data Sources and Hadoop-based data transfer mechanisms to pull or push data into Data Lake.

Chapter 3, Data Integration, Quality, and Enrichment, explores the processes that are performed on vast quantities of data in the Management Tier. You will get a deeper understanding of the key technology aspects and components such as profiling, validation, integration, cleansing, standardization, and enrichment using Hadoop ecosystem components.

Chapter 4, Data Discovery and Consumption, helps you understand how data can be discovered, packaged, and provisioned, for it to be consumed by the downstream systems. You will learn the key technology aspects, architectural guidance and tools for data discovery, and data provisioning functionalities.

Chapter 5, Data Governance, explores the details, need, and utility of data governance in a Data Lake environment. You will learn how to deal with metadata management, lineage tracking, data lifecycle management to govern the usability, security, integrity, and availability of the data through the data governance processes applied on the data in Data Lake. This chapter also explores how the current Data Lake can evolve in a futuristic setting.

主站蜘蛛池模板: 建瓯市| 富平县| 弥渡县| 红原县| 沙雅县| 海盐县| 会理县| 余庆县| 玉门市| 慈溪市| 巴东县| 礼泉县| 谢通门县| 莱阳市| 山西省| 思茅市| 长沙县| 黄石市| 腾冲县| 米脂县| 长岛县| 石门县| 新河县| 黄陵县| 天等县| 紫金县| 通渭县| 大埔区| 新龙县| 松江区| 绥阳县| 贵德县| 玉门市| 富源县| 乐安县| 雷山县| 淮滨县| 绥江县| 新昌县| 巴青县| 象州县|