官术网_书友最值得收藏!

The basics of web requests

The worldwide capacity to generate data is estimated to double in size every two years. Even though there is an interdisciplinary field known as data science that is entirely dedicated to the study of data, almost every programming task in software development also has something to do with collecting and analyzing data. A significant part of this is, of course, data collection. However, the data that we need for our applications is sometimes not stored nicely and cleanly in a database—sometimes, we need to collect the data we need from web pages.

For example, web scraping is a data extraction method that automatically makes requests to web pages and downloads specific information. Web scraping allows us to comb through numerous websites and collect any data we need in a systematic and consistent manner—the collected data can be analyzed later on by our applications or simply saved on our computers in various formats. An example of this would be Google, which programs and runs numerous web scrapers of its own to find and index web pages for the search engine.

The Python language itself provides a number of good options for applications of this kind. In this chapter, we will mainly work with the requests module to make client-side web requests from our Python programs. However, before we look into this module in more detail, we need to understand some web terminology in order to be able to effectively design our applications.

主站蜘蛛池模板: 湘潭县| 垦利县| 大田县| 大邑县| 达日县| 来安县| 肇源县| 扎囊县| 远安县| 营口市| 遵义县| 新竹县| 亳州市| 金秀| 西平县| 新绛县| 义马市| 柳州市| 新疆| 双鸭山市| 固安县| 南城县| 息烽县| 蒙阴县| 尤溪县| 沁水县| 农安县| 田阳县| 玛多县| 巴楚县| 大姚县| 吉首市| 德州市| 兴安盟| 沂水县| 达州市| 南康市| 合阳县| 伊金霍洛旗| 伽师县| 石屏县|