官术网_书友最值得收藏!

Search engines

One well-known use case for web scraping is indexing websites for the purpose of building a search engine. In this case, a web scraper would visit different websites and follow references to other websites in order to discover all of the content available on the internet. By collecting some of the content from the pages, you could respond to search queries by matching the terms to the contents of the pages you have collected. You could also suggest similar pages if you track how pages are linked together, and rank the most important pages by the number of connections they have to other sites.

Googlebot is the most famous example of a web scraper used to build a search engine. It is the first step in building the search engine as it downloads, indexes, and ranks each page on a website. It will also follow links to other websites, which is how it is able to index a substantial portion of the internet. According to Googlebot's documentation, the scraper attempts to reach each web page every few seconds, which requires them to reach estimates of well into billions of pages per day!

If your goal is to build a search engine, albeit on a much smaller scale, you will find enough tools in this book to collect the information you need. This book will not, however, cover indexing and ranking pages to provide relevant search results.

主站蜘蛛池模板: 宣化县| 廊坊市| 柏乡县| 池州市| 洛隆县| 泽库县| 玉溪市| 莒南县| 称多县| 永平县| 怀远县| 襄汾县| 广丰县| 宝应县| 绥江县| 专栏| 洮南市| 乐安县| 尖扎县| 兴仁县| 铜鼓县| 云和县| 东阳市| 兴国县| 香河县| 桐乡市| 曲阳县| 沧州市| 大冶市| 綦江县| 高要市| 西峡县| 凤台县| 弥勒县| 政和县| 镇巴县| 芜湖市| 嵩明县| 漳州市| 丹凤县| 永善县|