官术网_书友最值得收藏!

Chapter 4. Advanced Word2vec

In Chapter 3, Word2vec – Learning Word Embeddings, we introduced you to Word2vec, the basics of learning word embeddings, and the two common Word2vec algorithms: skip-gram and CBOW. In this chapter, we will discuss several topics related to Word2vec, focusing on these two algorithms and extensions.

First, we will explore how the original skip-gram algorithm was implemented and how it compares to its more modern variant, which we used in Chapter 3, Word2vec – Learning Word Embeddings. We will examine the differences between skip-gram and CBOW and look at the behavior of the loss over time of the two approaches. We will also discuss which method works better, using both our observation and the available literature.

We will discuss several extensions to the existing Word2vec methods that boost performance. These extensions include using more effective sampling techniques to sample negative examples for negative sampling and ignoring uninformative words in the learning process, among others. You will also learn a novel word embedding learning technique known as Global Vectors (GloVe) and the specific advantages that GloVe has over skip-gram and CBOW.

Finally, you will learn how to use Word2vec to solve a real-world problem: document classification. We will see this with a simple trick of obtaining document embeddings from word embeddings.

主站蜘蛛池模板: 巍山| 汤原县| 汕尾市| 商城县| 溧阳市| 广河县| 黄浦区| 丰都县| 厦门市| 许昌县| 盈江县| 克拉玛依市| 中山市| 广河县| 泸西县| 江口县| 皋兰县| 顺平县| 东阳市| 会昌县| 永丰县| 怀仁县| 饶阳县| 冷水江市| 平阴县| 衢州市| 周宁县| 门源| 林周县| 廉江市| 洞头县| 龙江县| 门头沟区| 洛南县| 盘山县| 昂仁县| 进贤县| 宜章县| 剑阁县| 武平县| 开远市|