官术网_书友最值得收藏!

Chapter 4. Advanced Word2vec

In Chapter 3, Word2vec – Learning Word Embeddings, we introduced you to Word2vec, the basics of learning word embeddings, and the two common Word2vec algorithms: skip-gram and CBOW. In this chapter, we will discuss several topics related to Word2vec, focusing on these two algorithms and extensions.

First, we will explore how the original skip-gram algorithm was implemented and how it compares to its more modern variant, which we used in Chapter 3, Word2vec – Learning Word Embeddings. We will examine the differences between skip-gram and CBOW and look at the behavior of the loss over time of the two approaches. We will also discuss which method works better, using both our observation and the available literature.

We will discuss several extensions to the existing Word2vec methods that boost performance. These extensions include using more effective sampling techniques to sample negative examples for negative sampling and ignoring uninformative words in the learning process, among others. You will also learn a novel word embedding learning technique known as Global Vectors (GloVe) and the specific advantages that GloVe has over skip-gram and CBOW.

Finally, you will learn how to use Word2vec to solve a real-world problem: document classification. We will see this with a simple trick of obtaining document embeddings from word embeddings.

主站蜘蛛池模板: 鲁山县| 大厂| 邯郸县| 肃北| 内乡县| 津南区| 彭泽县| 容城县| 玛多县| 禹城市| 韶关市| 邢台县| 洪洞县| 亚东县| 长沙市| 平舆县| 德保县| 宣汉县| 集安市| 新竹县| 宽城| 沅陵县| 金溪县| 灵宝市| 赤峰市| 安新县| 乌拉特中旗| 海林市| 开鲁县| 松江区| 仙居县| 麻城市| 吴江市| 梓潼县| 墨竹工卡县| 浪卡子县| 佛学| 手游| 邳州市| 托克逊县| 包头市|