In this article, wordvector is used to train doc2vec on quanteda’s tokens objects.
Prepare data
Download the corpus of news summaries.
# Load data
f <- tempfile()
download.file('https://www.dropbox.com/s/e19kslwhuu9yc2z/yahoo-news.RDS?dl=1',
f, mode = "wb")
library(quanteda)
## Package version: 4.4
## Unicode version: 15.1
## ICU version: 74.2
## Parallel computing: disabled
## See https://quanteda.io for tutorials and examples.
library(wordvector)
quanteda_options(verbose = TRUE)
# Construct corpus
dat <- readRDS(f)
dat$text <- paste0(dat$head, ". ", dat$body)
corp <- corpus(dat, text_field = 'text')
# Tokenize
toks <- tokens(corp, remove_punct = TRUE, remove_symbols = TRUE) %>%
tokens_remove(stopwords("en", "marimo"), padding = TRUE) %>%
tokens_select("^[a-zA-Z-]+$", valuetype = "regex", case_insensitive = FALSE,
padding = TRUE)
## Creating a tokens from a corpus object...
## ...starting tokenization
## ...tokenizing 1 of 7 blocks
## ...preserving hyphens
## ...preserving elisions
## ...preserving social media tags (#, @)
## ...tokenizing 2 of 7 blocks
## ...preserving hyphens
## ...preserving elisions
## ...preserving social media tags (#, @)
## ...tokenizing 3 of 7 blocks
## ...preserving hyphens
## ...preserving elisions
## ...preserving social media tags (#, @)
## ...tokenizing 4 of 7 blocks
## ...preserving hyphens
## ...preserving elisions
## ...preserving social media tags (#, @)
## ...tokenizing 5 of 7 blocks
## ...preserving hyphens
## ...preserving elisions
## ...preserving social media tags (#, @)
## ...tokenizing 6 of 7 blocks
## ...preserving hyphens
## ...preserving elisions
## ...preserving social media tags (#, @)
## ...tokenizing 7 of 7 blocks
## ...preserving hyphens
## ...preserving elisions
## ...preserving social media tags (#, @)
## ...removing separators, punctuation, symbols
## ...298,565 unique types
## ...complete, elapsed time: 94.7 seconds.
## Finished constructing tokens from 656,334 documents
## tokens_remove() changed from 45,194,192 tokens (656,334 documents) to 28,232,774 tokens (656,334 documents)
## tokens_keep() changed from 28,232,774 tokens (656,334 documents) to 26,564,509 tokens (656,334 documents)Train doc2vec
textmodel_doc2vec() supports both dm
(distributed memory) and dbow (distributed bag-of-words)
models.
# Set the number of processors
options(wordvector_threads = 16)
# Train doc2vec
dov <- textmodel_doc2vec(toks, dim = 50, type = "dm", min_count = 5, verbose = TRUE)
## Training distributed memory model with 50 dimensions
## ...using 16 threads for distributed computing
## ...initializing
## ...negative sampling in 10 iterations
## ......iteration 1 elapsed time: 17.64 seconds (alpha: 0.0455)
## ......iteration 2 elapsed time: 34.94 seconds (alpha: 0.0410)
## ......iteration 3 elapsed time: 52.53 seconds (alpha: 0.0365)
## ......iteration 4 elapsed time: 69.95 seconds (alpha: 0.0320)
## ......iteration 5 elapsed time: 88.79 seconds (alpha: 0.0272)
## ......iteration 6 elapsed time: 106.46 seconds (alpha: 0.0226)
## ......iteration 7 elapsed time: 123.50 seconds (alpha: 0.0182)
## ......iteration 8 elapsed time: 144.91 seconds (alpha: 0.0132)
## ......iteration 9 elapsed time: 164.26 seconds (alpha: 0.0084)
## ......iteration 10 elapsed time: 182.18 seconds (alpha: 0.0039)
## ...completeSince the distributed memory model has hidden layers for documents
and words, you can extract document and word vectors using
as.matrix()
Predict probability
If probabitliy() is applied to a fitted doc2vec model,
you receive the predicted probability of the words in each document.
head(probability(dov, c("bad", "good"), mode = "numeric", layer = "documents"))
## bad good
## text1 0.2177387 0.29853595
## text2 0.3888728 0.31369409
## text3 0.1390950 0.06860224
## text4 0.5840342 0.83973212
## text5 0.4504653 0.42143507
## text6 0.6455948 0.57977184