NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2320 most downloaded on PyPI
Python framework for fast Vector Space Modelling
Last release 11 months ago
18 Oct 2025
Ships fairly regularly
a new release about every 7 months
Nearly every release is documented
notes for 59 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
17 years old
80 releases · first in 2010
Removed deprecated usage of scipy.sparsetools to e.g., scipy.sparse.csc_matvecs , scipy.sparse.csc_matrix . ( hechth , julianpollmann , #3615 )
keyedvectors.py (e.g., get_mean_vector, sort_by_descending_frequency, save_word2vec_format). (julianpollmann, #3615)ldamodel.py update_dir_prior(). (julianpollmann, #3615)scipy.sparsetools to e.g., scipy.sparse.csc_matvecs, scipy.sparse.csc_matrix. (hechth, julianpollmann, #3615)nogil, noexcept, type definitions). (julianpollmann, #3615)One column per quarter.
Import deprecated scipy.linalg.triu from numpy.triu instead (__Luffy610__, #3524)
Fix incorrect conversion of cosine distance to cosine similarity ( monash849 , #3441 )
hs and negative in Word2Vec (gau-nernst, #3443)### :red_circle: Bug fixes * #3447: Remove unused FuzzyTM dependency, handle ImportError, by @mpenkov * #3441: Fix changed calculation of cosine dista
hs and negative in Word2Vec, by @gau-nernstfix deprecation warning from pytest by @martino-vic in #3354
Morfessor, tox and gensim.models.wrappers by @pabs3 in #3345wmdistance by @TLouf in #3327Full Changelog: 4.2.0...4.3.0
A number of incremental improvements, optimizations and bugfixes: CHANGELOG
A number of incremental improvements, optimizations and bugfixes: CHANGELOG
encoding parameter to TextDirectoryCorpus, by @Sandman-Renbuild_vocab and train to use correct argument names, by @HLassestr() method in WmdSimilarity, by @DingQKf prefix on f-strings fix, by @code-review-doctortest_translate_gc on OSX + py3.9, by @menshikh-ivDeprecated obsolete step parameter from doc2vec
This is a bugfix release that addresses left over compatibility issues with older versions of numpy and MacOS.
This is a bugfix release that addresses compatibility issues with older versions of numpy.
Gensim 4.1 brings two major new functionalities:
There are several minor changes that are not backwards compatible with previous versions of Gensim.
The affected functionality is relatively less used, so it is unlikely to affect most users, so we have opted to not require a major version bump.
Nevertheless, we describe them below.
We now handle both positive and negative keyword parameters consistently.
They may now be either:
So you can now simply do:
model.most_similar(positive='war', negative='peace')instead of the slightly more involved
model.most_similar(positive=['war'], negative=['peace'])Both invocations remain correct, so you can use whichever is most convenient.
If you were somehow expecting gensim to interpret the strings as a list of characters, e.g.
model.most_similar(positive=['w', 'a', 'r'], negative=['p', 'e', 'a', 'c', 'e'])then you will need to specify the lists explicitly in gensim 4.1.
step parameter from doc2vecWith the newer version, do this:
model.infer_vector(..., epochs=123)instead of this:
model.infer_vector(..., steps=123)Plus a large number of smaller improvements and fixes, as usual.
⚠️ If migrating from old Gensim 3.x, read the Migration guide first.
shrink_windows argument for Word2Vec, by @M-DemayDeprecated obsolete step parameter from doc2vec
This is a bugfix release that addresses compatibility issues with older versions of numpy.
Gensim 4.1 brings two major new functionalities:
There are several minor changes that are not backwards compatible with previous versions of Gensim.
The affected functionality is relatively less used, so it is unlikely to affect most users, so we have opted to not require a major version bump.
Nevertheless, we describe them below.
We now handle both positive and negative keyword parameters consistently.
They may now be either:
So you can now simply do:
model.most_similar(positive='war', negative='peace')instead of the slightly more involved
model.most_similar(positive=['war'], negative=['peace'])Both invocations remain correct, so you can use whichever is most convenient.
If you were somehow expecting gensim to interpret the strings as a list of characters, e.g.
model.most_similar(positive=['w', 'a', 'r'], negative=['p', 'e', 'a', 'c', 'e'])then you will need to specify the lists explicitly in gensim 4.1.
step parameter from doc2vecWith the newer version, do this:
model.infer_vector(..., epochs=123)instead of this:
model.infer_vector(..., steps=123)Plus a large number of smaller improvements and fixes, as usual.
⚠️ If migrating from old Gensim 3.x, read the Migration guide first.
shrink_windows argument for Word2Vec, by @M-DemayDeprecated obsolete step parameter from doc2vec
Gensim 4.1 brings two major new functionalities:
There are several minor changes that are not backwards compatible with previous versions of Gensim.
The affected functionality is relatively less used, so it is unlikely to affect most users, so we have opted to not require a major version bump.
Nevertheless, we describe them below.
We now handle both positive and negative keyword parameters consistently.
They may now be either:
So you can now simply do:
model.most_similar(positive='war', negative='peace')instead of the slightly more involved
model.most_similar(positive=['war'], negative=['peace'])Both invocations remain correct, so you can use whichever is most convenient.
If you were somehow expecting gensim to interpret the strings as a list of characters, e.g.
model.most_similar(positive=['w', 'a', 'r'], negative=['p', 'e', 'a', 'c', 'e'])then you will need to specify the lists explicitly in gensim 4.1.
step parameter from doc2vecWith the newer version, do this:
model.infer_vector(..., epochs=123)instead of this:
model.infer_vector(..., steps=123)Plus a large number of smaller improvements and fixes, as usual.
⚠️ If migrating from old Gensim 3.x, read the Migration guide first.
shrink_windows argument for Word2Vec, by @M-DemayBugfix release to address issues with Wheels on Windows:
⚠️ Gensim 4.0 contains breaking API changes! See the Migration guide to update your existing Gensim 3.x code and models.
Gensim 4.0 is a major release with lots of performance & robustness improvements, and a new website.
Massively optimized popular algorithms the community has grown to love: fastText, word2vec, doc2vec, phrases:
a. Efficiency
| model | 3.8.3: wall time / peak RAM / throughput | 4.0.0: wall time / peak RAM / throughput |
|---|---|---|
| fastText | 2.9h / 4.11 GB / 822k words/s | 2.3h / 1.26 GB / 914k words/s |
| word2vec | 1.7h / 0.36 GB / 1685k words/s | 1.2h / 0.33 GB / 1762k words/s |
In other words, fastText now needs 3x less RAM (and is faster); word2vec has 2x faster init (and needs less RAM, and is faster); detecting collocation phrases is 2x faster. (4.0 benchmarks)
b. Robustness. We fixed a bunch of long-standing bugs by refactoring the internal code structure (see 🔴 Bug fixes below)
c. Simplified OOP model for easier model exports and integration with TensorFlow, PyTorch &co.
These improvements come to you transparently aka "for free", but see Migration guide for some changes that break the old Gensim 3.x API. Update your code accordingly.
Dropped a bunch of externally contributed modules and wrappers: summarization, pivoted TFIDF, Mallet…
Code quality was not up to our standards. Also there was no one to maintain these modules, answer user questions, support them.
So rather than let them rot, we took the hard decision of removing these contributed modules from Gensim. If anyone's interested in maintaining them, please fork & publish into your own repo. They can live happily outside of Gensim.
Dropped Python 2. Gensim 4.0 is Py3.6+. Read our Python version support policy.
A new Gensim website – finally! 🙃
So, a major clean-up release overall. We're happy with this tighter, leaner and faster Gensim.
This is the direction we'll keep going forward: less kitchen-sink of "latest academic algorithms", more focus on robust engineering, targetting concrete NLP & document similarity use-cases.
max_final_vocab parameter in fastText constructor, by @mpenkovalpha parameter in LDA model, by @xh2save_facebook_model failure after update-vocab & other initialization streamlining, by @gojomoxml.etree.cElementTree, by @hugovksimilarities.index to the more appropriate similarities.annoy, by @piskvorkynum_words to topn in dtm_coherence, by @MeganStodelon_batch_begin and on_batch_end callbacks, by @mpenkovpattern dependency, by @mpenkovgensim.viz subpackage, by @mpenkov⚠️ Gensim 4.0 contains breaking API changes! See the Migration guide to update your existing Gensim 3.x code and models.
Gensim 4.0 is a major release with lots of performance & robustness improvements and a new website.
Massively optimized popular algorithms the community has grown to love: fastText, word2vec, doc2vec, phrases:
a. Efficiency
| model | 3.8.3: wall time / peak RAM / throughput | 4.0.0: wall time / peak RAM / throughput |
|---|---|---|
| fastText | 2.9h / 4.11 GB / 822k words/s | 2.3h / 1.26 GB / 914k words/s |
| word2vec | 1.7h / 0.36 GB / 1685k words/s | 1.2h / 0.33 GB / 1762k words/s |
In other words, fastText now needs 3x less RAM (and is faster); word2vec has 2x faster init (and needs less RAM, and is faster); detecting collocation phrases is 2x faster. (4.0 benchmarks)
b. Robustness. We fixed a bunch of long-standing bugs by refactoring the internal code structure (see 🔴 Bug fixes below)
c. Simplified OOP model for easier model exports and integration with TensorFlow, PyTorch &co.
These improvements come to you transparently aka "for free", but see Migration guide for some changes that break the old Gensim 3.x API. Update your code accordingly.
Dropped a bunch of externally contributed modules: summarization, pivoted TFIDF normalization, FIXME.
Code quality was not up to our standards. Also there was no one to maintain them, answer user questions, support these modules.
So rather than let them rot, we took the hard decision of removing these contributed modules from Gensim. If anyone's interested in maintaining them please fork into your own repo, they can live happily outside of Gensim.
Dropped Python 2. Gensim 4.0 is Py3.6+. Read our Python version support policy.
A new Gensim website – finally! 🙃
So, a major clean-up release overall. We're happy with this tighter, leaner and faster Gensim.
This is the direction we'll keep going forward: less kitchen-sink of "latest academic algorithms", more focus on robust engineering, targetting common concrete NLP & document similarity use-cases.
⚠️ Gensim 4.0 contains breaking API changes! See the Migration guide to update your existing Gensim 3.x code and models.
Gensim 4.0 is a major release with lots of performance & robustness improvements and a new website.
Massively optimized popular algorithms the community has grown to love: fastText, word2vec, doc2vec, phrases:
a. Efficiency
| model | 3.8.3: wall time / peak RAM / throughput | 4.0.0: wall time / peak RAM / throughput |
|---|---|---|
| fastText | 2.9h / 4.11 GB / 822k words/s | 2.3h / 1.26 GB / 914k words/s |
| word2vec | 1.7h / 0.36 GB / 1685k words/s | 1.2h / 0.33 GB / 1762k words/s |
In other words, fastText now needs 3x less RAM (and is faster); word2vec has 2x faster init (and needs less RAM, and is faster); detecting collocation phrases is 2x faster. (4.0 benchmarks)
b. Robustness. We fixed a bunch of long-standing bugs by refactoring the internal code structure (see 🔴 Bug fixes below)
c. Simplified OOP model for easier model exports and integration with TensorFlow, PyTorch &co.
These improvements come to you transparently aka "for free", but see Migration guide for some changes that break the old Gensim 3.x API. Update your code accordingly.
Dropped a bunch of externally contributed modules: summarization, pivoted TFIDF normalization, FIXME.
Code quality was not up to our standards. Also there was no one to maintain them, answer user questions, support these modules.
So rather than let them rot, we took the hard decision of removing these contributed modules from Gensim. If anyone's interested in maintaining them please fork into your own repo, they can live happily outside of Gensim.
Dropped Python 2. Gensim 4.0 is Py3.6+. Read our Python version support policy.
A new Gensim website – finally! 🙃
So, a major clean-up release overall. We're happy with this tighter, leaner and faster Gensim.
This is the direction we'll keep going forward: less kitchen-sink of "latest academic algorithms", more focus on robust engineering, targetting common concrete NLP & document similarity use-cases.
This 4.0.0beta pre-release is for users who want the cutting edge performance and bug fixes. Plus users who want to help out, by testing and providing feedback: code, documentation, workflows… Please let us know on the mailing list!
Install the pre-release with:
pip install --pre --upgrade gensimProduction stability is important to Gensim, so we're improving the process of upgrading already-trained saved models. There'll be an explicit model upgrade script between each 4.n to 4.(n+1) Gensim release. Check progress here.
max_final_vocab parameter in fastText constructor, by @mpenkovalpha parameter in LDA model, by @xh2save_facebook_model failure after update-vocab & other initialization streamlining, by @gojomoxml.etree.cElementTree, by @hugovksimilarities.index to the more appropriate similarities.annoy, by @piskvorkynum_words to topn in dtm_coherence, by @MeganStodelFix backward incompatibility for LdaModel ( @chinmayapancholi13 , #1327 )
Bugfix release to address issues with wheels on Windows due to Numpy binary incompatibility:
⚠️ Gensim 4.0 contains breaking API changes! See the Migration guide to update your existing Gensim 3.x code and models.
Gensim 4.0 is a major release with lots of performance & robustness improvements, and a new website.
Massively optimized popular algorithms the community has grown to love: fastText, word2vec, doc2vec, phrases:
a. Efficiency
| model | 3.8.3: wall time / peak RAM / throughput | 4.0.0: wall time / peak RAM / throughput |
|---|---|---|
| fastText | 2.9h / 4.11 GB / 822k words/s | 2.3h / 1.26 GB / 914k words/s |
| word2vec | 1.7h / 0.36 GB / 1685k words/s | 1.2h / 0.33 GB / 1762k words/s |
In other words, fastText now needs 3x less RAM (and is faster); word2vec has 2x faster init (and needs less RAM, and is faster); detecting collocation phrases is 2x faster. (4.0 benchmarks)
b. Robustness. We fixed a bunch of long-standing bugs by refactoring the internal code structure (see 🔴 Bug fixes below)
c. Simplified OOP model for easier model exports and integration with TensorFlow, PyTorch &co.
These improvements come to you transparently aka "for free", but see Migration guide for some changes that break the old Gensim 3.x API. Update your code accordingly.
Dropped a bunch of externally contributed modules and wrappers: summarization, pivoted TFIDF, Mallet…
Code quality was not up to our standards. Also there was no one to maintain these modules, answer user questions, support them.
So rather than let them rot, we took the hard decision of removing these contributed modules from Gensim. If anyone's interested in maintaining them, please fork & publish into your own repo. They can live happily outside of Gensim.
Dropped Python 2. Gensim 4.0 is Py3.6+. Read our Python version support policy.
A new Gensim website – finally! 🙃
So, a major clean-up release overall. We're happy with this tighter, leaner and faster Gensim.
This is the direction we'll keep going forward: less kitchen-sink of "latest academic algorithms", more focus on robust engineering, targetting concrete NLP & document similarity use-cases.
max_final_vocab parameter in fastText constructor, by @mpenkovalpha parameter in LDA model, by @xh2save_facebook_model failure after update-vocab & other initialization streamlining, by @gojomoxml.etree.cElementTree, by @hugovksimilarities.index to the more appropriate similarities.annoy, by @piskvorkynum_words to topn in dtm_coherence, by @MeganStodelon_batch_begin and on_batch_end callbacks, by @mpenkovpattern dependency, by @mpenkovgensim.viz subpackage, by @mpenkov⚠️ Gensim 4.0 contains breaking API changes! See the Migration guide to update your existing Gensim 3.x code and models.
Gensim 4.0 is a major release with lots of performance & robustness improvements and a new website.
Massively optimized popular algorithms the community has grown to love: fastText, word2vec, doc2vec, phrases:
a. Efficiency
| model | 3.8.3: wall time / peak RAM / throughput | 4.0.0: wall time / peak RAM / throughput |
|---|---|---|
| fastText | 2.9h / 4.11 GB / 822k words/s | 2.3h / 1.26 GB / 914k words/s |
| word2vec | 1.7h / 0.36 GB / 1685k words/s | 1.2h / 0.33 GB / 1762k words/s |
In other words, fastText now needs 3x less RAM (and is faster); word2vec has 2x faster init (and needs less RAM, and is faster); detecting collocation phrases is 2x faster. (4.0 benchmarks)
b. Robustness. We fixed a bunch of long-standing bugs by refactoring the internal code structure (see 🔴 Bug fixes below)
c. Simplified OOP model for easier model exports and integration with TensorFlow, PyTorch &co.
These improvements come to you transparently aka "for free", but see Migration guide for some changes that break the old Gensim 3.x API. Update your code accordingly.
Dropped a bunch of externally contributed modules: summarization, pivoted TFIDF normalization, FIXME.
Code quality was not up to our standards. Also there was no one to maintain them, answer user questions, support these modules.
So rather than let them rot, we took the hard decision of removing these contributed modules from Gensim. If anyone's interested in maintaining them please fork into your own repo, they can live happily outside of Gensim.
Dropped Python 2. Gensim 4.0 is Py3.6+. Read our Python version support policy.
A new Gensim website – finally! 🙃
So, a major clean-up release overall. We're happy with this tighter, leaner and faster Gensim.
This is the direction we'll keep going forward: less kitchen-sink of "latest academic algorithms", more focus on robust engineering, targetting common concrete NLP & document similarity use-cases.
⚠️ Gensim 4.0 contains breaking API changes! See the Migration guide to update your existing Gensim 3.x code and models.
Gensim 4.0 is a major release with lots of performance & robustness improvements and a new website.
Massively optimized popular algorithms the community has grown to love: fastText, word2vec, doc2vec, phrases:
a. Efficiency
| model | 3.8.3: wall time / peak RAM / throughput | 4.0.0: wall time / peak RAM / throughput |
|---|---|---|
| fastText | 2.9h / 4.11 GB / 822k words/s | 2.3h / 1.26 GB / 914k words/s |
| word2vec | 1.7h / 0.36 GB / 1685k words/s | 1.2h / 0.33 GB / 1762k words/s |
In other words, fastText now needs 3x less RAM (and is faster); word2vec has 2x faster init (and needs less RAM, and is faster); detecting collocation phrases is 2x faster. (4.0 benchmarks)
b. Robustness. We fixed a bunch of long-standing bugs by refactoring the internal code structure (see 🔴 Bug fixes below)
c. Simplified OOP model for easier model exports and integration with TensorFlow, PyTorch &co.
These improvements come to you transparently aka "for free", but see Migration guide for some changes that break the old Gensim 3.x API. Update your code accordingly.
Dropped a bunch of externally contributed modules: summarization, pivoted TFIDF normalization, FIXME.
Code quality was not up to our standards. Also there was no one to maintain them, answer user questions, support these modules.
So rather than let them rot, we took the hard decision of removing these contributed modules from Gensim. If anyone's interested in maintaining them please fork into your own repo, they can live happily outside of Gensim.
Dropped Python 2. Gensim 4.0 is Py3.6+. Read our Python version support policy.
A new Gensim website – finally! 🙃
So, a major clean-up release overall. We're happy with this tighter, leaner and faster Gensim.
This is the direction we'll keep going forward: less kitchen-sink of "latest academic algorithms", more focus on robust engineering, targetting common concrete NLP & document similarity use-cases.
This 4.0.0beta pre-release is for users who want the cutting edge performance and bug fixes. Plus users who want to help out, by testing and providing feedback: code, documentation, workflows… Please let us know on the mailing list!
Install the pre-release with:
pip install --pre --upgrade gensimProduction stability is important to Gensim, so we're improving the process of upgrading already-trained saved models. There'll be an explicit model upgrade script between each 4.n to 4.(n+1) Gensim release. Check progress here.
max_final_vocab parameter in fastText constructor, by @mpenkovalpha parameter in LDA model, by @xh2save_facebook_model failure after update-vocab & other initialization streamlining, by @gojomoxml.etree.cElementTree, by @hugovksimilarities.index to the more appropriate similarities.annoy, by @piskvorkynum_words to topn in dtm_coherence, by @MeganStodel⚠️ 3.8.x will be the last Gensim version to support Py2.7. Starting with 4.0.0, Gensim will only support Py3.5 and above.
This is primarily a bugfix release to bring back Py2.7 compatibility to gensim 3.8.
lxml.etree.cElementTree (PR #2777, @tirkarthi)Remove
gensim.models.FastText.load_fasttext_format: use load_facebook_vectors to load embeddings only (faster, less CPU/memory usage, does not support training continuation) and load_facebook_model to load full model (slower, more CPU/memory intensive, supports training continuation)gensim.models.wrappers.fasttext (obsoleted by the new native gensim.models.fasttext implementation)gensim.examplesgensim.nosygensim.scripts.word2vec_standalonegensim.scripts.make_wiki_lemmagensim.scripts.make_wiki_onlinegensim.scripts.make_wiki_online_lemmagensim.scripts.make_wiki_online_nodebuggensim.scripts.make_wiki (all of these obsoleted by the new native gensim.scripts.segment_wiki implementation)Move
gensim.scripts.make_wikicorpus ➡ gensim.scripts.make_wiki.pygensim.summarization ➡ gensim.models.summarizationgensim.topic_coherence ➡ gensim.models._coherencegensim.utils ➡ gensim.utils.utils (old imports will continue to work)gensim.parsing.* ➡ gensim.utils.text_utilssmart_open version for compatibility with Py2.7Remove
gensim.models.FastText.load_fasttext_format: use load_facebook_vectors to load embeddings only (faster, less CPU/memory usage, does not support training continuation) and load_facebook_model to load full model (slower, more CPU/memory intensive, supports training continuation)gensim.models.wrappers.fasttext (obsoleted by the new native gensim.models.fasttext implementation)gensim.examplesgensim.nosygensim.scripts.word2vec_standalonegensim.scripts.make_wiki_lemmagensim.scripts.make_wiki_onlinegensim.scripts.make_wiki_online_lemmagensim.scripts.make_wiki_online_nodebuggensim.scripts.make_wiki (all of these obsoleted by the new native gensim.scripts.segment_wiki implementation)Move
gensim.scripts.make_wikicorpus ➡ gensim.scripts.make_wiki.pygensim.summarization ➡ gensim.models.summarizationgensim.topic_coherence ➡ gensim.models._coherencegensim.utils ➡ gensim.utils.utils (old imports will continue to work)gensim.parsing.* ➡ gensim.utils.text_utilsRemove
gensim.models.FastText.load_fasttext_format: use load_facebook_vectors to load embeddings only (faster, less CPU/memory usage, does not support training continuation) and load_facebook_model to load full model (slower, more CPU/memory intensive, supports training continuation)gensim.models.wrappers.fasttext (obsoleted by the new native gensim.models.fasttext implementation)gensim.examplesgensim.nosygensim.scripts.word2vec_standalonegensim.scripts.make_wiki_lemmagensim.scripts.make_wiki_onlinegensim.scripts.make_wiki_online_lemmagensim.scripts.make_wiki_online_nodebuggensim.scripts.make_wiki (all of these obsoleted by the new native gensim.scripts.segment_wiki implementation)Move
gensim.scripts.make_wikicorpus ➡ gensim.scripts.make_wiki.pygensim.summarization ➡ gensim.models.summarizationgensim.topic_coherence ➡ gensim.models._coherencegensim.utils ➡ gensim.utils.utils (old imports will continue to work)gensim.parsing.* ➡ gensim.utils.text_utilsgensim.downloader to run offline, by introducing a local file cache (mpenkov, #2545)gensim.downloader target directory configurable (mpenkov, #2456)nmslib indexer (masa3141, #2417)smart_open deprecation warning globally (itayB, #2530)topn=0 versus topn=None bug in most_similar, accept topn of any integer type (Witiko, #2497)CHANGELOG.md (mpenkov, #2482)gensim.similarities.termsim module (Witiko, #2485)Support section in README (piskvorky, #2542)Remove
gensim.models.FastText.load_fasttext_format: use load_facebook_vectors to load embeddings only (faster, less CPU/memory usage, does not support training continuation) and load_facebook_model to load full model (slower, more CPU/memory intensive, supports training continuation)gensim.models.wrappers.fasttext (obsoleted by the new native gensim.models.fasttext implementation)gensim.examplesgensim.nosygensim.scripts.word2vec_standalonegensim.scripts.make_wiki_lemmagensim.scripts.make_wiki_onlinegensim.scripts.make_wiki_online_lemmagensim.scripts.make_wiki_online_nodebuggensim.scripts.make_wiki (all of these obsoleted by the new native gensim.scripts.segment_wiki implementation)Move
gensim.scripts.make_wikicorpus ➡ gensim.scripts.make_wiki.pygensim.summarization ➡ gensim.models.summarizationgensim.topic_coherence ➡ gensim.models._coherencegensim.utils ➡ gensim.utils.utils (old imports will continue to work)gensim.parsing.* ➡ gensim.utils.text_utilsDoc2Vec.docvecs comment (gojomo, #2472)WordEmbeddingsKeyedVectors.most_similar (Witiko, #2461)matutils.unitvec always return float norm when requested (Witiko, #2419)Remove
gensim.models.FastText.load_fasttext_format: use load_facebook_vectors to load embeddings only (faster, less CPU/memory usage, does not support training continuation) and load_facebook_model to load full model (slower, more CPU/memory intensive, supports training continuation)gensim.models.wrappers.fasttext (obsoleted by the new native gensim.models.fasttext implementation)gensim.examplesgensim.nosygensim.scripts.word2vec_standalonegensim.scripts.make_wiki_lemmagensim.scripts.make_wiki_onlinegensim.scripts.make_wiki_online_lemmagensim.scripts.make_wiki_online_nodebuggensim.scripts.make_wiki (all of these obsoleted by the new native gensim.scripts.segment_wiki implementation)Move
gensim.scripts.make_wikicorpus ➡ gensim.scripts.make_wiki.pygensim.summarization ➡ gensim.models.summarizationgensim.topic_coherence ➡ gensim.models._coherencegensim.utils ➡ gensim.utils.utils (old imports will continue to work)gensim.parsing.* ➡ gensim.utils.text_utilsgensim.models.fasttext.load_facebook_model function: load full model (slower, more CPU/memory intensive, supports training continuation)gensim.models.fasttext.load_facebook_vectors function: load embeddings only (faster, less CPU/memory usage, does not support training continuation)To achieve consistency with the reference implementation from Facebook,
a FastText model will now always report any word, out-of-vocabulary or
not, as being in the model, and always return some vector for any word
looked-up. Specifically:
'any_word' in ft_model will always return True. Previously, itTrue only if the full word was in the vocabulary. (To test if awv.vocab'any_word' in ft_model.wv.vocab will return False if the fullft_model['any_word'] will always return a vector. Previously, itKeyError for OOV words when the model had no vectorsThe gensim.models.FastText.load_fasttext_format function (deprecated) now loads the entire model contained in the .bin file, including the shallow neural network that enables training continuation.
Loading this NN requires more CPU and RAM than previously required.
Since this function is deprecated, consider using one of its alternatives (see below).
Furthermore, you must now pass the full path to the file to load, including the file extension.
Previously, if you specified a model path that ends with anything other than .bin, the code automatically appended .bin to the path before loading the model.
This behavior was confusing, so we removed it.
Remove:
gensim.models.FastText.load_fasttext_format: use load_facebook_vectors to load embeddings only (faster, less CPU/memory usage, does not support training continuation) and load_facebook_model to load full model (slower, more CPU/memory intensive, supports training continuation)FastText.load_fasttext_model (@mpenkov, #2340)Doc2Vec.infer_vector (@tobycheese, #2347)LdaSeqModel (@horpto, #2360)process_result_queue from cycle in LdaMulticore (@horpto, #2358)LdaModel.do_mstep (@horpto, #2344)FastTextKeyedVectors using KeyedVectors (missing attribute compatible_hash) (@menshikh-iv, #2349)WordEmbeddingsKeyedVectors.most_similar (@Witiko, #2356)flake8==3.7.1 (@horpto, #2365)FastText documentation (@mpenkov, #2353)Any*Vec docstrings (@tobycheese, #2345)poincare documentation to indicate the relation format (@AMR-KELEG, #2357)Remove
gensim.models.wrappers.fasttext (obsoleted by the new native gensim.models.fasttext implementation)gensim.examplesgensim.nosygensim.scripts.word2vec_standalonegensim.scripts.make_wiki_lemmagensim.scripts.make_wiki_onlinegensim.scripts.make_wiki_online_lemmagensim.scripts.make_wiki_online_nodebuggensim.scripts.make_wiki (all of these obsoleted by the new native gensim.scripts.segment_wiki implementation)Move
gensim.scripts.make_wikicorpus ➡ gensim.scripts.make_wiki.pygensim.summarization ➡ gensim.models.summarizationgensim.topic_coherence ➡ gensim.models._coherencegensim.utils ➡ gensim.utils.utils (old imports will continue to work)gensim.parsing.* ➡ gensim.utils.text_utilsFast Online NMF (@anotherbugmaster, #2007)
Benchmark wiki-english-20171001
| Model | Perplexity | Coherence | L2 norm | Train time (minutes) |
|---|---|---|---|---|
| LDA | 4727.07 | -2.514 | 7.372 | 138 |
| NMF | 975.74 | -2.814 | 7.265 | 73 |
| NMF (with regularization) | 985.57 | -2.436 | 7.269 | 441 |
Simple to use (same interface as LdaModel)
Note truncated.
Your coding agent can read these notes before it upgrades. Set up the MCP server →