Gensim Data is the official storage repository for downloadable text corpora and pretrained NLP models used through Gensim's downloader API. It provides stable distribution for research resources that might otherwise disappear or change at their original locations. Large datasets and model files are stored as immutable GitHub release attachments. Users normally access the catalog through Python or the Gensim command-line downloader rather than cloning the repository itself. Downloads are cached in a local gensim-data directory for later reuse. The collection includes corpora such as text8 and 20 Newsgroups alongside pretrained word-vector resources such as GloVe. Every dataset retains its own licensing terms, so users must review individual usage conditions.

Features

  • Pretrained NLP model distribution
  • Downloadable text corpus collection
  • Gensim Python downloader integration
  • Command-line dataset downloading
  • Immutable versioned release storage
  • Automatic local dataset caching

Project Samples

Project Activity

See All Activity >

License

GNU Library or Lesser General Public License version 3.0 (LGPLv3)

Follow Gensim-data

Gensim-data Web Site

Other Useful Business Software
$300 Free Credits to Build on Google Cloud Icon
$300 Free Credits to Build on Google Cloud

New customers can spin up VMs, build with AI, and query data at no cost.

Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
Learn More
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Gensim-data!

Additional Project Details

Programming Language

Python

Related Categories

Python Natural Language Processing (NLP) Tool

Registered

3 days ago