Wordvectors is a collection of pretrained word embeddings for more than 30 languages. It was created to make multilingual vector representations easier to obtain, especially for languages with fewer readily available resources than English. The repository provides models trained using both Word2Vec and fastText. Training corpora are constructed from Wikipedia database dumps using language-specific preprocessing when necessary. Scripts are included for corpus creation and for training new embeddings with either supported algorithm. Available languages include Chinese, Japanese, Korean, Spanish, French, German, Hindi, Russian, Vietnamese, Thai, and many others, with model metadata covering vector, corpus, and vocabulary sizes.

Features

  • Pretrained multilingual word embeddings
  • More than 30 supported languages
  • Word2Vec model support
  • fastText model support
  • Wikipedia-based corpus generation
  • Reusable vector training scripts

Project Samples

Project Activity

See All Activity >

Categories

Libraries

License

MIT License

Follow wordvectors

wordvectors Web Site

Other Useful Business Software
Veeam Data Platform v13.1 Icon
Veeam Data Platform v13.1

Move workloads across hypervisors and clouds with no vendor lock-in. Try VDP free today.

Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
Try Now
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of wordvectors!

Additional Project Details

Programming Language

Python

Related Categories

Python Libraries

Registered

2026-08-28