Less Code, Lower Barrier, Faster Deployment
NLTK Source
Curated list of datasets and tools for post-training
Information hub for our project training the largest possible LLMs
Language modeling in a sentence representation space
Reading book source
The Classical Language Toolkit
A full spaCy pipeline and models for scientific/biomedical documents
Video-based AI memory library. Store millions of text chunks in MP4
Topic Modelling for Humans
Indexing and query tools for very large text corpora
A New Axis of Sparsity for Large Language Models
The simplest, fastest repository for training/finetuning models
Your Fully-Automated Personal AI Assistant
Web application for Markdown note taking
Code release for Cut and Learn for Unsupervised Object Detection
Chinese XLNet pre-trained model
Omnilingual ASR Open-Source Multilingual SpeechRecognition
Style-Bert-VITS2: Bert-VITS2 with more controllable voice styles
SOTA discrete acoustic codec models with 40/75 tokens per second
A fast TTS architecture with conditional flow matching
Traditional Mandarin LLMs for Taiwan
A subtitle generator for Japanese Adult Videos.
Aligns tokens in two versions of a text with differing tokenization.
Chinese Llama-3 LLMs) developed from Meta Llama 3