tiktoken is a fast BPE tokeniser for use with OpenAI's models
This repo contains the code for 1D tokenizer and generator
Tokenizer-Free TTS for Multilingual Speech Generation
Long-form streaming TTS system for multi-speaker dialogue generation
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
The best ChatGPT that $100 can buy
Python library and CLI tool to interface with Google Translate
Large Language Model Principles and Practice Tutorial from Scratch
Unified Multimodal Understanding and Generation Models
The official PyTorch implementation of Google's Gemma models
MOSS‑TTS Family open‑source speech and sound generation model
Minimal, clean code for the Byte Pair Encoding (BPE) algorithm
Audiocraft is a library for audio processing and generation
Code for the paper Language Models are Unsupervised Multitask Learners
Custom BLEURT model for evaluating text similarity using PyTorch
Robust BERT-based model for English with improved MLM training