tiktoken is a fast BPE tokeniser for use with OpenAI's models
This repo contains the code for 1D tokenizer and generator
Tokenizer-Free TTS for Multilingual Speech Generation
Long-form streaming TTS system for multi-speaker dialogue generation
The best ChatGPT that $100 can buy
Audio Language Models are Few-Shot Learners
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
Python library and CLI tool to interface with Google Translate
LLM-based Reinforcement Learning audio edit model
Unified Multimodal Understanding and Generation Models
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
Large Language Model Principles and Practice Tutorial from Scratch
The official PyTorch implementation of Google's Gemma models
Qwen3-Coder is the code version of Qwen3
MOSS‑TTS Family open‑source speech and sound generation model
Minimal, clean code for the Byte Pair Encoding (BPE) algorithm
Audiocraft is a library for audio processing and generation
Data loaders and abstractions for text and NLP
Code for the paper Language Models are Unsupervised Multitask Learners
Inference code for Llama models
Open-source pre-training implementation of Google's LaMDA in PyTorch
An implementation of model parallel GPT-2 and GPT-3-style models
GPT2 for Multiple Languages, including pretrained models
A large annotated semantic parsing corpus for developing NL interfaces