tiktoken is a fast BPE tokeniser for use with OpenAI's models
This repo contains the code for 1D tokenizer and generator
Tokenizer-Free TTS for Multilingual Speech Generation
Long-form streaming TTS system for multi-speaker dialogue generation
Audio Language Models are Few-Shot Learners
TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
The best ChatGPT that $100 can buy
Python library and CLI tool to interface with Google Translate
LLM-based Reinforcement Learning audio edit model
A Foundation Model for the Language of Financial Markets
Large Language Model Principles and Practice Tutorial from Scratch
Unified Multimodal Understanding and Generation Models
The official PyTorch implementation of Google's Gemma models
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
Qwen3-Coder is the code version of Qwen3
MOSS‑TTS Family open‑source speech and sound generation model
Minimal, clean code for the Byte Pair Encoding (BPE) algorithm
Audiocraft is a library for audio processing and generation
Data loaders and abstractions for text and NLP
Code for the paper Language Models are Unsupervised Multitask Learners
Inference code for Llama models
Open-source pre-training implementation of Google's LaMDA in PyTorch
State of the art faster Transformer with Tensorflow 2.0
An implementation of model parallel GPT-2 and GPT-3-style models
GPT2 for Multiple Languages, including pretrained models