Faster Whisper transcription with CTranslate2
State-of-the-art Parameter-Efficient Fine-Tuning
MobileLLM Optimizing Sub-billion Parameter Language Models
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
MTEB: Massive Text Embedding Benchmark
SWE-agent takes a GitHub issue and tries to automatically fix it
Agent Skill for generating 2D sprite sheets and map, transparent PNG
A SOTA open-source image editing model
Han Language Processing
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Stable Diffusion web UI
1 min voice data can also be used to train a good TTS model
Multilingual sentence & image embeddings with BERT
State-of-the-art (SoTA) text-to-video pre-trained model
Agent S: an open agentic framework that uses computers like a human
Best Practices on Recommendation Systems
A game engine powered by python and panda3d
Offline inference engine for art, real-time voice conversations
A PyTorch-based Speech Toolkit
OCR expert VLM powered by Hunyuan's native multimodal architecture
Management of Yandex Station and other smart home devices
Synthetic data generators for tabular and time-series data
Image inpainting tool powered by SOTA AI Model
Image inpainting tool powered by SOTA AI Model
Industrial-level controllable zero-shot text-to-speech system