Open-source framework for intelligent speech interaction
Large Multimodal Models for Video Understanding and Editing
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
On-device Speech-to-Intent engine powered by deep learning
Benchmarking synthetic data generation methods
Making Enterprise Data Intelligent and Responsive for AI
AIMET is a library that provides advanced quantization and compression
Powering Amazon custom machine learning chips
An advanced paper search agent powered by large language models
Open-weight, large-scale hybrid-attention reasoning model
Large-language-model & vision-language-model based on Linear Attention
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
LLM-based Reinforcement Learning audio edit model
Capable of understanding text, audio, vision, video
MuA multi-agent reinforcement learning environment
An industrial grade federated learning framework
LLM
Private AI platform for agents, enterprise search and RAG pipelines
Implementation of Phenaki Video, which uses Mask GIT
Foundational model for human-like, expressive TTS
Play couplet with seq2seq model
The Python code to reproduce illustrations from Machine Learning Book
User toolkit for analyzing and interfacing with Large Language Models
CodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)
Memory-efficient and performant finetuning of Mistral's models