"Big Model" trains a visual multimodal VLM with 26M parameters
Controllable & emotion-expressive zero-shot TTS
Generate blog articles from video or audio
Implementation of "MobileCLIP" CVPR 2024
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Video understanding codebase from FAIR for reproducing video models
Renderer for the harmony response format to be used with gpt-oss
CLIP, Predict the most relevant text snippet given an image
Ling is a MoE LLM provided and open-sourced by InclusionAI
Simple, unified interface to multiple Generative AI providers
A Python library for audio
BertViz: Visualize Attention in NLP Models (BERT, GPT2, BART, etc.)
Intelligent companion for seamless AI engineering and research
Universal LLM Deployment Engine with ML Compilation
All-in-one WebUI for AI generative image and video creation
ID-based RAG FastAPI: Integration with Langchain and PostgreSQL
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
LLM-based Reinforcement Learning audio edit model
A mcp server for vikingdb store and search
Repo of Qwen2-Audio chat & pretrained large audio language model
MOSS‑TTS Family open‑source speech and sound generation model
Bidirectional token-classification model for identifiable info
Project Lyra: Open Generative 3D World Models