Open Source Speech Language Model
Spark-TTS Inference Code
The most powerful local music generation model
Multimodal embedding and reranking models built on Qwen3-VL
Multimodal-Driven Architecture for Customized Video Generation
Efficient few-shot learning with Sentence Transformers
MOSS-TTS-Nano is an open-source multilingual tiny speech generation
Foundation model for image generation
Controllable and fast Text-to-Speech for over 7000 languages
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A python library that makes AMR parsing, generation and visualization
Build Vision Agents quickly with any model or video provider
Capable of understanding text, audio, vision, video
Foundational video generation model with 13.6B parameters
A Multi-Modal World Model for Reconstructing, Generating, Simulation
FAIR Sequence Modeling Toolkit 2
The open-source data curation platform for LLMs
A Repo For Document AI
Scalable data pre processing and curation toolkit for LLMs
Pretrained model hub for Keras 3
Extract schema, statistics and entities from datasets
Stable Diffusion web UI
A very simple framework for state-of-the-art NLP
Cloud-native open source data warehouse for analytics and AI queries
Sample code and notebooks for Generative AI on Google Cloud