Build cross-modal and multimodal applications on the cloud
Zep: A long-term memory store for LLM / Chatbot applications
A single Gradio + React WebUI with extensions for ACE-Step
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Official Repo For "Sa2VA: Marrying SAM2 with LLaVA
LLM-based agent for general purpose software engineering tasks
Large Multimodal Models for Video Understanding and Editing
Speech-AI-Forge is a project developed around TTS generation model
Jupyter notebooks that demonstrate how to build models using SageMaker
Recognition and resolution of numbers, units, date/time, etc.
Miso TTS is an 8 billion, highly emotive text-to-speech model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
An advanced paper search agent powered by large language models
GUI Exploration Lab. One of the best GUI agent solutions
Open-weight, large-scale hybrid-attention reasoning model
Large-language-model & vision-language-model based on Linear Attention
Extensible, parallel implementations of t-SNE
Capable of understanding text, audio, vision, video
Open source framework for deep learning satellite and aerial imagery
Private AI platform for agents, enterprise search and RAG pipelines
Debug, evaluate, and monitor your LLMapps, RAG systems, and agentic AI
mlpack: a scalable C++ machine learning library
Simple and distributed Machine Learning
E2M converts various file types (doc, docx, epub, html, htm, url
Fundamentals of Machine Learning and Deep Learning