Generate Any 3D Scene in Seconds
FAIR Sequence Modeling Toolkit 2
1B text generation model based on the HRM architecture
Backlog-row-first content production system for teams
Machine Learning Engineering Open Book
Your Fully-Automated Personal AI Assistant
Tiny vision language model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Multi-modal large language model designed for audio understanding
State-of-the-art (SoTA) text-to-video pre-trained model
Open-source framework for intelligent speech interaction
Large Multimodal Models for Video Understanding and Editing
Chat with your documents using local AI
An open sourced end-to-end VLM-based GUI Agent
kaldi-asr/kaldi is the official location of the Kaldi project
Autonomous LLM agent for end-to-end data science workflows
Context-aware desktop AI assistant that understands screen content
General-purpose image editing model that delivers high-fidelity
Self-learning data agent that grounds its answers in layers of content
A long-running autonomous coding agent powered by the Claude Agent
This repository contains the official implementation of FastVLM
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Semantic cache for LLMs. Fully integrated with LangChain
An advanced paper search agent powered by large language models
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training