Motion-controllable Video Generation via Latent Trajectory Guidance
"Big Model" trains a visual multimodal VLM with 26M parameters
Automatically translates the text of a video based on a subtitle file
Nerlnet is a framework for research and development
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
The Library for LLM-based multi-agent applications
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
Framework to easily create LLM powered bots over any dataset
A unified, comprehensive and efficient recommendation library
A python library for easy manipulation and forecasting of time series
Personal notes from Wu Enda's machine learning course
Open-source MCP server that gives your coding agent
Collection of reference environments, offline reinforcement learning
Build MLOps Pipelines in Minutes
AI logo animation skill: turn raster logos into smooth SVG animation
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Official Repo For "Sa2VA: Marrying SAM2 with LLaVA
LLM-based agent for general purpose software engineering tasks
Multi-modal large language model designed for audio understanding
Large Multimodal Models for Video Understanding and Editing
The official PyTorch implementation of Google's Gemma models
Capable of understanding text, audio, vision, video
Advanced evolutionary computation library built on top of PyTorch
An Easy-to-use, Scalable and High-performance RLHF Framework