Multi-modal large language model designed for audio understanding
A neural network that transforms a design mock-up into static websites
The standard data-centric AI package for data quality and ML
code for Mesh R-CNN, ICCV 2019
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Build cross-modal and multimodal applications on the cloud
Chinese and English multimodal conversational language model
A Codex Skill for generating high-density, editable PowerPoints
LISA: Reasoning Segmentation via Large Language Model
Skywork-R1V is an advanced multimodal AI model series
Refer and Ground Anything Anywhere at Any Granularity
Language modeling in a sentence representation space
Stable Diffusion WebUI optimized for AMD GPUs with editing tools
Implementation of Phenaki Video, which uses Mask GIT
Plug-n-play module turning text-to-image models into animation
Run GGUF models easily with a UI or API. One File. Zero Install.
Mice speech to text with MX Cinnamon OS ISO
Open source framework for deep learning satellite and aerial imagery
A Python application to add watermarks (text or image) to PDF files
Open source demo platform where you can easily showcase your AI models
Autoregressive Model Beats Diffusion
AI-powered tool to quickly remove watermarks from images flawlessly
User toolkit for analyzing and interfacing with Large Language Models
Overcoming Data Limitations for High-Quality Video Diffusion Models
dashAI: an interactive platform for training, evaluating and deploying