ChatRWKV is like ChatGPT but powered by RWKV
Unified Multimodal Understanding and Generation Models
Open-source platform for building enterprise-grade agents
Code to accompany "A Method for Animating Children's Drawings"
Multimodal Diffusion with Representation Alignment
Apache-2.0 open-source image generation and editing model family
StreamSpeech is a seamless model for offline speech recognition
Library for OCR-related tasks powered by Deep Learning
Generate blog articles from video or audio
Audio Language Models are Few-Shot Learners
Pre & Post-training & Dataset & Evaluation & Depoly & RAG
Repo of Qwen2-Audio chat & pretrained large audio language model
Toolkit for conversational AI
Hypernetworks that adapt LLMs for specific benchmark tasks
Tutorial tailored for Chinese babies on rapid fine-tuning
Implement a concise and clear Deep Search Agent from 0
RGBD video generation model conditioned on camera input
code for Mesh R-CNN, ICCV 2019
Capable of understanding text, audio, vision, video
Official Repo For "Sa2VA: Marrying SAM2 with LLaVA
Building a Secure and Interoperable Future for AI-Driven Payments
A neural network that transforms a design mock-up into static websites
A lightweight audio-to-MIDI converter with pitch bend detection
Marrying Grounding DINO with Segment Anything & Stable Diffusion
Open source demo platform where you can easily showcase your AI models