StreamSpeech is a seamless model for offline speech recognition
Toolkit for conversational AI
Generate blog articles from video or audio
A wiki system with complex functionality for simple integration
Audio Language Models are Few-Shot Learners
Pre & Post-training & Dataset & Evaluation & Depoly & RAG
Repo of Qwen2-Audio chat & pretrained large audio language model
Hypernetworks that adapt LLMs for specific benchmark tasks
Tutorial tailored for Chinese babies on rapid fine-tuning
Implement a concise and clear Deep Search Agent from 0
RGBD video generation model conditioned on camera input
code for Mesh R-CNN, ICCV 2019
Python project template generator with batteries included
Capable of understanding text, audio, vision, video
Official Repo For "Sa2VA: Marrying SAM2 with LLaVA
Building a Secure and Interoperable Future for AI-Driven Payments
A neural network that transforms a design mock-up into static websites
A lightweight audio-to-MIDI converter with pitch bend detection
Marrying Grounding DINO with Segment Anything & Stable Diffusion
Open source demo platform where you can easily showcase your AI models
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation
Powerful open source image generation model
polizei, afd, linke, linke gewalt, polizeigewalt, demo gegen rechts
My MSX programs and some additional .cas tools
BTRFS based NAS and private cloud storage solution