The best free open source website change detection and restock service
An open source implementation of CLIP
RAG-Anything: All-in-One RAG Framework
Foundation model for image generation
Marrying Grounding DINO with Segment Anything & Stable Diffusion
PersonaPlex code
Style-Bert-VITS2: Bert-VITS2 with more controllable voice styles
Crowdsourcing platform for full text transcription and tagging
Towards Human-Sounding Speech
A full spaCy pipeline and models for scientific/biomedical documents
Temporal-Consistent Diffusion Model for Real-World Video
Framework for building AI-powered interactive digital humans and agent
End-to-end speech processing toolkit
Code and models for ICML 2024 paper, NExT-GPT
CineCLI is a cross-platform command-line movie browser
Extract audio and video content and organize it into a Markdown note
StreamSpeech is a seamless model for offline speech recognition
NeuTTS model built from small LLM backbones
On-device TTS model by Neuphonic
Open source personal AI Assistant for Linux, Windows and Mac
Large-language-model & vision-language-model based on Linear Attention
Foundational video generation model with 13.6B parameters
Fast multimodal LLM for real-time voice interaction and AI apps
StarVector is a foundation model for SVG generation
Chat with it via text and voice