State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Automatic Speech Recognition with Word-level Timestamps
Python Audio Analysis Library: Feature Extraction, Classification
Multilingual speech recognition and audio understanding model
Automagically synchronize subtitles with video
Python library for audio and music analysis
Label Studio is a multi-type data labeling and annotation tool
Swing Music is a beautiful, self-hosted music player
Automated Music Discovery and Collection Manager
Robust Speech Recognition via Large-Scale Weak Supervision
EPUB to audiobook converter, optimized for Audiobookshelf
Have a natural, spoken conversation with AI
A Web UI for easy subtitle using whisper model
A python tool that uses GPT-4, FFmpeg, and OpenCV
A nearly-live implementation of OpenAI's Whisper
Build Vision Agents quickly with any model or video provider
Get your documents ready for gen AI
Adversarial Robustness Toolbox (ART) - Python Library for ML security
Marrying Grounding DINO with Segment Anything & Stable Diffusion
A lightweight audio-to-MIDI converter with pitch bend detection
Just pull anything
JamPilot — jam along with anything, in real time
Record, edit and replay supervised Windows desktop macros
Audio Transcription software for Linux (Vlc) with a foot pedal
Audio Transcription software for Linux (Gstreamer) with a foot pedal