State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Educational framework exploring multi-agent orchestration
Repo of Qwen2-Audio chat & pretrained large audio language model
Tooling for the Common Objects In 3D dataset
Implementation of Phenaki Video, which uses Mask GIT
Fast, flexible and easy to use probabilistic modelling in Python
MARS5 speech model (TTS) from CAMB.AI
Usable Implementation of "Bootstrap Your Own Latent" self-supervised
Toloka-Kit is a Python library for working with Toloka API
High-Resolution Image Synthesis with Latent Diffusion Models
Unlimited, private and free Speech-To-Text program
AI-powered tool to quickly remove watermarks from images flawlessly
AI-powered quiz solver for Windows. Free to use, easy to set up.
Open-source AI video pipeline, fully automated with MCP
Two Integrated Text To Speech Engines uses MMS & Silero
Edge TTS Desktop turns text into speech through edge-tts.
Video+code lecture on building nanoGPT from scratch
Inference Llama 2 in one file of pure C
AI Suite for upscaling, interpolating & restoring images/videos
The Brutalist Market Analyzer
It's possible for machines to become self-aware.
AI-powered semantic indexing: automating the creation of book indexes
The python App/Skrypt automaticly add important events into calendar.
Towards Human-Level Text-to-Speech through Style Diffusion
A desktop weather app powered by AI