AI-powered semantic indexing: automating the creation of book indexes
Chat & pretrained large vision language model
A Pioneering Open-Source Alternative to GPT-4o
ktrain is a Python library that makes deep learning AI more accessible
Chat & pretrained large audio language model proposed by Alibaba Cloud
Towards Real-World Vision-Language Understanding
Implementation of Make-A-Video, new SOTA text to video generator
Turn your website into a GIF
Minimal, clean code for the Byte Pair Encoding (BPE) algorithm
Virtual AI anchor that combines state-of-the-art technology
High-quality multi-lingual text-to-speech library by MyShell.ai
A library for transfer learning by reusing parts of TensorFlow models
Multi-Voice and Prompt-Controlled TTS Engine
Official code for Style Aligned Image Generation via Shared Attention
Text-to-Image generation. The repo for NeurIPS 2021 paper
Embed images and sentences into fixed-length vectors
Generate 3D objects conditioned on text or images
Convert an image to text to spot intelligible words.
Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis
CLIP + FFT/DWT/RGB = text to image/video
Let us control diffusion models
Run the Stable Diffusion releases in a Docker container
The first Chinese LLaMA2 model in the open source community
Alfred workflow using ChatGPT, DALL·E 2 and other models for chatting
Multimodal AI Story Teller, built with Stable Diffusion, GPT, etc.