Real-World Centric Foundation GUI Agents
SOTA discrete acoustic codec models with 40/75 tokens per second
Multilingual sentence & image embeddings with BERT
Question and Answer based on Anything
Phi-3.5 for Mac: Locally-run Vision and Language Models
InvokeAI is a leading creative engine for Stable Diffusion models
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
Fast-stable-diffusion + DreamBooth
This repo contains the code for 1D tokenizer and generator
Scalable generative AI framework built for researchers and developers
Infinite Worlds with Versatile Interactions
AI-Powered Personalized Learning Assistant
Easy-to-use and high-performance NLP and LLM framework
Fast and customizable framework for automatic ML model creation
⚡ Building applications with LLMs through composability ⚡
State-of-the-art diffusion models for image and audio generation
Agent Skill for generating 2D sprite sheets and map, transparent PNG
Open source libraries and APIs to build custom preprocessing pipelines
Multi-modal large language model designed for audio understanding
Large Multimodal Models for Video Understanding and Editing
Dealing with all unstructured data, such as reverse image search
Algorithms for outlier, adversarial and drift detection
A python tool that uses GPT-4, FFmpeg, and OpenCV
Audio foundation model excelling in audio understanding
Open-weight, large-scale hybrid-attention reasoning model