The Triton Inference Server provides an optimized cloud
Easiest and laziest way for building multi-agent LLMs applications
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
Ready-to-use OCR with 80+ supported languages
A Pythonic framework to simplify AI service building
State-of-the-art diffusion models for image and audio generation
Library for serving Transformers models on Amazon SageMaker
LLM training code for MosaicML foundation models
Easy-to-use Speech Toolkit including Self-Supervised Learning model
Library for OCR-related tasks powered by Deep Learning
Replace OpenAI GPT with another LLM in your app
Bring the notion of Model-as-a-Service to life
A unified framework for scalable computing
Large Language Model Text Generation Inference
A library for accelerating Transformer models on NVIDIA GPUs
Powering Amazon custom machine learning chips
Open-source tool designed to enhance the efficiency of workloads
A high-performance ML model serving framework, offers dynamic batching
Unified Model Serving Framework
Low-latency REST API for serving text-embeddings
Standardized Serverless ML Inference Platform on Kubernetes
Trainable, memory-efficient, and GPU-friendly PyTorch reproduction
Lightweight Python library for adding real-time multi-object tracking
GUI shell for running local LLM on desktop
Openai style api for open large language models