Uncover insights, surface problems, monitor, and fine tune your LLM
The Triton Inference Server provides an optimized cloud
Lightweight Python library for adding real-time multi-object tracking
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V)
Multilingual Automatic Speech Recognition with word-level timestamps
Superduper: Integrate AI models and machine learning workflows
A high-performance ML model serving framework, offers dynamic batching
Phi-3.5 for Mac: Locally-run Vision and Language Models
Libraries for applying sparsification recipes to neural networks
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
A set of Docker images for training and serving models in TensorFlow
Open-Source AI Camera. Empower any camera/CCTV
State-of-the-art diffusion models for image and audio generation
Bring the notion of Model-as-a-Service to life
State-of-the-art Parameter-Efficient Fine-Tuning
Optimizing inference proxy for LLMs
LMDeploy is a toolkit for compressing, deploying, and serving LLMs
Neural Network Compression Framework for enhanced OpenVINO
Sparsity-aware deep learning inference runtime for CPUs
Large Language Model Text Generation Inference
Standardized Serverless ML Inference Platform on Kubernetes
Trainable models and NN optimization tools
Efficient few-shot learning with Sentence Transformers
Fast inference engine for Transformer models