Low-latency REST API for serving text-embeddings
Trainable, memory-efficient, and GPU-friendly PyTorch reproduction
LLM training code for MosaicML foundation models
PyTorch extensions for fast R&D prototyping and Kaggle farming
Libraries for applying sparsification recipes to neural networks
Gaussian processes in TensorFlow
Optimizing inference proxy for LLMs
Neural Network Compression Framework for enhanced OpenVINO
Efficient few-shot learning with Sentence Transformers
Pytorch domain library for recommendation systems
Data manipulation and transformation for audio signal processing
A Pythonic framework to simplify AI service building
Adversarial Robustness Toolbox (ART) - Python Library for ML security
Phi-3.5 for Mac: Locally-run Vision and Language Models
A Unified Library for Parameter-Efficient Learning
Scripts for fine-tuning Meta Llama3 with composable FSDP & PEFT method
DoWhy is a Python library for causal inference
State-of-the-art Parameter-Efficient Fine-Tuning
Simplifies the local serving of AI models from any source
Multilingual Automatic Speech Recognition with word-level timestamps
A high-performance ML model serving framework, offers dynamic batching
Unified Model Serving Framework
Probabilistic reasoning and statistical analysis in TensorFlow
Integrate, train and manage any AI models and APIs with your database
Replace OpenAI GPT with another LLM in your app