High-speed Large Language Model Serving for Local Deployment
157 models, 30 providers, one command to find what runs on hardware
Sparsity-aware deep learning inference runtime for CPUs
Fast State-of-the-Art Static Embeddings
A high-quality rapid TTS voice cloning model
A high-performance ML model serving framework, offers dynamic batching
Elegant and Performant Deep Learning
Generate audiobooks from e-books
Fast ML inference & training for ONNX models in Rust
Lightning-fast, on-device TTS, running natively via ONNX
Real-time NVIDIA GPU dashboard
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
ChatGLM2-6B: An Open Bilingual Chat LLM
A Powerful Desktop Full-Text Search Engine, Just Like Local Google.
Easy-to-use deep learning framework with 3 key features
fast C++ library for GPU linear algebra & scientific computing
Calculate token/s & GPU memory requirement for any LLM
Generative Adversarial Networks for Efficient and High Fidelity Speech
Deep learning library featuring a higher-level API for TensorFlow
Caffe2 is a lightweight, modular, and scalable deep learning framework
Tiny pre-trained IBM model for multivariate time series forecasting