CLI proxy that reduces LLM token consumption
The production toolkit for LLMs. Observability, prompt management
Real-time multi-AI collaboration: Claude, Codex & Gemini
Completely free, private, UI based Tech Documentation MCP server
Uncertainty Quantification for Language Models, is a Python package
Performance-optimized AI inference on your GPUs
A Telegram bot for Large Language Models
State of the art LLM and coding model
95% token savings. 155x faster queries. 16 languages
A course of learning LLM inference serving on Apple Silicon
Run Mixtral-8x7B models in Colab or consumer desktops
Calculate token/s & GPU memory requirement for any LLM
Flagship MoE model for long-context agents and complex coding
Omnimodal AI model for agents, coding, and long-context tasks
Unified multimodal Gemma model for local coding and reasoning
Flagship Poolside model for agentic coding and software engineering