OpenRouter
OpenRouter is an AI model routing platform that gives developers access to hundreds of models through a single unified API. It connects users with models from providers such as OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek, Qwen, xAI, and many others. The platform supports text, image, video, and audio generation while allowing developers to use one API key and a consistent interface across providers. OpenRouter can route requests based on price, performance, and availability, with fallback options that help maintain service when a provider experiences downtime. It also offers configurable data policies so organizations can control which providers receive prompts and how requests are handled. Developers can purchase credits, choose from more than 500 active models across over 80 providers, and integrate OpenRouter using an OpenAI-compatible API.
Learn more
Photon
Photon is Moondream’s official high-performance inference engine, designed to run vision-language models efficiently across cloud, desktop, and edge environments while delivering real-time performance for production AI systems. It is built as a custom inference layer tightly integrated with the Moondream model architecture, using optimized scheduling, native image processing, and purpose-built CUDA kernels to maximize speed and efficiency. This co-designed approach allows Photon to significantly reduce latency compared to traditional VLM setups, enabling responsive interactions on edge devices and real-time throughput on server-grade hardware. It supports deployment across a wide range of NVIDIA GPUs, from embedded systems like Jetson devices to high-end multi-GPU servers, making it adaptable for diverse operational needs. It includes production-ready features such as automatic batching, prefix caching, and memory-efficient attention mechanisms.
Learn more
BaseRT
BaseRT is a high-performance LLM inference runtime for Apple Silicon that lets developers pull models from Hugging Face, chat with them locally, or serve an OpenAI-compatible API from one CLI. Accelerated by hand-written Metal kernels, it is designed to deliver fast prefill and decode performance on M-series Macs, with published benchmarks showing up to 6.4× faster prefill than llama.cpp, 3.9× faster than MLX, and up to 1.33× faster decode. The basert CLI handles model downloading, conversion, interactive chat, serving, completion, benchmarking, inspection, and bundle signing. Its server supports chat, completions, embeddings, transcription, tool calls, continuous batching, paged KV caching, and prefix caching, while supported models can process text, vision, and audio. BaseRT uses its own .base model format with Q2–Q8 affine quantization, optional AWQ calibration, and signed bundles, and can convert GGUF, Hugging Face, and MLX checkpoints.
Learn more
Macyou
Macyou rents dedicated Apple Silicon Macs for AI workloads. Users configure a Mac (M4 Mac mini to Mac Studio M3 Ultra with 256 GB unified memory), pick a pre-configured stack — local LLMs via Ollama (Llama, Qwen, Mistral, DeepSeek), agent frameworks (CrewAI, LangGraph), or ML dev environments (MLX, Jupyter) — and get a running deployment in about 5 minutes. Every deployment exposes an OpenAI-compatible API, so existing OpenAI SDK code works by changing base_url; access also includes SSH with root and a browser-based remote desktop. Each customer gets a dedicated physical machine with full-disk encryption and a disk wipe between tenants, hosted in a GDPR-friendly jurisdiction. Pricing is a fixed monthly fee per machine with no per-token charges; Thunderbolt 5 clustering pools unified memory across nodes for larger models. Published, measured inference benchmarks (raw JSON, CC BY 4.0) show real tokens per second per chip.
Learn more