Low-latency REST API for serving text-embeddings
GitLab automatic code review tool based on large models
A high-throughput and memory-efficient inference and serving engine
Linkedin Automation Tool
Compress tool outputs, logs, files, and RAG chunks
High-performance inference framework for large language models
LLM inference server with continuous batching & SSD caching
LightLLM is a Python-based LLM (Large Language Model) inference
TokenSpeed is a speed-of-light LLM inference engine
Parallax is a distributed model serving framework
Easy token price estimates for 400+ LLMs. TokenOps
The first AI agent that builds permissionless integrations
Weaving the Digital Agent Galaxy
A high-performance ML model serving framework, offers dynamic batching
LangChain powered shell command generator and runner CLI
User toolkit for analyzing and interfacing with Large Language Models
AIlice is a fully autonomous, general-purpose AI agent
Serving multiple LoRA finetuned LLM as one
Serving LangChain LLM apps automagically with FastApi
AI-powered CLI git wrapper, boilerplate code generator, chat history