Redundancy-aware KV Cache Compression for Reasoning Models
AI Agent Source Code Deep Research Report
Agent framework and applications built upon Qwen>=3.0
Running large language models on a single GPU
Unified web UI for training and running open models locally
Persistent context and multi-instance coordination
A Python library for audio
Gradient boosting framework based on decision tree algorithms
Developer friendly Natural Language Processing
ReFT: Representation Finetuning for Language Models
The repository provides code for running inference with SAM 2
Self-evolving autonomous agent framework
High-speed Large Language Model Serving for Local Deployment
One brain, many harnesses. Portable .agent/ folder
Drag & drop UI to build your customized LLM flow
LLM inference in C/C++
A Web UI for easy subtitle using whisper model
A step-by-step guide to build your own AI agent
Benchmarking synthetic data generation methods
Memory-efficient and performant finetuning of Mistral's models
LLM training in simple, raw C/CUDA
Turns Data and AI algorithms into production-ready web applications
14-stage Fusion Pipeline for LLM token compression
Designed for training LLM/VLM agents via RL
Speech recognition module for Python