Unified KV Cache Compression Methods for Auto-Regressive Models
Real-time multi-AI collaboration: Claude, Codex & Gemini
Cache-Augmented Generation: A Simple, Efficient Alternative to RAG
High-performance inference framework for large language models
LightLLM is a Python-based LLM (Large Language Model) inference
StarVector is a foundation model for SVG generation
Accessible large language models via k-bit quantization for PyTorch
A lightweight vLLM implementation built from scratch
slime is an LLM post-training framework for RL Scaling
Knowledge Graph Generation from Any Text
GPT4V-level open-source multi-modal model based on Llama3-8B
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
LLM training code for MosaicML foundation models
Leveraging BERT and c-TF-IDF to create easily interpretable topics
Data Lake for Deep Learning. Build, manage, and query datasets
Integrating LLMs into structured NLP pipelines
Capable of understanding text, audio, vision, video
Text-space optimizer that trains reusable natural-language skills
A straightforward method for training your LLM
A Telegram bot for Large Language Models
A Survey of Large Language Models
Language-model investigation agent with a terminal UI
950 line, minimal, extensible LLM inference engine built from scratch
Seamlessly integrate LLMs into scikit-learn
A security scanner for custom LLM applications