DepGraph: Towards Any Structural Pruning
High-performance inference framework for large language models
Code and models for ICML 2024 paper, NExT-GPT
Run PyTorch LLMs locally on servers, desktop and mobile
High-performance Inference and Deployment Toolkit for LLMs and VLMs
A system for agentic LLM-powered data processing and ETL
Power CLI and Workflow manager for LLMs (core package)
CV, NLP, LLM project applications, and advanced engineering deployment
A Next-Generation Training Engine Built for Ultra-Large MoE Models
Structured data extraction and instruction calling with ML, LLM
Robust recipes to align language models with human and AI preferences
Open Source Deep Research Alternative to Reason and Search
Accessible large language models via k-bit quantization for PyTorch
Accelerate local LLM inference and finetuning
slime is an LLM post-training framework for RL Scaling
In-depth tutorials on LLMs, RAGs and real-world AI agent applications
Large Audio Language Model built for natural interactions
95% token savings. 155x faster queries. 16 languages
Advanced techniques for RAG systems
Refer and Ground Anything Anywhere at Any Granularity
Set of tools to assess and improve LLM security
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
Research code artifacts for Code World Model (CWM)
Low-latency REST API for serving text-embeddings