High accuracy RAG for answering questions from scientific documents
Netease Youdao's open-source embedding and reranker models
High-performance inference server for text embeddings models API layer
AI-Powered Wiki Generator for GitHub/Gitlab/Bitbucket Repositories
A New Axis of Sparsity for Large Language Models
An Efficient Web-enhanced Question Answering System
Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
The beginning of scalable pixel-native search
Document content and metadata extraction microservice
"VideoRAG: Chat with Your Videos
Evaluation and Tracking for LLM Experiments
Fast State-of-the-Art Static Embeddings
Minimal Python framework for scalable AI inference servers fast
Open source RAG framework for building scalable modular AI apps
Ready-to-run cloud templates for RAG
Document Index for Vectorless, Reasoning-based RAG
Making RAG Simpler with Small and Open-Sourced Language Models
SimpleMem: Efficient Lifelong Memory for LLM Agents
Build production-ready AI agents in both Python and Typescript
Low-latency AI inference engine optimized for mobile devices
Ship AI Agents to Google Cloud in minutes, not months
AI-powered document analysis and tagging for Paperless-ngx
A collection of scientific methods, processes, algorithms
Learning to Reason with Search for LLMs via Reinforcement Learning
In-depth tutorials on LLMs, RAGs and real-world AI agent applications