[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences
Natural Gradient Boosting for Probabilistic Prediction
Evaluate your LLM's response with Prometheus and GPT4
A Tree Search Library with Flexible API for LLM Inference-Time Scaling
Open-source AI marketing skills for Claude Code
Autonomous harness engineering
An Efficient Web-enhanced Question Answering System
Open source platform for the machine learning lifecycle
Agent Zero AI framework
The open source post-building layer for agents
Codex plugin that turns attached object images into code-only
The highest-scoring AI memory system ever benchmarked
A batteries-included library for building AI-powered software
Language Model Reinforcement Learning Environments frameworks
Zheng Xi (Efonda Fund Manager) Investment Research Agent Skill
AI agent to evaluate and score resumes
"VideoRAG: Chat with Your Videos
The most accurate natural language detection library for Python
Uncertainty Quantification for Language Models, is a Python package
Requirement-driven evaluation harness for AI agents and LLM
Workflow that turns every post into a calibrated experiment
A specialized Claude Code workspace for creating long-form
Open-source evaluation toolkit of large multi-modality models (LMMs)
Multimodal embedding and reranking models built on Qwen3-VL
CLIP, Predict the most relevant text snippet given an image