Mercury 2.5Inception
|
RankLLMCastorini
|
|||||
Related Products
|
||||||
About
Mercury 2.5 is Inception’s most capable production model yet and a significant step up in quality over Mercury 2 while maintaining the same low-latency serving profile. It is the most capable diffusion LLM on the market and, according to Inception, the largest diffusion language model ever trained. Mercury 2.5 delivers a 40% increase in intelligence over Mercury 2, with performance comparable to cost-optimized frontier models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. It generates at 1,107 tokens per second on widely available NVIDIA GPUs and supports a 260K-token context window. Capabilities include tunable reasoning, parallel tool calls, and schema-aligned JSON. The model is designed for latency-sensitive workloads where many model calls may happen inside a single interaction. In search agents and RAG pipelines, it can support planning, query rewriting, reranking, fact structuring, source summarization, and answer checking while keeping calls fast.
|
About
RankLLM is a Python toolkit for reproducible information retrieval research using rerankers, with a focus on listwise reranking. It offers a suite of rerankers, pointwise models like MonoT5, pairwise models like DuoT5, and listwise models compatible with vLLM, SGLang, or TensorRT-LLM. Additionally, it supports RankGPT and RankGemini variants, which are proprietary listwise rerankers. It includes modules for retrieval, reranking, evaluation, and response analysis, facilitating end-to-end workflows. RankLLM integrates with Pyserini for retrieval and provides integrated evaluation for multi-stage pipelines. It also includes a module for detailed analysis of input prompts and LLM responses, addressing reliability concerns with LLM APIs and non-deterministic behavior in Mixture-of-Experts (MoE) models. The toolkit supports various backends, including SGLang and TensorRT-LLM, and is compatible with a wide range of LLMs.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
Developers and enterprises seeking a high-speed diffusion language model for search, RAG, voice agents, coding assistants, and latency-sensitive AI applications
|
Audience
Academic researchers and developers seeking a solution offering tools for implementing and evaluating listwise reranking with large language models
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
No information available.
Free Version
Free Trial
|
Pricing
Free
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationInception
United States
www.inceptionlabs.ai/blog/introducing-mercury-2-5
|
Company InformationCastorini
Canada
github.com/castorini/rank_llm/
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
||||||
|
|
||||||
Categories |
Categories |
|||||
Integrations
Gemini
Gemini Enterprise
JSON
Llama
Mistral AI
NVIDIA TensorRT
OpenAI
Python
Qwen
RankGPT
|
Integrations
Gemini
Gemini Enterprise
JSON
Llama
Mistral AI
NVIDIA TensorRT
OpenAI
Python
Qwen
RankGPT
|
|||||
|
|
|