BenchGen
BenchGen is the learning infrastructure for AI agents: an open platform where developers discover benchmarks and RL environments, evaluate their complete agent system — model and harness together — against verifiable rewards, and export clean trajectory data for fine-tuning. One loop: benchmark → evaluate → fine-tune → re-evaluate.
Learn more
FutureHouse
FutureHouse is a nonprofit AI research lab focused on automating scientific discovery in biology and other complex sciences. FutureHouse features superintelligent AI agents designed to assist scientists in accelerating research processes. It is optimized for retrieving and summarizing information from scientific literature, achieving state-of-the-art performance on benchmarks like RAG-QA Arena's science benchmark. It employs an agentic approach, allowing for iterative query expansion, LLM re-ranking, contextual summarization, and document citation traversal to enhance retrieval accuracy. FutureHouse also offers a framework for training language agents on challenging scientific tasks, enabling agents to perform tasks such as protein engineering, literature summarization, and molecular cloning. Their LAB-Bench benchmark evaluates language models on biology research tasks, including information extraction, database retrieval, etc.
Learn more
GLM-5.3-Flash
GLM-5.3-Flash is Z.ai’s natively multimodal model in the GLM-5 series (previously previewed as Ox Alpha), designed to deliver strong coding, agentic, visual, and knowledge-work performance at relatively low inference cost. It uses 320 billion total parameters with 18 billion active parameters, along with a hybrid architecture that combines sparse and linear attention to reduce the cost of long-context processing. The model supports context lengths of up to one million tokens and was trained on a 30-trillion-token multimodal corpus. GLM-5.3-Flash can reason across text, images, documents, interfaces, dashboards, and other visual information while using that feedback to refine its own outputs. Z.ai reports substantial gains over GLM-5.2 on coding and agentic benchmarks, including DeepSWE and AutomationBench, while approaching higher-cost frontier models on several evaluations.
Learn more
GLM-4.7
GLM-4.7 is an advanced large language model designed to significantly elevate coding, reasoning, and agentic task performance. It delivers major improvements over GLM-4.6 in multilingual coding, terminal-based tasks, and real-world software engineering benchmarks such as SWE-bench and Terminal Bench. GLM-4.7 supports “thinking before acting,” enabling more stable, accurate, and controllable behavior in complex coding and agent workflows. The model also introduces strong gains in UI and frontend generation, producing cleaner webpages, better layouts, and more polished slides. Enhanced tool-using capabilities allow GLM-4.7 to perform more effectively in web browsing, automation, and agent benchmarks. Its reasoning and mathematical performance has improved substantially, showing strong results on advanced evaluation suites. GLM-4.7 is available via Z.ai, API platforms, coding agents, and local deployment for flexible adoption.
Learn more