Autonomous Agents (LLMs) research papers. Updated Daily
Medical imaging toolkit for deep learning
Refer and Ground Anything Anywhere at Any Granularity
Qwen3-VL, the multimodal large language model series by Alibaba Cloud
Think with AI beyond the chat box
Project Lyra: Open Generative 3D World Models
State-of-the-art (SoTA) text-to-video pre-trained model
A high performance implementation of HDBSCAN clustering
Models for object and human mesh reconstruction
Visual Causal Flow
Claw3D is an open source 3D engine built on OpenClaw
A Codex Skill for generating high-density, editable PowerPoints
A complete AI agency at your fingertips
Qwen2.5-VL is the multimodal large language model series
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Open-source 2D IDE for managing AI agents in native CLIs
HeavyDB (formerly MapD/OmniSciDB)
Build your own AI application system for free
Unsupervised Learning for Image Registration
Video understanding codebase from FAIR for reproducing video models
Unifying 3D Mesh Generation with Language Models
Foundational Models for State-of-the-Art Speech and Text Translation
Gracefully face hCaptcha challenge with multimodal llms
A Systematic Framework for Interactive World Modeling
code for Mesh R-CNN, ICCV 2019