Advancing Open-source World Models
Global weather forecasting model using graph neural networks and JAX
code for Mesh R-CNN, ICCV 2019
Generating Immersive, Explorable, and Interactive 3D Worlds
Open image model at the forefront of design
A Pragmatic VLA Foundation Model
Hunyuan Translation Model Version 1.5
CLIP, Predict the most relevant text snippet given an image
State-of-the-art (SoTA) text-to-video pre-trained model
Miso TTS is an 8 billion, highly emotive text-to-speech model
HY-Motion model for 3D character animation generation
CogView4, CogView3-Plus and CogView3(ECCV 2024)
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
RGBD video generation model conditioned on camera input
Netease Youdao's open-source embedding and reranker models
An Efficient Agentic Model for Computer Use
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Revolutionizing Database Interactions with Private LLM Technology
Pokee Deep Research Model Open Source Repo
DeepMind model for tracking arbitrary points across videos & robotics
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Renderer for the harmony response format to be used with gpt-oss
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
The official PyTorch implementation of Google's Gemma models