Tool for exploring and debugging transformer model behaviors
General-purpose image editing model that delivers high-fidelity
Foundation Models for Time Series
Hackable and optimized Transformers building blocks
CogView4, CogView3-Plus and CogView3(ECCV 2024)
A Multi-Modal World Model for Reconstructing, Generating, Simulation
A Systematic Framework for Interactive World Modeling
Global weather forecasting model using graph neural networks and JAX
GPT4V-level open-source multi-modal model based on Llama3-8B
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Generating Immersive, Explorable, and Interactive 3D Worlds
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Inference code for scalable emulation of protein equilibrium ensembles
Audio foundation model excelling in audio understanding
Revolutionizing Database Interactions with Private LLM Technology
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Tiny vision language model
Open-source image generative foundation model
Video Object and Interaction Deletion
Z80-μLM is a 2-bit quantized language model
Multimodal-Driven Architecture for Customized Video Generation
4M: Massively Multimodal Masked Modeling
Official implementation of DreamCraft3D
New family of code large language models (LLMs)
Controllable & emotion-expressive zero-shot TTS