Tiny vision language model
Lets make video diffusion practical
High-Resolution Image Synthesis with Latent Diffusion Models
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
GLM-4 series: Open Multilingual Multimodal Chat LMs
LTX-Video Support for ComfyUI
Advancing Open-source World Models
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Python inference and LoRA trainer package for the LTX-2 audio–video
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Official inference repo for FLUX.2 models
CodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)
CLIP, Predict the most relevant text snippet given an image
An experimental version of DeepSeek model
Diversity-driven optimization and large-model reasoning ability
Image generation model with single-stream diffusion transformer
Designed for text embedding and ranking tasks
Run the full 2.78-trillion-parameter Kimi K3 model
4M: Massively Multimodal Masked Modeling
Repo of Qwen2-Audio chat & pretrained large audio language model
Moonshot's most powerful AI model
Infinite Worlds with Versatile Interactions
Inference code for scalable emulation of protein equilibrium ensembles
Codex plugin that turns attached object images into code-only