CogView4, CogView3-Plus and CogView3(ECCV 2024)
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Multimodal Diffusion with Representation Alignment
Qwen3 is the large language model series developed by Qwen team
Multimodal-Driven Architecture for Customized Video Generation
GLM-4 series: Open Multilingual Multimodal Chat LMs
Uncommon Objects in 3D dataset
Open Source Speech Language Model
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
A Family of Open Sourced Music Foundation Models
Strong, Economical, and Efficient Mixture-of-Experts Language Model
Agentic, Reasoning, and Coding (ARC) foundation models
CLIP, Predict the most relevant text snippet given an image
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
A series of math-specific large language models of our Qwen2 series
GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image
Chinese and English multimodal conversational language model
RGBD video generation model conditioned on camera input
Analyze computation-communication overlap in V3/R1
Open Multilingual Multimodal Chat LMs
Official code for Style Aligned Image Generation via Shared Attention
GLIDE: a diffusion-based text-conditional image synthesis model
A library for Multilingual Unsupervised or Supervised word Embeddings
CTC-based forced aligner for audio-text in 158 languages