Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Multimodal-Driven Architecture for Customized Video Generation
Clean and efficient FP8 GEMM kernels with fine-grained scaling
FlashMLA: Efficient Multi-head Latent Attention Kernels
Chinese and English multimodal conversational language model
An experimental version of DeepSeek model
A Powerful Native Multimodal Model for Image Generation
AI cognitive-enhancement Skills based on Anthropic's J-space
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
High-Fidelity and Controllable Generation of Textured 3D Assets
Distribution TN 365 KDE moderne et stable !
Runtime extension of Proximus enabling Deployment on AMD Ryzen™ AI
Learning to Act by Watching Unlabeled Online Videos
Metric monocular depth estimation (vision model)