A Customizable Image-to-Video Model based on HunyuanVideo
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
Ollama Telegram bot, with advanced configuration
Long-form streaming TTS system for multi-speaker dialogue generation
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Ollama Python library
Controllable & emotion-expressive zero-shot TTS
DeepSeek Coder: Let the Code Write Itself
This repos contains notebooks for the Advanced Solutions Lab
A command-line productivity tool powered by AI large language models
The best way to use Hermes Agent from the web or from your phone
Unified Multimodal Understanding and Generation Models
Cosmos-RL is a flexible and scalable Reinforcement Learning framework
Open-source multi-speaker long-form text-to-speech model
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
New family of code large language models (LLMs)
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Renderer for the harmony response format to be used with gpt-oss
Designed for text embedding and ranking tasks
Internet-scale Neural Networks
No-code multi-agent framework to build LLM Agents, workflows
4M: Massively Multimodal Masked Modeling
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Open-source large language model family from Tencent Hunyuan