ERNIE-ImageBaidu
|
ModelScopeAlibaba Cloud
|
|||||
Related Products
|
||||||
About
ERNIE-Image is an open text-to-image generation model developed by Baidu, designed to deliver high-quality visuals with strong instruction accuracy and controllability. It is built on a single-stream Diffusion Transformer (DiT) architecture with around 8 billion parameters, allowing it to achieve state-of-the-art performance among open-weight image models while remaining relatively efficient. The model includes a built-in prompt enhancement system that expands simple user inputs into richer, structured descriptions, improving the quality and consistency of generated images. ERNIE-Image is optimized for complex instruction following, enabling accurate rendering of text within images, structured layouts, and multi-element compositions, making it particularly suitable for use cases like posters, comics, and multi-panel designs. It supports multilingual prompts, including English, Chinese, and Japanese, broadening accessibility and usability across regions.
|
About
This model is based on a multi-stage text-to-video generation diffusion model, which inputs a description text and returns a video that matches the text description. Only English input is supported.
This model is based on a multi-stage text-to-video generation diffusion model, which inputs a description text and returns a video that matches the text description. Only English input is supported.
The text-to-video generation diffusion model consists of three sub-networks: text feature extraction, text feature-to-video latent space diffusion model, and video latent space to video visual space. The overall model parameters are about 1.7 billion. Support English input. The diffusion model adopts the Unet3D structure, and realizes the function of video generation through the iterative denoising process from the pure Gaussian noise video.
|
|||||
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Designers, marketers, and content creators who need precise, high-quality AI image generation with strong control over layout, text, and visual composition
|
Audience
Users interested in an open source text-to-video AI video generation model
|
|||||
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Not Supported
|
|||||
API
Offers API
Not Supported
|
API
Offers API
Not Supported
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
No information available.
Free Version
Not Supported
Free Trial
Not Supported
|
Pricing
Free
Free Version
Supported
Free Trial
Not Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
|||||
Company InformationBaidu
Founded: 2000
China
ernie.baidu.com/blog/posts/ernie-image/
|
Company InformationAlibaba Cloud
China
modelscope.cn/
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
||||||
|
|
||||||
|
|
|
|||||
Categories |
Categories |
|||||
Integrations
01.AI
Not Supported
GLM-4.5
Not Supported
Qwen
Not Supported
Qwen-Image
Not Supported
Qwen2
Not Supported
Qwen2-VL
Not Supported
Qwen2.5
Not Supported
Qwen2.5-1M
Not Supported
Qwen2.5-Coder
Not Supported
Qwen2.5-Max
Not Supported
|
Integrations
01.AI
Supported
GLM-4.5
Supported
Qwen
Supported
Qwen-Image
Supported
Qwen2
Supported
Qwen2-VL
Supported
Qwen2.5
Supported
Qwen2.5-1M
Supported
Qwen2.5-Coder
Supported
Qwen2.5-Max
Supported
|
|||||
|
|
|