Marengo

Marengo

TwelveLabs
ModelScope

ModelScope

Alibaba Cloud
+
+

Related Products

  • LTX
    182 Ratings
    Visit Website
  • 4K Video Downloader
    12,893 Ratings
    Visit Website
  • Google AI Studio
    40 Ratings
    Visit Website
  • Adobe Firefly
    25,030 Ratings
    Visit Website
  • Haast
    4 Ratings
    Visit Website
  • LALAL.AI
    5,355 Ratings
    Visit Website
  • Muzaic
    2 Ratings
    Visit Website
  • 3Q
    14 Ratings
    Visit Website
  • Screencapt
    140 Ratings
    Visit Website
  • pCloud Business
    189 Ratings
    Visit Website

About

Marengo is a multimodal video foundation model that transforms video, audio, image, and text inputs into unified embeddings, enabling powerful “any-to-any” search, retrieval, classification, and analysis across vast video and multimedia libraries. It integrates visual frames (with spatial and temporal dynamics), audio (speech, ambient sound, music), and textual content (subtitles, overlays, metadata) to create a rich, multidimensional representation of each media item. With this embedding architecture, Marengo supports robust tasks such as search (text-to-video, image-to-video, video-to-audio, etc.), semantic content discovery, anomaly detection, hybrid search, clustering, and similarity-based recommendation. The latest versions introduce multi-vector embeddings, separating representations for appearance, motion, and audio/text features, which significantly improve precision and context awareness, especially for complex or long-form content.

About

This model is based on a multi-stage text-to-video generation diffusion model, which inputs a description text and returns a video that matches the text description. Only English input is supported. This model is based on a multi-stage text-to-video generation diffusion model, which inputs a description text and returns a video that matches the text description. Only English input is supported. The text-to-video generation diffusion model consists of three sub-networks: text feature extraction, text feature-to-video latent space diffusion model, and video latent space to video visual space. The overall model parameters are about 1.7 billion. Support English input. The diffusion model adopts the Unet3D structure, and realizes the function of video generation through the iterative denoising process from the pure Gaussian noise video.

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Audience

Media companies, AI researchers, and platforms searching for a tool to build smart search engines, content discovery tools, recommendation systems, or video-analysis workflows

Audience

Users interested in an open source text-to-video AI video generation model

Support

Phone Support
24/7 Live Support
Online

Support

Phone Support
24/7 Live Support
Online

API

Offers API

API

Offers API

Screenshots and Videos

Screenshots and Videos

Pricing

$0.042 per minute
Free Version
Free Trial

Pricing

Free
Free Version
Free Trial

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation
Webinars
Live Online
In Person

Training

Documentation
Webinars
Live Online
In Person

Company Information

TwelveLabs
Founded: 2021
United States
www.twelvelabs.io/product/models-overview#marengo

Company Information

Alibaba Cloud
China
modelscope.cn/

Alternatives

VideoPoet

VideoPoet

Google

Alternatives

FLUX 3

FLUX 3

Black Forest Labs
Kaggle

Kaggle

Google
MiniMax H3

MiniMax H3

MiniMax

Categories

Categories

Integrations

01.AI
CodeQwen
GLM-4.5
Qwen-7B
Qwen-Image
Qwen2.5
Qwen2.5-1M
Qwen2.5-Coder
Qwen2.5-Max
Qwen2.5-VL
Qwen3
Qwen3.6
Qwen3.6-27B
Qwen3.6-35B-A3B
Qwen3.7-Max
Qwen3.7-Plus
Qwen3.8-2.4T-A95B
Qwen3.8-Flash-Next
Qwen3.8-Max
TwelveLabs

Integrations

01.AI
CodeQwen
GLM-4.5
Qwen-7B
Qwen-Image
Qwen2.5
Qwen2.5-1M
Qwen2.5-Coder
Qwen2.5-Max
Qwen2.5-VL
Qwen3
Qwen3.6
Qwen3.6-27B
Qwen3.6-35B-A3B
Qwen3.7-Max
Qwen3.7-Plus
Qwen3.8-2.4T-A95B
Qwen3.8-Flash-Next
Qwen3.8-Max
TwelveLabs
Claim Marengo and update features and information
Claim Marengo and update features and information
Claim ModelScope and update features and information
Claim ModelScope and update features and information