Open source AI model for generating full songs from lyrics prompts
Knowledge Graph Generation from Any Text
AI-Powered Data Processing: Use LOTUS to process all of your datasets
Gracefully face hCaptcha challenge with multimodal llms
Tools for merging pretrained large language models
Proofs, cases, concept supplements, and reference explanations
Flexible Photo Recrafting While Preserving Your Identity
Build Vision Agents quickly with any model or video provider
A fast TTS architecture with conditional flow matching
Agent framework and applications built upon Qwen>=3.0
Python binding to the Apache Tika™ REST services
State-of-the-art (SoTA) text-to-video pre-trained model
Large Multimodal Models for Video Understanding and Editing
Benchmarking synthetic data generation methods
An advanced paper search agent powered by large language models
Large-language-model & vision-language-model based on Linear Attention
Capable of understanding text, audio, vision, video
Full stack AI software engineer
Benchmarking Multimodal Agents for Open-Ended Tasks
Visual Automation IDE — automate anything you see on screen
Constrained Value Alignment via Safe Reinforcement Learning
AI-powered PC monitoring that explains. Not shows numbers/spikes.
Automatically Visualize any dataset, any size
Plug-n-play module turning text-to-image models into animation
An opinionated CLI to transcribe Audio files w/ Whisper on-device