Uncommon Objects in 3D dataset
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Automatic Speech Recognition with Word-level Timestamps
A dataset consists of 15,140 ChatGPT prompts from Reddit
Open source AI model for generating full songs from lyrics prompts
MiniMax H3 is a general-purpose, omni-modal generative system
Qwen3 is the large language model series developed by Qwen team
High-Performance Face Recognition Library on PaddlePaddle & PyTorch
Automatic subtitle synchronization tool
Super timeline all the things
Recipes to train reward model for RLHF
Multimodal-Driven Architecture for Customized Video Generation
Handwritten Text Recognition (HTR) system implemented with TensorFlow
The Triton Inference Server provides an optimized cloud
Pretrained (Language) Models for Probabilistic Time Series Forecasting
Pluggable SOTA multi-object tracking modules for segmentation
A trainable PyTorch reproduction of AlphaFold 3
tiktoken is a fast BPE tokeniser for use with OpenAI's models
HivisionIDPhotos: a lightweight and efficient AI ID photos tools
Genome modeling and design across all domains of life
SOTA Open Source TTS
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Semi-Structured Agentic Framework. Workflows build themselves
A Survey of Large Language Models
A Unified Framework for Image Customization