Qwen2.5-VL is the multimodal large language model series
Open source framework for deep learning satellite and aerial imagery
High-Performance Face Recognition Library on PaddlePaddle & PyTorch
A large-scale model of medical consultation in Chinese
LongBench v2 and LongBench (ACL 25'&24')
Hypernetworks that adapt LLMs for specific benchmark tasks
LISA: Reasoning Segmentation via Large Language Model
Leaderboard Comparing LLM Performance at Producing Hallucinations
Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Pretrained time-series foundation model developed by Google Research
Set of tools to assess and improve LLM security
Code to accompany "A Method for Animating Children's Drawings"
A simple screen parsing tool towards pure vision based GUI agent
The open-source tool for building high-quality datasets
Style-Bert-VITS2: Bert-VITS2 with more controllable voice styles
Controllable and fast Text-to-Speech for over 7000 languages
An Efficient, Scalable, Multi-Modality RL Training Framework
DeepMind model for tracking arbitrary points across videos & robotics
code for Mesh R-CNN, ICCV 2019
Implementation of the Surya Foundation Model for Heliophysics
Easy-to-use and powerful NLP library with Awesome model zoo
Create HTML profiling reports from pandas DataFrame objects
Multi-Agent daTa geneRation Infra and eXperimentation framework
Chinese and English multimodal conversational language model
Foundational model for human-like, expressive TTS