Document Image Parsing via Heterogeneous Anchor Prompting”
Implementation of Vision Transformer, a simple way to achieve SOTA
Build cross-modal and multimodal applications on the cloud
A self-hosted PaaS for deploying and managing web apps
Toolkit for running TensorFlow training scripts on SageMaker
CLI tool to build, test, debug, and deploy Serverless applications
Data Lake for Deep Learning. Build, manage, and query datasets
Segmentation models with pretrained backbones. PyTorch
LINE Messaging API SDK for Python
A Django content management system focused on flexibility & UX
Design clear, theme-specific GitHub README homepages
An AI-agent skill that turns Markdown into paste-ready WeChat article
Multi-source content processor for NotebookLM
Harmonized and Coherent Human Image Animation
Multimodal embedding and reranking models built on Qwen3-VL
High-quality implementations of standard and SOTA methods
Free, high-quality text-to-speech API endpoint to replace OpenAI
A set of Docker images for training and serving models in TensorFlow
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
A Codex Skill for generating high-density, editable PowerPoints
Learning agent trained in a diffusion world model
Python crawler and API for downloading JMComic albums and images
LISA: Reasoning Segmentation via Large Language Model
Skywork-R1V is an advanced multimodal AI model series