Implementation of Vision Transformer, a simple way to achieve SOTA
Build cross-modal and multimodal applications on the cloud
Data Lake for Deep Learning. Build, manage, and query datasets
Multi-source content processor for NotebookLM
Multimodal embedding and reranking models built on Qwen3-VL
Free, high-quality text-to-speech API endpoint to replace OpenAI
A set of Docker images for training and serving models in TensorFlow
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
A Codex Skill for generating high-density, editable PowerPoints
Learning agent trained in a diffusion world model
LISA: Reasoning Segmentation via Large Language Model
Skywork-R1V is an advanced multimodal AI model series
Framework for building neural networks
Refer and Ground Anything Anywhere at Any Granularity
Machine Learning Pipelines for Kubeflow
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
A neural network that transforms a design mock-up into static websites
code for Mesh R-CNN, ICCV 2019
Language modeling in a sentence representation space
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Create HTML profiling reports from pandas DataFrame objects
A python library for self-supervised learning on images
Jittor is a high-performance deep learning framework
Python SDK for the Computer Use model Lux, developed by OpenAGI
Official Repo For "Sa2VA: Marrying SAM2 with LLaVA