"Big Model" trains a visual multimodal VLM with 26M parameters
The official repo of Qwen chat & pretrained large language model
The best free open source website change detection and restock service
Edit videos with Claude Code
A lightweight approach to removing Google web service dependency
OCR expert VLM powered by Hunyuan's native multimodal architecture
Lightweight Markdown-only skills for autonomous ML research
ChatGPT extension for scientific research work
Extract one time password (OTP) secrets from QR codes
21 Lessons, Get Started Building with Generative AI
State-of-the-art diffusion models for image and audio generation
Main repository for the Sphinx documentation builder
Repo of Qwen2-Audio chat & pretrained large audio language model
Skills, a Chinese software copyright application material generator
Make any agent harness multimodal-native
Faster and easier training and deployments
MOSS‑TTS Family open‑source speech and sound generation model
Diffusion Transformer with Fine-Grained Chinese Understanding
LLM-based Reinforcement Learning audio edit model
Concatenate a directory full of files into a single prompt
Python library for scraping and analyzing online news articles easily
tensorboard for pytorch (and chainer, mxnet, numpy, etc.)
Multilingual sentence & image embeddings with BERT
LLM training code for MosaicML foundation models
AutoGluon: AutoML for Image, Text, and Tabular Data