ControlNet is a neuralnetwork architecture designed to add conditional control to text-to-image diffusion models. Rather than training from scratch, ControlNet “locks” the weights of a pre-trained diffusion model and introduces a parallel trainable branch that learns additional conditions—like edges, depth maps, segmentation, human pose, scribbles, or other guidance signals.
Efficient Image Captioning code in Torch, runs on GPU
NeuralTalk2 is a Torch-based image-captioning system that generates natural-language descriptions for images with neural networks. It improves on the original NeuralTalk implementation through batching, GPU acceleration, and a more efficient training pipeline. The model combines convolutional neuralnetwork image features with a recurrent neuralnetwork language model. It supports fine-tuning the underlying CNN instead of relying only on fixed visual features. ...