Multimodal Transformer for document image understanding and layout
Flexible text-to-text transformer model for multilingual NLP tasks
Portuguese ASR model fine-tuned on XLSR-53 for 16kHz audio input
CLIP ViT-bigG/14: Zero-shot image-text model trained on LAION-2B
Program to pull out specific motifs from bam sequence files
Instruction-tuned 7B language model for chat and complex tasks
T5-Small: Lightweight text-to-text transformer for NLP tasks
Summarization model fine-tuned on CNN/DailyMail articles
CTC-based forced aligner for audio-text in 158 languages
Vision-language-action model for robot control via images and text
CLIP model fine-tuned for zero-shot fashion product classification