CTC-based forced aligner for audio-text in 158 languages
Qwen3-Next: 80B instruct LLM with ultra-long context up to 1M tokens
Hermes 4 FP8: hybrid reasoning Llama-3.1-405B model by Nous Research
Multimodal 7B model for image, video, and text understanding tasks
Instruction-tuned 1.2B LLM for multilingual text generation by Meta
CLIP model fine-tuned for zero-shot fashion product classification