Search Results for "cursive image text" - Page 34

Showing 865 open source projects for "cursive image text"

View related business solutions
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    CLIP-ViT-bigG-14-laion2B-39B-b160k

    CLIP-ViT-bigG-14-laion2B-39B-b160k

    CLIP ViT-bigG/14: Zero-shot image-text model trained on LAION-2B

    CLIP-ViT-bigG-14-laion2B-39B-b160k is a powerful vision-language model trained on the English subset of the LAION-5B dataset using the OpenCLIP framework. Developed by LAION and trained by Mitchell Wortsman on Stability AI’s compute infrastructure, it pairs a ViT-bigG/14 vision transformer with a text encoder to perform contrastive learning on image-text pairs. This model excels at zero-shot image classification, image-to-text and text-to-image retrieval, and can be adapted for tasks such as image captioning or generation guidance. It achieves an impressive 80.1% top-1 accuracy on ImageNet-1k without any fine-tuning, showcasing its robustness in open-domain settings. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Qwen2.5-VL-3B-Instruct

    Qwen2.5-VL-3B-Instruct

    Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video

    Qwen2.5-VL-3B-Instruct is a 3.75 billion parameter multimodal model by Qwen, designed to handle complex vision-language tasks in both image and video formats. As part of the Qwen2.5 series, it supports image-text-to-text generation with capabilities like chart reading, object localization, and structured data extraction. The model can serve as an intelligent visual agent capable of interacting with digital interfaces and understanding long-form videos by dynamically sampling resolution and frame rate. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Krea 2 Turbo

    Krea 2 Turbo

    Fast 12B image model for high-quality text-to-image generation

    Krea 2 Turbo is Krea AI’s flagship open-weight text-to-image diffusion model optimized for fast, high-quality image generation. Built on a 12-billion-parameter Diffusion Transformer (DiT), it is a distilled and post-trained version of the Krea 2 Raw checkpoint, enabling photorealistic and artistic image synthesis in as few as eight inference steps. Designed for creative professionals, developers, and researchers, it supports concept art, design exploration, marketing assets, illustrations, and commercial visual production. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    layoutlm-base-uncased

    layoutlm-base-uncased

    Multimodal Transformer for document image understanding and layout

    layoutlm-base-uncased is a multimodal transformer model developed by Microsoft for document image understanding tasks. It incorporates both text and layout (position) features to effectively process structured documents like forms, invoices, and receipts. This base version has 113 million parameters and is pre-trained on 11 million documents from the IIT-CDIP dataset. LayoutLM enables better performance in tasks where the spatial arrangement of text plays a crucial role. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Fully Managed MySQL, PostgreSQL, and SQL Server Icon
    Fully Managed MySQL, PostgreSQL, and SQL Server

    Automatic backups, patching, replication, and failover. Focus on your app, not your database.

    Cloud SQL handles your database ops end to end, so you can focus on your app.
    Start Free
  • 5
    Gemma 4 12B

    Gemma 4 12B

    Unified multimodal Gemma model for local coding and reasoning

    Gemma 4 12B is Google DeepMind’s unified open-weight multimodal model designed for efficient local reasoning, coding, and multimodal understanding. Unlike other Gemma 4 models that rely on separate encoders, the 12B Unified model uses an encoder-free architecture that projects raw image patches and audio waveforms directly into the language model’s embedding space, reducing multimodal latency and simplifying fine-tuning. It supports text, image, audio, and video inputs with text output, making it useful for transcription, image understanding, video analysis, coding, and agentic workflows. The model has 11.95B parameters, 48 layers, a 256K-token context window, and support for over 140 languages. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    fashion-clip

    fashion-clip

    CLIP model fine-tuned for zero-shot fashion product classification

    FashionCLIP is a domain-adapted CLIP model fine-tuned specifically for the fashion industry, enabling zero-shot classification and retrieval of fashion products. Developed by Patrick John Chia and collaborators, it builds on the CLIP ViT-B/32 architecture and was trained on over 800K image-text pairs from the Farfetch dataset. The model learns to align product images and descriptive text using contrastive learning, enabling it to perform well across various fashion-related tasks without additional supervision. FashionCLIP 2.0, the latest version, uses the laion/CLIP-ViT-B-32-laion2B-s34B-b79K checkpoint for improved accuracy, achieving better F1 scores across multiple benchmarks compared to earlier versions. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    translategemma-4b-it

    translategemma-4b-it

    Lightweight multimodal translation model for 55 languages

    translategemma-4b-it is a lightweight, state-of-the-art open translation model from Google, built on the Gemma 3 family and optimized for high-quality multilingual translation across 55 languages. It supports both text-to-text translation and image-to-text extraction with translation, enabling workflows such as OCR-style translation of signs, documents, and screenshots. With a compact ~5B parameter footprint and BF16 support, the model is designed to run efficiently on laptops, desktops, and private cloud infrastructure, making advanced translation accessible without heavy hardware requirements. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    A software application for the optical recognition, the superimposition and the collation of early music prints
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    FLUX.1-Krea-dev

    FLUX.1-Krea-dev

    Text-to-image model optimized for artistic quality and safe generation

    FLUX.1-Krea-dev is a 12 billion parameter rectified flow transformer for text-to-image generation, developed by Black Forest Labs in collaboration with Krea. It delivers aesthetic, high-quality outputs focused on photography and visual coherence, making it a strong competitor to closed-source models. Trained using guidance distillation, it offers efficient inference while preserving creative fidelity. The model is distributed under a non-commercial license, with conditions to prevent misuse and support ethical AI development. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • 10
    Nex-N2-mini

    Nex-N2-mini

    Compact agentic model for coding, tools, and productivity tasks

    ...It uses adaptive thinking to decide when deeper reasoning is needed and coherent thinking to keep reasoning consistent across tasks and modalities. Nex-N2-mini supports image-text-to-text workflows, explicit reasoning traces, robust function calling, and deployment through Transformers, vLLM, SGLang, Docker, and quantized local apps. It performs strongly across agentic, coding, search, and reasoning benchmarks, including SWE-Bench, Terminal-Bench, BrowseComp, Toolathlon, and GPQA.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Nex-N2-Pro

    Nex-N2-Pro

    Large agentic model for coding, tools, research, and execution

    ...The model is built on Qwen3.5-397B-A17B and is designed as the high-quality counterpart to Nex-N2-mini, trading higher compute needs for stronger reasoning and agent performance. It supports image-text-to-text workflows, explicit reasoning traces, robust function calling, and deployment through Transformers, vLLM, SGLang, Docker, and quantized local apps. Nex-N2-Pro performs strongly across agentic, coding, search, and reasoning benchmarks, including Terminal-Bench, SWE-Bench Pro, BrowseComp, Toolathlon, WideSearch, GPQA Diamond, and GDPval.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    Hy3

    Hy3

    Open code agent for Lean 4 proofs and formal software verification

    ...The model uses 128 experts with four active for each token and supports a 256K-token context window, making it suitable for extended formal reasoning and large verification tasks. Leanstral accepts text and image inputs and produces text output, enabling multimodal workflows around mathematics, code, and specifications. It supports configurable reasoning effort, allowing users to disable reasoning or enable high-effort reasoning for complex prompts. This updated version of the original Leanstral focuses on performant, cost-effective formal coding and theorem-proving workflows.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    Leanstral 1.5

    Leanstral 1.5

    Open code agent for Lean 4 proofs and formal software verification

    ...The model uses 128 experts with four active for each token and supports a 256K-token context window, making it suitable for extended formal reasoning and large verification tasks. Leanstral accepts text and image inputs and produces text output, enabling multimodal workflows around mathematics, code, and specifications. It supports configurable reasoning effort, allowing users to disable reasoning or enable high-effort reasoning for complex prompts. This updated version of the original Leanstral focuses on performant, cost-effective formal coding and theorem-proving workflows.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Gemma 4

    Gemma 4

    Google’s flagship dense multimodal model for coding and reasoning

    Gemma 4 is Google DeepMind’s flagship dense open-weight multimodal model, designed for high-end reasoning, coding, agentic workflows, and multimodal understanding. The model contains approximately 30.7B parameters and supports text and image inputs with text generation output, while also processing video as image-frame sequences. Built as the most capable model in the Gemma 4 family, it combines strong reasoning performance with a large 256K-token context window and configurable thinking modes. Gemma 4 31B supports native function calling, structured outputs, and more than 140 languages, making it suitable for enterprise assistants, coding agents, document analysis, and multilingual applications. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Command A+

    Command A+

    4-bit Command A+ model for enterprise agents and multilingual tasks

    Command A+ 05-2026 W4A4 is a 4-bit quantized version of Cohere’s open-source Command A+ model, optimized for enterprise-grade agentic, multilingual, and reasoning-heavy workloads. It supports text and image inputs, generates text outputs, and uses a sparse Mixture-of-Experts Transformer architecture with 218B total parameters and 25B active parameters. The W4A4 release applies 4-bit weight and activation quantization mainly to MoE experts, preserving attention components at full precision to reduce quality loss while improving speed, latency, and hardware efficiency. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    NoobAI XL 1.1

    NoobAI XL 1.1

    Open, non-commercial SDXL model for quality image generation

    NoobAI XL 1.1 is a diffusion-based text-to-image generative model developed by Laxhar Dream Lab, fine-tuned from NoobAI XL 1.0 and built upon Illustrious-xl. It leverages the latest Danbooru and e621 datasets, using native tag captions to enhance visual fidelity, style accuracy, and prompt responsiveness. The model introduces refined quality tagging, ranking images by percentile to ensure results reflect modern aesthetic preferences.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    themaCreator

    themaCreator

    themaCreator - create posts from files

    themaCreator is an auto posts creator which allows you to generate full posts from files. It has image and file uploading features, IMDB grabber, extract all information from files and more. The program helps to create posts from files.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Qwen3.6-27B

    Qwen3.6-27B

    Dense multimodal Qwen model for coding, agents, and long context

    Qwen3.6-27B is an open-weight multimodal model built to deliver strong real-world coding, agent, and long-context performance in a dense 27B-parameter architecture. It combines a causal language model with a vision encoder and supports text, image, and video inputs, making it suitable for both software workflows and broader multimodal tasks. The model emphasizes stability and practical developer utility, with major improvements in agentic coding, frontend generation, and repository-level reasoning. It also introduces thinking preservation, allowing it to retain reasoning traces from earlier turns to improve consistency, reduce repeated computation, and support iterative agent workflows. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    GLM-5.3-Flash

    GLM-5.3-Flash

    Efficient 320B multimodal MoE model for coding and autonomous agents

    ...It also uses Manifold-Constrained Hyper-Connections (mHC) to improve scaling efficiency and was pretrained on a 30-trillion-token multimodal corpus. GLM-5.3-Flash supports text and image inputs and is particularly optimized for coding and autonomous agent workloads, approaching larger frontier models on related benchmarks while improving over GLM-5.2. It supports local deployment through SGLang, vLLM, TokenSpeed, and KTransformers and is released under the MIT license.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Muse Glimmer

    Muse Glimmer

    Local multimodal 30B model for autonomous agents, coding, and tools

    Muse Glimmer-30B is Meta Superintelligence Lab’s open-weight multimodal model built specifically for autonomous agentic tasks on consumer hardware. Distilled from the larger Muse Spark, it combines multi-step reasoning, reliable tool use, coding, failure recovery, and image understanding in a dense 29.6B-parameter architecture with a dedicated 1.8B-parameter perception encoder. It supports more than 100 languages and a 131K+ token context window, allowing agents to maintain coherent plans across extended workflows. Muse Glimmer can interpret screenshots, charts, documents, and images alongside text, while configurable reasoning strength lets developers balance response quality and speed. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Qwen3.6-35B-A3B-FP8

    Qwen3.6-35B-A3B-FP8

    FP8 Qwen model for efficient multimodal coding and agent tasks

    Qwen3.6-35B-A3B-FP8 is an FP8-quantized version of Qwen3.6 designed to deliver nearly the same performance as the original model while improving deployment efficiency. It is a multimodal open-weight model that combines a causal language model with a vision encoder, supporting text, image, and video inputs. Built for stability and real-world developer use, it emphasizes agentic coding, repository-level reasoning, and productive long-context workflows. A key capability is thinking preservation, which allows the model to retain reasoning traces from earlier messages, helping reduce repeated computation and improving consistency in iterative tasks. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Ministral 3 8B Base 2512

    Ministral 3 8B Base 2512

    Versatile 8B-base multimodal LLM, flexible foundation for custom AI

    Ministral 3 8B Base 2512 is a mid-sized, dense model in the Ministral 3 series, designed as a general-purpose foundation for text and image tasks. It pairs an 8.4B-parameter language model with a 0.4B-parameter vision encoder, enabling unified multimodal capabilities out of the box. As a “base” model (i.e., not fine-tuned for instruction or reasoning), it offers a flexible starting point for custom downstream tasks or fine-tuning. The model supports a large 256k token context window, making it capable of handling long documents or extended dialogues. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Krea 2 Raw

    Krea 2 Raw

    Base Krea image model for LoRA training and fine-tuning

    Krea 2 Raw is Krea AI’s base open-weight text-to-image diffusion checkpoint, designed primarily for fine-tuning, LoRA training, and post-training rather than direct inference. It is part of the Krea 2 model family and uses a 12-billion-parameter Diffusion Transformer architecture to generate images from natural-language prompts. Unlike Krea 2 Turbo, which is distilled and optimized for faster direct generation, Raw is the foundational checkpoint before additional post-training and fine-tuning. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    Qwen2.5-VL-7B-Instruct

    Qwen2.5-VL-7B-Instruct

    Multimodal 7B model for image, video, and text understanding tasks

    Qwen2.5-VL-7B-Instruct is a multimodal vision-language model developed by the Qwen team, designed to handle text, images, and long videos with high precision. Fine-tuned from Qwen2.5-VL, this 7-billion-parameter model can interpret visual content such as charts, documents, and user interfaces, as well as recognize common objects. It supports complex tasks like visual question answering, localization with bounding boxes, and structured output generation from documents. The model is also...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    JFileManipulator aims to be a collection of tools, which manipulate files in some way or the other, and is aimed at manipulating said files in batches. It for example allows you to write a text to a number of image files, all for a group of files.
    Downloads: 0 This Week
    Last Update:
    See Project