OmniParser
OmniParser is a comprehensive method for parsing user interface screenshots into structured elements, significantly enhancing the ability of multimodal models like GPT-4 to generate actions accurately grounded in corresponding regions of the interface. It reliably identifies interactable icons within user interfaces and understands the semantics of various elements in a screenshot, associating intended actions with the correct screen regions. To achieve this, OmniParser curates an interactable icon detection dataset containing 67,000 unique screenshot images labeled with bounding boxes of interactable icons derived from DOM trees. Additionally, a collection of 7,000 icon-description pairs is used to fine-tune a caption model that extracts the functional semantics of detected elements. Evaluations on benchmarks such as SeeClick, Mind2Web, and AITW demonstrate that OmniParser outperforms GPT-4V baselines, even when using only screenshot inputs without additional information.
Learn more
Manifest
Manifest by Omfang is an action layer for AI agents that turns any webpage into a structured map of what an agent can actually do. Developers send a URL to the Manifest API and receive clean JSON describing the page state, available buttons, forms, inputs, navigation links, required fields, and dependencies between actions. Manifest loads the page in a real Chromium browser after JavaScript runs, reads the accessibility tree used by screen readers, and cross-references it with the underlying DOM to capture details the accessibility layer can miss, including input types, placeholders, disabled states, and prerequisites for submission. Once actions are identified, their locators are resolved deterministically rather than generated by a model; ambiguous matches are marked uncertain instead of guessed. This gives agents a machine-readable interface for reasoning about what to click, fill, or submit without relying on screenshots, brittle CSS paths, etc.
Learn more
Thesys Agent Builder
Thesys Agent Builder is a no-code platform for creating and publishing interactive AI applications that respond with structured user interfaces (charts, cards, tables, forms, slides, reports, and more) instead of simple text, powered by the underlying C1 Generative UI engine. You can upload data (files, URLs, databases), connect to tools, configure the agent’s behavior and tone with natural language instructions, pick styles to match your brand, and publish a live agent that can be used on the web or embedded in your site without writing frontend code. It lets you build AI agents that can parse and visualize data, answer questions with interactive insights, generate visual summaries and reports, and provide rich, action-oriented responses from connected data sources in just minutes. It emphasizes real interactivity and practical utility by turning conversations into actionable UI elements that help users explore, analyze, and act on information.
Learn more
netTerrain DCIM
netTerrain is an automated and interactive visual diagraming and reporting solution that renders real-word views of your entire IT ecosystem—from data centers to networks, fiber, and cloud. netTerrain's interactive maps and reports replace scattered documentation and guesswork with clarity: cut costs, troubleshoot faster, prevent downtime, reduce field visits, and instantly find and share vital information.
You get high-level overviews and details on capacity, power, security patches, work orders, and more. With actionable insights, you can now visualize and understand any element in your IT ecosystem and make the correct business decisions every time!
Learn more