Alternatives to Petuum
Compare Petuum alternatives for your business or organization using the curated list below. SourceForge ranks the best alternatives to Petuum in 2026. Compare features, ratings, user reviews, pricing, and more from Petuum competitors and alternatives in order to make an informed decision for your business.
-
1
DataSet
DataSet
DataSet retains live, searchable real-time insights. Store indefinitely using DataSet-hosted or customer-managed, low-cost S3 storage. Ingest structured, semi-structured, and unstructured data faster than ever before. A limitless enterprise infrastructure for live data queries, analytics, insights, and retention, with no data schema requirements. The technology of choice for engineering, DevOps, IT, and security teams to unlock the power of data. Sub-second query performance powered by a patented parallel processing architecture. Work quicker and smarter to make better business decisions. Ingest hundreds of terabytes effortlessly. No rebalancing nodes, storage management, or resource reallocation. Scale on a limitless flexible platform. An efficient cloud-native architecture minimizes cost and maximizes output. Benefit from a predictable cost model with unmatched performance.Starting Price: $0.99 per GB per day -
2
PanBI
PanApps
Gain valuable business insight by analyzing your structured, semi-structured and unstructured data using interactive visual interfaces. Personalized interactive dashboards provide rich insight to data. Organize and retrieve collections by user preferences and permissions. Add comments against datasets and share for public view or private use. Dynamically build visually appealing informative charts and maps. -
3
BigBI
BigBI
BigBI enables data specialists to build their own powerful big data pipelines interactively & efficiently, without any coding! BigBI unleashes the power of Apache Spark enabling: Scalable processing of real Big Data (up to 100X faster) Integration of traditional data (SQL, batch files) with modern data sources including semi-structured (JSON, NoSQL DBs, Elastic, Hadoop), and unstructured (Text, Audio, video), Integration of streaming data, cloud data, AI/ML & graphs -
4
Nirveda Cognition
Nirveda Cognition
Make Smarter, Faster & More Informed Decisions. Enterprise Document Intelligence Platform to turn data into Actionable Insights. Our versatile platform uses cognitive Machine Learning and Natural Language Processing algorithms to automatically classify, extract, enrich, and integrate relevant, timely, and accurate information from your documents. The solution is delivered as a service to lower the cost of ownership and accelerate time to value. How It Works. CLASSIFY. Ingest structured, semi-structured, or unstructured documents. Identify and classify documents based on semantic understanding of language and visual cues. Extract. Extracts words, short phrases, and sections of text from printed, handwritten, and tabular data. Detects the presence of a signature or page annotation. Easily review and make corrections to the extracted data. AI uses human corrections to learn and improve. Enrich. Customizable data verification, validation, standardization and normalization. -
5
CADfix
International TechneGroup
CADfix is the leading software solution for CAD model translation, repair, healing, defeaturing, and simplification. CADfix tackles the ever-present problems of 3D model data exchange and re-use between different engineering applications, removing the barriers that prevent the re-use of solid models in design, analysis, and manufacturing systems. CADfix delivers breakthroughs in geometry processing, tackling some of the toughest 3D geometry issues affecting the industry. ITI works closely with customers to develop leading-edge novel geometry processing technology that will have a significant impact on engineering process efficiency. Take a look at some of the applications below to see how CADfix meets the demands of today’s leading engineering enterprises. CADfix delivers breakthroughs in geometry processing, tackling some of the toughest 3D geometry issues affecting the industry. ITI works closely with customers to develop leading-edge novel geometry processing technology. -
6
KlearStack
KlearStack
KlearStack offers template-less, automated invoice processing, and thus removes the drudgery of manual entry from unstructured documents. Our mission is to automate the tedious manual processes and exhausting data entry, so that humans are freed for more intelligent and creative tasks! To help organizations make their unstructured data a competitive advantage by unlocking the useful information from unstructured and free-form semi-structured documents. KlearStack’s artificial intelligence today provides best solutions to automate the following processes that involve unstructured documents: Invoice Automation Purchase Order Automation Receipt Capture Consumer Durable Loans Multi-Vendor Trade Finance Process Automation Two Wheeler Loan Automation Used Cars Loan Process Automation With our proprietary template-less AI/ML technology, you don't need to spend hundreds or thousands of days on designing and maintaining templates anymore! Improve productivity by up-to 200 -
7
Recraft
Recraft
Recraft is an AI-powered image generation platform designed to create high-quality visuals with strong design aesthetics. It enables users to generate photorealistic images, vectors, and design assets from simple prompts. The platform stands out for its ability to produce vector graphics directly, making it useful for professional design work. Recraft focuses on delivering visually consistent and stylistically refined outputs without requiring extensive training. Users can easily create and reuse custom styles by uploading reference images. It also includes tools for editing, upscaling, and refining images within a single platform. The system is built to support creative workflows for branding, marketing, and visual content creation. Overall, Recraft helps designers and creators produce polished visuals quickly and efficiently.Starting Price: $10/month -
8
BDB Platform
Big Data BizViz
BDB is a modern data analytics and BI platform which can skillfully dive deep into your data to provide actionable insights. It is deployable on the cloud as well as on-premise. Our exclusive microservices based architecture has the elements of Data Preparation, Predictive, Pipeline and Dashboard designer to provide customized solutions and scalable analytics to different industries. BDB’s strong NLP based search enables the user to unleash the power of data on desktop, tablets and mobile as well. BDB has various ingrained data connectors, and it can connect to multiple commonly used data sources, applications, third party API’s, IoT, social media, etc. in real-time. It lets you connect to RDBMS, Big data, FTP/ SFTP Server, flat files, web services, etc. and manage structured, semi-structured as well as unstructured data. Start your journey to advanced analytics today. -
9
Lityx
Lityx
Empower your team to deliver AI-based business solutions efficiently and at scale. Cloud-based, comprehensive, no-code machine learning. Accelerate team production. Efficiently capture actionable insights. Leverage your data to predict and optimize behaviors. Rapidly scale and take models into production. Gurobi Optimizer allows users to tackle even the toughest challenges. LityxIQ is the powerful yet easy-to-use, no-code AutoML platform built by data scientists for all your team members. A cloud-based SaaS, it works seamlessly with the tools you already have, integrating easily with other systems, source data platforms, data lakes and data warehouses, as well as the leading visualization tools. And fueled by the world’s fastest solver, you can tackle even the toughest analytics job. Our Solution Accelerators capability increases efficiency by reducing time to value. -
10
LlamaIndex
LlamaIndex
LlamaIndex is a “data framework” to help you build LLM apps. Connect semi-structured data from API's like Slack, Salesforce, Notion, etc. LlamaIndex is a simple, flexible data framework for connecting custom data sources to large language models. LlamaIndex provides the key tools to augment your LLM applications with data. Connect your existing data sources and data formats (API's, PDF's, documents, SQL, etc.) to use with a large language model application. Store and index your data for different use cases. Integrate with downstream vector store and database providers. LlamaIndex provides a query interface that accepts any input prompt over your data and returns a knowledge-augmented response. Connect unstructured sources such as documents, raw text files, PDF's, videos, images, etc. Easily integrate structured data sources from Excel, SQL, etc. Provides ways to structure your data (indices, graphs) so that this data can be easily used with LLMs. -
11
Cloudera Data Warehouse
Cloudera
Cloudera Data Warehouse is a cloud-native, self-service analytics solution that lets IT rapidly deliver query capabilities to BI analysts, enabling users to go from zero to query in minutes. It supports all data types, structured, semi-structured, unstructured, real-time, and batch, and scales cost-effectively from gigabytes to petabytes. It is fully integrated with streaming, data engineering, and AI services, and enforces a unified security, governance, and metadata framework across private, public, or hybrid cloud deployments. Each virtual warehouse (data warehouse or mart) is isolated and automatically configured and optimized, ensuring that workloads do not interfere with each other. Cloudera leverages open source engines such as Hive, Impala, Kudu, and Druid, along with tools like Hue and more, to handle diverse analytics, from dashboards and operational analytics to research and discovery over vast event or time-series data. -
12
SOLIXCloud CDP
Solix Technologies
SOLIXCloud CDP delivers cloud data management as-a-service for modern data-driven enterprises. Built on opensource, cloud native technologies SOLIXCloud CDP helps companies manage and process all of their structured, semi-structured and unstructured data for advanced anaytics, compliance, infrastructure optimization and data security. With features such as Solix Connect for data ingestion, Solix Data Governance, Solix Metadata Management and Solix Search, SOLIXCloud CDP offers a comprehensive cloud data management application framework to build and run data-driven applications such as SQL data warehouse, machine learning and artifitial intelligience while fulfilling the ever growing data management requirements of complex data regulations, data retention and consumer data privacy. -
13
Stambia
Stambia
In a context where data is at the heart of organizations, data integration has become a key factor in the success of digital transformation. No digital transformation without movement or transformation of data. Organizations must meet several challenges. Be able to remove the silos in the information systems. Agile and fast processing of growing data volumes and very different types of information (structured, semi-structured or unstructured data) Manage massive loads as well as ingest the data in real-time (streaming), for the most relevant decisions. Control the infrastructure costs of the data. In this context, Stambia responds by providing a unified solution for any type of data processing, which can be deployed both in the cloud and on site, and which guarantees control and optimization of the costs of ownership and transformation of the data.Starting Price: $20,000 one-time fee -
14
Pienso
Pienso
Creating a topic model from scratch takes advanced programming know-how. This expertise is expensive, and supersedes the knowledge that matters most: familiarity with your data. Labeling your own training data is slow, tedious, and costly. Farming it out to workers paid a low wage is faster and cheaper, but compromises accuracy and nuance. Either approach leaves you stuck with a fixed taxonomy that's hard to evolve. It’s time to stop tagging. Free subject matter experts to model and analyze their own data. You've got mountains of text data, filled with insights just waiting to be mined. And Pienso is here to help. Pienso is designed to train models with your own data, because we know that works best. Whether your data is unstructured or semi-structured, long or short, Pienso can help you parse it into insight. -
15
Alibaba Cloud Data Lake Formation
Alibaba Cloud
A data lake is a centralized repository used for big data and AI computing. It allows you to store structured and unstructured data at any scale. Data Lake Formation (DLF) is a key component of the cloud-native data lake framework. DLF provides an easy way to build a cloud-native data lake. It seamlessly integrates with a variety of compute engines and allows you to manage the metadata in data lakes in a centralized manner and control enterprise-class permissions. Systematically collects structured, semi-structured, and unstructured data and supports massive data storage. Uses an architecture that separates computing from storage. You can plan resources on demand at low costs. This improves data processing efficiency to meet the rapidly changing business requirements. DLF can automatically discover and collect metadata from multiple engines and manage the metadata in a centralized manner to solve the data silo issues. -
16
SOLIXCloud
Solix Technologies
Data volume keeps growing, but not all data has equal value. Cloud data management enables forward thinking companies to reduce the cost of managing enterprise data and still provide security, compliance, performance and easy access. As content ages, it loses value, but organizations can still monetize their less current data through modern SaaS-based solutions. SOLIXCloud delivers all of the capabilities required to strike the perfect balance between historical and current data management. With a complete suite of compliance features for structured, unstructured, and semi-structured data, SOLIXCloud offers a fully managed service for all enterprise data. Solix metadata management is an end-to-end framework to explore all enterprise metadata and lineage from a centralized repository and business glossary. -
17
Blox.ai
Blox.ai
Business data is usually present in different formats, across sources. A lot of business data is unstructured and semi-structured. IDP (Intelligent Document Processing) leverages AI, along with programmable automation (such as repetitive tasks), to convert data into usable, structured formats, and for consumption by downstream systems.Using Natural Language Processing (NLP), Computer Vision (CV), Optical Character Recognition (OCR) and machine learning tools, Blox.ai identifies, labels and extracts relevant data from any type of document. The AI then maps this extracted information into a structured format while configuring a model which can be applied to all similar document types. The Blox.ai stack is set up to reconcile the data based on business requirements and to push the output to downstream systems automatically.Starting Price: $650 -
18
IRI Data Protector Suite
IRI, The CoSort Company
The IRI Data Protector suite contains multiple data masking products which can be licensed standalone or in a discounted bundle to profile, classify, search, mask, and audit PII and other sensitive information in structured, semi-structured, and unstructured data sources. Apply their many masking functions consistently for referential integrity: IRI FieldShield® Structured Data Masking FieldShield classifies, finds, de-identifies, risk-scores, and audits PII in databases, flat files, JSON, etc. IRI DarkShield® Semi & Unstructured Data Masking DarkShield classifies, finds, and deletes PII in text, pdf, Parquet, C/BLOBs, MS documents, logs, NoSQL DBs, images, and faces. IRI CellShield® Excel® Data Masking CellShield finds, reports on, masks, and audits changes to PII in Excel columns and values LAN-wide or in the cloud. IRI Data Masking as a Service IRI DMaaS engineers in the US and abroad do the work of classifying, finding, masking, and risk-scoring PII for you. -
19
Azure Table Storage
Microsoft
Use Azure Table storage to store petabytes of semi-structured data and keep costs down. Unlike many data stores—on-premises or cloud-based—Table storage lets you scale up without having to manually shard your dataset. Availability also isn’t a concern: using geo-redundant storage, stored data is replicated three times within a region—and an additional three times in another region, hundreds of miles away. Table storage is excellent for flexible datasets—web app user data, address books, device information, and other metadata—and lets you build cloud applications without locking down the data model to particular schemas. Because different rows in the same table can have a different structure—for example, order information in one row, and customer information in another—you can evolve your application and table schema without taking it offline. Table storage embraces a strong consistency model. -
20
Alibaba Cloud Data Integration
Alibaba
Alibaba Cloud Data Integration is a comprehensive data synchronization platform that facilitates both real-time and offline data exchange across various data sources, networks, and locations. It supports data synchronization between more than 400 pairs of disparate data sources, including RDS databases, semi-structured storage, non-structured storage (such as audio, video, and images), NoSQL databases, and big data storage. The platform also enables real-time data reading and writing between data sources such as Oracle, MySQL, and DataHub. Data Integration allows users to schedule offline tasks by setting specific trigger times, including year, month, day, hour, and minute, simplifying the configuration of periodic incremental data extraction. It integrates seamlessly with DataWorks data modeling, providing an operations and maintenance integrated workflow. The platform leverages the computing capability of Hadoop clusters to synchronize HDFS data to MaxCompute. -
21
Etlworks
Etlworks
Etlworks is a modern, cloud-first, any-to-any data integration platform that scales with the business. It can connect to business applications, databases, and structured, semi-structured, and unstructured data of any type, shape, and size. You can create, test, and schedule very complex data integration and automation scenarios and data integration APIs in no time, right in the browser, using an intuitive drag-and-drop interface, scripting languages, and SQL. Etlworks supports real-time change data capture (CDC) from all major databases, EDI transformations, and many other fundamental data integration tasks. Most importantly, it really works as advertised.Starting Price: $300 per month -
22
Axis AI
Axis Technical Group
There’s a wide range of solutions available today for automatically extracting data from structured and semi-structured content and documents, such as databases, websites, or paper-based forms, all of which can be easily read by machines using templates or sets of predefined or custom rules. However, some businesses such as real estate, healthcare, energy, and others still rely heavily on unstructured documents. These are inconsistent in layout or form, or contain key information in English-language sentences, paragraphs, or randomly throughout the documents, making them virtually impossible for machines to understand. Axis AI offers a far better choice with a revolutionary solution for classifying and extracting information from unstructured content. Using proprietary algorithms, including those used to perform Natural Language Processing (NLP), Axis AI reads and extracts data from sentences, paragraphs, or entire pages written in natural English. -
23
IBM Streams
IBM
IBM Streams evaluates a broad range of streaming data — unstructured text, video, audio, geospatial and sensor — helping organizations spot opportunities and risks and make decisions in real-time. Make sense of your data, turning fast-moving volumes and varieties into insight with IBM® Streams. Streams evaluate a broad range of streaming data — unstructured text, video, audio, geospatial and sensor — helping organizations spot opportunities and risks as they happen. Combine Streams with other IBM Cloud Pak® for Data capabilities, built on an open, extensible architecture. Help enable data scientists to collaboratively build models to apply to stream flows, plus, analyze massive amounts of data in real-time. Acting upon your data and deriving true value is easier than ever. -
24
IBM Guardium Data Protection supports a zero trust approach to security. It discovers and classifies sensitive data from across the enterprise, providing real-time data activity monitoring and advanced user behavior analytics to help discover unusual activity around sensitive data. Guardium Data Protection is built on a scalable architecture, which provides full visibility into structured, semi-structured and unstructured data activity across all major data repositories—stored on-premises, in private and public cloud and in containers. Using a single interface, you can set access policies, monitor user access to protected data and discover, investigate and remediate vulnerabilities and threats as they occur in real time across your data environment.
-
25
DOCBrains
AGI Brains
Documents being an integral part of almost every industry, The majority of such document dominated industries are moving towards automated digital transformation. The actual pain areas are the processing structure of such complex, unstructured and semi-structured documents and Invoices. DOCBrains can automatically fetch files from various sources (Dropbox, Google Drive, Network Drive, email attachments) for you, Or upload your business documents via a secured encrypted environment into the bot. Our document processor engine best practice to ensure each relevant data gets into consideration for further processing using various ICR, OCR and AI algorithms. Document processing activity is truly fast, efficient and with 100% accuracy. Data extraction, validation and export for further processing are the three steps effectively built and implemented in the system. -
26
SISA Radar
SISA Information Security
Helping organizations improve data protection with data discovery, file analysis and classification. Secure your entire data ecosystem with SISA Radar data discovery and data classification. Organize and classify sensitive data based on the criticality and business needs. Gain contextual information to improve sensitive data management. Gain visibility into structured, semi-structured and unstructured sensitive data. Protect data from unauthorized access. Meet compliance standards of PCI DSS, GDPR, CCPA, POPIA, PDPA, APRA and other privacy regulations Create and customize your own data classification scheme. Embrace a scalable and future-proof approach to next-gen data security. A single platform to discover, identify and contextualize sensitive data. A proprietary data discovery algorithm for faster detection and lower false positives. -
27
Web Data Miner
Knowlesys Software
The Web is the largest database of public resources in the world. At present, there are at least 100 million websites with over 80 billion webpages. The number of webpages increases dramatically every single second. You can explore lots of valuable information in these webpages, including the list and contact information of potential customers, price list of competing products, real-time financial news, public opinions information, word-out-mouth information, supply and demand, scientific periodicals, forum posts, blogs and articles, and latest news. The key information, however, exists in the massive HTML webpages of websites in the form of semi-structures. As a result, the information can hardly be gathered and directly utilized. -
28
PlaNet
Optimal Solutions
Graduate from the world of Spreadsheets. Optimize your supply, manufacturing, and distribution network. Generate optimal S&OP plans in minutes. Allow us to demonstrate PlaNet’s capability on your dataset. We can give you the data templates in Excel. Manufacturers should undertake the exercise of right-sizing and right-locating their facilities periodically in order to be end-to-end cost-optimal. They should also have a robust sales & operations planning (S&OP) process in place for optimal demand fulfillment. Spreadsheet-based models cannot solve these complex problems. PlaNet allows you to leapfrog the optimization technology’s learning curve and lets you focus on your customers and business. Right-sizing and right-locating for you supply chain network for strategic capacity optimization. Manufacturing companies need to periodically reassess if its production assets are optimally sized and located, if its raw materials are sourced from the right locations. -
29
Ace Cloud Hosting
Ace Cloud Hosting
With over 15 years of experience, we're leaders in cloud-based technologies, offering Application Hosting, Managed Security Services, Public Cloud, and Hosted Virtual Desktop Solutions. Our commitment to innovation has garnered us accolades, such as the Best Outsourced Technology Provider in the CPA Practice Advisor Reader's Choice Award 2023 and the Most Innovative Cloud Solutions Provider in the Global Business Awards. We proudly serve 17,000+ customers and are trusted to tackle their toughest challenges, develop strategies, implement managed services, and modernize and secure their cloud-based applications and infrastructure. Our solutions simplify complexity, reduce costs, and ensure information is available, accessible, and adaptable anywhere, anytime, on any device. Join us as we push technology's boundaries and create a better future for all. -
30
Figma Weave
Figma
Figma Weave is a node-based creative workflow platform that brings AI models and professional editing tools into one visual workspace. The platform helps creators turn artistic ideas into scalable workflows without giving up control over quality, composition, or output. Figma Weave supports models from providers such as Google, Kling, OpenAI, Bytedance, Black Forest Labs, Runway, Luma, LTX, Wan, Grok, Recraft, and Bria. It also includes editing tools such as inpaint, outpaint, crop, masking, upscaling, depth extraction, image description, relighting, layers, type, blends, and compositing. Teams can build workflows and turn them into simplified tools so creative processes can be reused and scaled. Built for designers, artists, creative teams, and enterprises, Figma Weave combines AI generation with professional creative control in a single platform.Starting Price: $19 per month -
31
Evidently AI
Evidently AI
The open-source ML observability platform. Evaluate, test, and monitor ML models from validation to production. From tabular data to NLP and LLM. Built for data scientists and ML engineers. All you need to reliably run ML systems in production. Start with simple ad hoc checks. Scale to the complete monitoring platform. All within one tool, with consistent API and metrics. Useful, beautiful, and shareable. Get a comprehensive view of data and ML model quality to explore and debug. Takes a minute to start. Test before you ship, validate in production and run checks at every model update. Skip the manual setup by generating test conditions from a reference dataset. Monitor every aspect of your data, models, and test results. Proactively catch and resolve production model issues, ensure optimal performance, and continuously improve it.Starting Price: $500 per month -
32
Solvas Digitize
Alter Domus Data Solutions Inc.
Solvas Digitize is an intelligent document processing solution designed to help financial organizations manage complex documentation with greater accuracy and efficiency. By fully automating document intake, data extraction, validation, and reconciliation, it transforms unstructured, semi-structured, and structured documents into clean, ready-to-use information. The system centralizes every step of the workflow, allowing teams to control extraction quality, resolve missing data quickly, and eliminate manual errors. Its above-industry-average accuracy delivers reliable digitized data that supports faster, more strategic decision-making. As a managed service, Solvas Digitize combines advanced technology with expert support, reducing operational burden and eliminating the need for large capital investments. It is built to handle high-volume, high-complexity documents across investor reporting, accounting, compliance, and portfolio management use cases. -
33
Xceptor
Xceptor
Xceptor is a highly configurable, enterprise-grade data and process automation platform tailored for financial services. It automates the end-to-end journey from data ingestion, across structured, semi-structured, and unstructured formats like PDFs, emails, faxes, and forms, through intelligent AI-powered extraction, transformation, normalization, validation, enrichment, reconciliation, and workflow orchestration. It supports solutions for pre- and post-trade processing, confirmations, reconciliations, tax document tagging, client onboarding, and regulatory reporting, while maintaining governance with audit trails, exception management, real-time dashboards, role-based access, and confidence scoring. Xceptor’s low‑code engine and AI modules allow business users to configure data transformations and workflows without extensive technical expertise, enabling fast adaptation to new regulations, seamless integration with existing systems. -
34
Imperva Data Security Fabric
Imperva
Protect data at scale with an enterprise-class, multicloud, hybrid security solution for all data types. Extend data security across multicloud, hybrid, and on-premises environments. Discover and classify structured, semi-structured, & unstructured. Prioritize data risk for both incident context and additional data capabilities. Centralize data management via a single data service or dashboard. Protect against data exposure and avoid breaches. Simplify data-centric security, compliance, and governance. Unify the view and gain insights to at-risk data and users. Supervise Zero Trust posture and policy enforcement. Save time and money with automation and workflows. Support for hundreds of file shares and data repositories including public, private, datacenter and third-party cloud services. Cover both your immediate needs & future integrations as you transform and extend use cases in the cloud. -
35
Unstructured
Unstructured
80% of enterprise data exists in difficult-to-use formats like HTML, PDF, CSV, PNG, PPTX, and more. Unstructured effortlessly extracts and transforms complex data for use with every major vector database and LLM framework. Unstructured allows data scientists to pre-process data at scale so they spend less time collecting and cleaning, and more time modeling and analyzing. Our enterprise-grade connectors capture data wherever it lives, so we can transform it into AI-friendly JSON files for companies who are eager to fold AI into their business. You can count on Unstructured to deliver data that's curated, clean of artifacts, and most importantly, LLM-ready. -
36
IRI DarkShield
IRI, The CoSort Company
IRI DarkShield is a powerful data masking tool that can (simultaneously) find and anonymize Personally Identifiable Information (PII) "hidden" in semi-structured and unstructured files and database columns / collections. DarkShield jobs are configured, logged, and run from IRI Workbench or a restful RPC (web services) API to encrypt, redact, blur, etc., the PII it finds in: * NoSQL & RDBs * PDFs * Parquet * JSON, XML & CSV * Excel & Word * BMP, DICOM, GIF, JPG & TIFF DarkShield is one of 3 data masking products in the IRI Data Protector Suite, and comes with IRI Voracity data management platform subscriptions. DarkShield bridges the gap between structured and unstructured data masking, allowing users to secure data in a consistent manner across disparate silos and formats by using the same masking functions as FieldShield and CellShield EE. DarkShield also handles data in RDBs and flat-files, too, but there are more capabilities that FieldShield offers for those sources.Starting Price: $5000 -
37
QnA Maker
Microsoft
From data to bot in minutes. Build, train and publish a sophisticated bot using FAQ pages, support websites, product manuals, SharePoint documents or editorial content through an easy-to-use UI or via REST APIs. No code knowledge or coding experience required. Create and publish a bot in teams, or elsewhere without writing a single line of code. You can also add personality to your bot using pre-built chit-chat datasets. Extract question-answer pairs from semi-structured content, including FAQ pages, support websites, excel files, SharePoint documents, product manuals and policies. Design complex multi-turn conversations easily through QnA Maker portal or using REST APIs. Improve your ranking model through suggestions based on your bot's usage and user feedback. QnA Maker service is a hosted model, so you can choose product tiers according to your size and throughput needs and feel secure with all components within your Azure compliance boundary. -
38
TabFM
Google
TabFM is a zero-shot foundation model for tabular data, designed to simplify classification and regression workflows that traditionally require manual model training, hyperparameter tuning, and domain-specific feature engineering. Built specifically for tables, TabFM reframes tabular prediction as an in-context learning problem: instead of fitting a new supervised model to each dataset, it takes historical training examples and target testing rows together as one unified prompt, then interprets relationships between columns and rows at inference time. Because tables are two-dimensional and orderless, TabFM uses a hybrid architecture that combines alternating row and column attention, row compression, and a dedicated Transformer for in-context learning over compressed row embeddings. This design lets the model capture complex feature interactions and dependencies while keeping prediction computationally efficient for larger datasets.Starting Price: Free -
39
SpaceQuant
SpaceQuant
Modern property underwriting and origination platform eliminating data entry and leveraging data insight at scale. Increase loan originations by cutting loan closing times. Beat your competition with an ultra-fast quote and certainty of execution. Shorten the time to quote and close a deal by 80%. Enhance transparency and compliance, every data point is easily traceable to its location in source documents. A comprehensive system of alerts and automated reconciliations dramatically increases the accuracy and consistency of data. Mitigate risk with greater intelligence and data transparency. Get access to critical information in real time and uncover potential property performance issues well in advance. Your employees can focus on critical underwriting decisions, not data entry. SpaceQuant uses its proprietary AI technology to extract and analyze unstructured and semi-structured data from property financial documents, rent rolls, operating statements, budgets, info memorandums, etc. -
40
SiMX TextConverter
SiMX
SiMX TextConverter is a powerful and yet easy-to-use software tool for extracting and mining data from a wide variety of unstructured, semi-structured and structured data sources. It offers the best of both worlds: a flexible and intuitive visual interface for professionals with limited technical expertise, as well as, advanced functionality for professional programmers. TextConverter lets you capture, structure, transform and consolidate information from virtually any source and makes it available for business analysis via relational databases and flat files. It also includes analytical reporting capabilities for data mining and monitoring and controlling the data processing configuration process. TextConverter provides significant savings for customers across many industries including financial, insurance, healthcare, industrial and more through automation of extracting, reverse engineering and loading data from numerous text-based reports coming from disparate systems.Starting Price: $950.00/one-time -
41
IBM InfoSphere® Information Governance Catalog is a web-based tool that allows you to explore, understand and analyze information. You can create, manage and share a common business language, document and enact policies and rules, and track data lineage. Combine with IBM Watson® Knowledge Catalog to leverage existing curated data sets and extend your on-premises Information Governance Catalog investment to the cloud. A knowledge catalog allows you to put collected metadata into the hands of knowledge workers so data science and analytics communities can get easy access to the best assets for their purpose while still adhering to enterprise governance requirements. Provides a common business language and vocabulary to enable a deeper understanding of all your data assets, structured, semi-structured and unstructured. Documents governance policies and enacts rules to help you define how information should be structured, stored, transformed and moved.
-
42
Towhee
Towhee
You can use our Python API to build a prototype of your pipeline and use Towhee to automatically optimize it for production-ready environments. From images to text to 3D molecular structures, Towhee supports data transformation for nearly 20 different unstructured data modalities. We provide end-to-end pipeline optimizations, covering everything from data decoding/encoding, to model inference, making your pipeline execution 10x faster. Towhee provides out-of-the-box integration with your favorite libraries, tools, and frameworks, making development quick and easy. Towhee includes a pythonic method-chaining API for describing custom data processing pipelines. We also support schemas, making processing unstructured data as easy as handling tabular data.Starting Price: Free -
43
NanoGPT
NanoGPT
NanoGPT is private pay-per-use AI for every workflow, giving users access to chat, image, video, audio, speech, and embedding models from one platform. It is built to reduce friction for people who want access to strong models without managing many subscriptions or provider accounts, while keeping conversation history local by default and offering private options for sensitive use. NanoGPT brings together models from major providers such as ChatGPT, Claude, Gemini, DeepSeek, Llama, DALL-E, Stable Diffusion, Flux, Recraft, and more, so users can switch between tools depending on the task. It supports conversations, coding, creative writing, image generation, video generation, audio creation, text-to-speech, web search, file uploads, and model comparison in the same interface. Its model pages let users browse and discover AI language models for conversations, coding, and creative writing, as well as image models for creative projects. -
44
UBIAI
UBIAI
Leverage UBIAI's powerful labeling platform to train and deploy your custom NLP model faster than ever! When dealing with semi-structured text such as invoices or contracts, preserving document layout is key to training a high-performance model. Combining natural language processing and computer vision, UBIAI’s OCR feature allows you to perform NER, relation extraction, and classification annotation directly on native PDF documents, scanned images or pictures from your phone without losing any layout information, resulting in a significant boost of your NLP model performance. With UBIAI text annotation tool you can perform named entity recognition (NER), relation extraction and document classification all in the same interface. Unlike other tools, UBIAI enables you to create nested and overlapping entities containing multiple relations.Starting Price: $299 per month -
45
DagsHub
DagsHub
DagsHub is a collaborative platform designed for data scientists and machine learning engineers to manage and streamline their projects. It integrates code, data, experiments, and models into a unified environment, facilitating efficient project management and team collaboration. Key features include dataset management, experiment tracking, model registry, and data and model lineage, all accessible through a user-friendly interface. DagsHub supports seamless integration with popular MLOps tools, allowing users to leverage their existing workflows. By providing a centralized hub for all project components, DagsHub enhances transparency, reproducibility, and efficiency in machine learning development. DagsHub is a platform for AI and ML developers that lets you manage and collaborate on your data, models, and experiments, alongside your code. DagsHub was particularly designed for unstructured data for example text, images, audio, medical imaging, and binary files.Starting Price: $9 per month -
46
Sotero
Sotero
Sotero is the first cloud-native, zero trust data security platform that consolidates your entire security stack into one easy-to-manage environment. The Sotero data security platform employs an intelligent data security fabric that ensures your sensitive data is never left unprotected. Sotero automatically secures all your data instances and applications, regardless of source, location or lifecycle stage (at rest, in transit, or in use). With Sotero, you can move from a fragmented, complex data security stack to one unified data security fabric that provides 360° management of your entire data security ecosystem. You’re no longer forced to go to point solutions to know who is accessing your data. You get governance, auditability, visibility, and 100% control via a single pane. The Sotero platform protects any data asset wherever it resides – whether the data is a relational database, unstructured, semi-structured, structured, on-premise or in the cloud. -
47
DocVu.AI
DocVu.AI
AI and ML in DocVu.AI process loads of images into a neatly ordered set of digital documents and data. DocVu.AI seamlessly integrates into your existing systems landscape. With our deep-seated mortgage expertise and preconfigured templates, onboarding is a breeze. DocVu.AI uses the power of AI and machine learning to transform information on documents into data that machines can process. This transformation of data is for structured, semi-structured, and unstructured data. DocVu.AI can process tables, long-form text, signatures, and handwriting into digital information. DocVu.AI is much more than an Intelligent document processing engine; DocVu.AI's in-build architectural flexibility enables DocVu.AI to address unique conditions of large and small enterprises. This inherent flexibility in the process, coupled with the range of data that DocVu.AI can process accurately, has made DocVu.AI the number one choice for over 50 banks in the US. -
48
PartiQL
PartiQL
PartiQL's extensions to SQL are easy to understand, treat nested data as first class citizens and compose seamlessly with each other and SQL. This enables intuitive filtering, joining and aggregation on the combination of structured, semistructured and nested datasets. PartiQL enables unified query access across multiple data stores and data formats by separating the syntax and semantics of a query from the underlying format of the data or the data store that is being accessed. It enables users to interact with data with or without regular schema. PartiQL syntax, semantics, the embeddable reference interpreter, CLI, test framework, and tests are licensed under the Apache License, version 2.0, allowing you to freely use, copy, and distribute your changes under the terms of your choice. -
49
NLWeb
Microsoft
NLWeb is an open project developed by Microsoft that aims to make it simple to create a rich, natural language interface for websites using the model of their choice and their own data. Our goal is for NLWeb, short for Natural Language Web, to be the fastest and easiest way to effectively turn your website into an AI app, allowing users to query the contents of the site by directly using natural language, just like with an AI assistant or Copilot. Every NLWeb instance is also a Model Context Protocol (MCP) server, allowing websites to make their content discoverable and accessible to agents and other participants in the MCP ecosystem if they choose. NLWeb leverages semi-structured formats like Schema.org, RSS, and other data that websites already publish, combining them with LLM-powered tools to create natural language interfaces usable by both humans and AI agents. -
50
Amazon Bio Discovery
Amazon
Amazon Bio Discovery is an AI-powered application designed to accelerate early-stage drug discovery by combining computational biology models with real-world laboratory testing in a unified, “lab-in-the-loop” workflow. It provides scientists with direct access to a broad catalog of biological foundation models trained on large-scale biological datasets, enabling them to generate and evaluate potential drug candidates such as antibodies with greater speed and precision. Through an integrated AI agent, users can interact in natural language to select appropriate models, configure experiments, and optimize inputs without requiring advanced coding or infrastructure expertise. It allows researchers to build multi-step pipelines that combine different models, benchmark their performance, and reuse workflows across teams, improving collaboration between computational biologists and lab scientists.