Docker image used to run data processing workloads
A free, open-source, and cross-platform big data analytics framework
Simple and distributed Machine Learning
R interface for Apache Spark
Scalable and Flexible Gradient Boosting
Python Stream Processing
Monitor the stability of a Pandas or Spark dataframe
Scalable master data management and identity resolution
Apache Polaris, the interoperable, open source catalog
Series (one-dimensional) and dataframes (two-dimensional)
A Scala API for Apache Beam and Google Cloud Dataflow
Apache IoTDB
Distributed Big Data Orchestration Service
A distributed and extensible workflow scheduler platform
Apache Spark to Apache Cassandra connector
A graph database that supports more than 100+ billion data
A portable SCADA/IoT platform centered on the MongoDB database server.
A scalable, unified data and AI engineering platform for enterprise
Memory optimized analytics database, based on Apache Spark
Julia binding for Apache Spark
Enterprise-strength marketing and product analytics platform
Use SQL to build ELT pipelines on a data lakehouse
Harmonious distributed data analysis in Rust
SZT‑bigdata is an open source project
World's first open source data quality & data preparation project