DataHub is an open source metadata platform that helps organizations discover, understand, and trust their data assets at scale. It models data as a richly connected graph spanning datasets, dashboards, pipelines, ML features, and services, so users can explore relationships like lineage and ownership across tools and domains. The platform focuses on continuous metadata ingestion from many sources, treating metadata as a stream that stays fresh as systems change. A modern web UI and search layer make it easy for analysts, engineers, and stewards to find assets, read documentation, and see context such as usage, schema history, and downstream impacts. Governance features such as glossaries, tags, ownership, and policies bring order and accountability without blocking day-to-day work. APIs and SDKs enable teams to integrate DataHub into catalogs, CI pipelines, and automation so metadata becomes part of the development lifecycle rather than an afterthought.
Features
- Pluggable metadata ingestion for databases, warehouses, BI tools, ML platforms, and more
- Rich data lineage and impact analysis across datasets, jobs, and dashboards
- Powerful search and filtering with entity pages showing schema, docs, and usage
- Business glossary, tags, and ownership to drive governance and discoverability
- Versioned metadata with history, diffs, and programmatic APIs for automation
- Role-based access and policies for safe collaboration at enterprise scale