Data Catalog and Metadata Management at Scale

By Priya Raghavan • • 9 min read

Every organization that has grown past a handful of databases faces the same problem: nobody can find the data they need, nobody knows whether they can trust it, and nobody understands how it flows from source to report. Data catalogs solve this by creating a searchable, annotated inventory of every dataset, pipeline, and metric across the platform. But implementing a catalog at scale is more than deploying a tool — it requires deliberate metadata modeling, automated ingestion, lineage tracking, and organizational buy-in.

This guide covers the architecture of modern data catalogs, compares leading open-source platforms, walks through practical implementation patterns, and shares lessons from scaling metadata management across hundreds of data sources. For related organizational strategies, see our article on Data Mesh organizational patterns.

Why Metadata Management Matters

Metadata is often described as "data about data," but that undersells its importance. In practice, metadata is the connective tissue that makes a data platform navigable. Without it, analysts spend an estimated 30-40% of their time searching for data rather than analyzing it. Data engineers duplicate pipelines because they cannot discover existing ones. Compliance teams cannot answer auditors' questions about data lineage or access patterns.

The costs compound as organizations scale. A company with 50 data sources and 500 tables might manage with tribal knowledge and a shared spreadsheet. At 500 sources and 50,000 tables, that approach collapses. Metadata management at this scale requires automation, standardization, and tooling that treats metadata as a first-class product.

A data catalog is not a project you finish. It is an ongoing capability that grows with your data platform, continuously ingesting new metadata and maintaining the accuracy of what it already knows.

The Four Pillars of Metadata

Effective metadata management addresses four distinct categories:

  • Technical metadata — Schemas, column types, partitioning strategies, storage formats, and physical locations. This is the foundation that automated crawlers extract.
  • Business metadata — Human-readable descriptions, business glossary terms, domain ownership, and use case documentation. This requires human curation.
  • Operational metadata — Pipeline run histories, freshness timestamps, row counts, data quality scores, and SLA compliance. This comes from orchestration and monitoring systems.
  • Social metadata — Usage statistics, popular queries, user endorsements, questions, and tribal knowledge captured in comments. This emerges from catalog adoption.

Open-Source Catalog Landscape

The open-source data catalog space has matured significantly. Three platforms dominate production deployments, each with distinct architectural philosophies.

LinkedIn DataHub

DataHub uses a metadata graph built on a generalized entity-relationship model. Every metadata entity — datasets, dashboards, pipelines, users, tags — is a node in the graph, connected by typed relationships. The ingestion framework supports push-based (via REST or Kafka) and pull-based (scheduled crawlers) metadata collection.

DataHub's architecture separates the metadata store (MySQL or PostgreSQL), search index (Elasticsearch), and graph index (Neo4j or its built-in graph) into independently scalable components. This makes it suitable for organizations with tens of thousands of datasets.

Amundsen

Originally built at Lyft, Amundsen focuses on data discovery with a search-first user experience. It uses Neo4j as its primary metadata store, which naturally represents relationships like ownership, lineage, and tagging. The frontend emphasizes table-level detail pages with integrated usage statistics from query log analysis.

Amundsen's databuilder framework provides ETL-style extractors for common data sources. It is simpler to deploy than DataHub but offers fewer extensibility points for custom metadata types.

OpenMetadata

OpenMetadata takes a schema-first approach, defining all metadata entities through JSON Schema. This makes the metadata model explicit, versionable, and extensible without code changes. It provides a built-in data quality framework and integrates with dbt for test result ingestion.

FeatureDataHubAmundsenOpenMetadata
Metadata ModelGeneralized entitiesFixed typesJSON Schema
Search BackendElasticsearchElasticsearchElasticsearch
Graph StoreNeo4j / built-inNeo4jMySQL
LineageColumn-levelTable-levelColumn-level
Data QualityVia integrationsLimitedBuilt-in
API-FirstYes (REST + GraphQL)REST onlyYes (REST + SDK)
DeploymentDocker / KubernetesDocker / KubernetesDocker / Kubernetes

Metadata Ingestion Architecture

The most critical engineering challenge in metadata management is ingestion: extracting metadata from diverse sources and loading it into the catalog in a normalized, consistent format. The ingestion layer must handle schema changes gracefully, operate incrementally to avoid full-scan overhead, and preserve provenance so administrators know which crawler or API call produced each metadata record.

Push vs. Pull Ingestion

Pull-based ingestion uses scheduled crawlers that connect to data sources, extract schema information, and push it to the catalog. This works well for databases and warehouses that expose metadata through information schemas or system catalogs. A typical DataHub ingestion recipe for a PostgreSQL source looks like this:

# datahub_postgres_recipe.yaml
source:
  type: postgres
  config:
    host_port: "analytics-db.internal:5432"
    database: warehouse
    username: "${DATAHUB_PG_USER}"
    password: "${DATAHUB_PG_PASSWORD}"
    include_tables: true
    include_views: true
    profiling:
      enabled: true
      profile_table_level_only: false
    lineage_via_query_log:
      enabled: true
      lookback_days: 30

sink:
  type: datahub-rest
  config:
    server: "http://datahub-gms:8080"

pipeline_name: postgres_warehouse_ingestion
schedule: "0 */6 * * *"

Push-based ingestion embeds metadata emission directly in data pipelines. When an Airflow DAG processes a table, it publishes metadata events to a Kafka topic that the catalog consumes. This provides real-time metadata freshness but requires instrumenting every pipeline:

from datahub.emitter.kafka_emitter import DatahubKafkaEmitter
from datahub.metadata.schema_classes import (
    DatasetPropertiesClass,
    MetadataChangeProposalWrapper,
)

emitter = DatahubKafkaEmitter(
    config={"connection": {"bootstrap": "kafka:9092"}}
)

# Emit metadata after pipeline completes
dataset_urn = "urn:li:dataset:(urn:li:dataPlatform:snowflake,analytics.orders,PROD)"
properties = DatasetPropertiesClass(
    description="Cleaned order records joined with customer data",
    customProperties={
        "pipeline": "orders_etl_v3",
        "last_row_count": "14523897",
        "freshness": "2026-10-01T08:30:00Z",
    },
)

event = MetadataChangeProposalWrapper(
    entityUrn=dataset_urn,
    aspect=properties,
)
emitter.emit(event)

The best approach combines both: pull-based crawlers provide baseline coverage with periodic full scans, while push-based emitters keep high-priority datasets current in real time.

Data Lineage Tracking

Lineage — the record of how data flows from source to destination through transformations — is the most valuable and most difficult metadata to capture. Without lineage, an analyst looking at a revenue dashboard cannot trace the numbers back to the source tables, cannot identify which ETL job produced the aggregation, and cannot assess whether upstream data quality issues affect the report.

Column-Level Lineage

Table-level lineage shows that table B depends on table A. Column-level lineage shows that B.total_revenue is derived from A.unit_price * A.quantity. The latter is dramatically more useful for impact analysis but harder to extract.

SQL-based lineage extraction parses query logs to identify source-destination mappings. Tools like sqllineage and DataHub's SQL parser can handle standard SQL dialects, but struggle with dynamic SQL, stored procedures, and non-SQL transformations like Python-based Spark jobs.

For comprehensive lineage in complex environments, OpenLineage provides a standardized event format that orchestrators and processing engines can emit natively. Apache Airflow, Spark, and dbt all have OpenLineage integrations that emit lineage events as jobs execute:

# OpenLineage event structure (emitted by Spark)
{
  "eventType": "COMPLETE",
  "eventTime": "2026-10-01T08:30:00Z",
  "run": {"runId": "a1b2c3d4-e5f6-7890-abcd-ef1234567890"},
  "job": {
    "namespace": "spark://analytics-cluster",
    "name": "orders_daily_aggregate"
  },
  "inputs": [
    {
      "namespace": "postgres://analytics-db",
      "name": "raw.orders",
      "facets": {
        "columnLineage": {
          "fields": {
            "order_date": {"inputFields": [{"namespace": "postgres://analytics-db", "name": "raw.orders", "field": "created_at"}]},
            "revenue": {"inputFields": [
              {"namespace": "postgres://analytics-db", "name": "raw.orders", "field": "unit_price"},
              {"namespace": "postgres://analytics-db", "name": "raw.orders", "field": "quantity"}
            ]}
          }
        }
      }
    }
  ],
  "outputs": [
    {"namespace": "snowflake://warehouse", "name": "analytics.daily_orders"}
  ]
}

Governance and Access Control

A data catalog becomes a governance tool when it can enforce policies, not just document them. Modern catalogs integrate with access control systems to implement tag-based policies: a dataset tagged pii automatically requires elevated permissions, and the catalog can show users what they have access to before they waste time querying something they cannot reach.

Automated Classification

Manual tagging does not scale. Automated classifiers scan column names and sample values to detect sensitive data patterns. A column named ssn containing nine-digit numbers gets auto-tagged as PII. A column with values matching email regex patterns gets tagged as personal_contact. These classifiers run during metadata ingestion and propose tags for human review.

The classification pipeline typically operates in three stages:

  1. Pattern matching — Regular expressions on column names and sampled values catch the obvious cases: emails, phone numbers, IP addresses, credit card patterns.
  2. Statistical profiling — Cardinality analysis, value distributions, and null rates identify columns that behave like identifiers or sensitive fields even when naming is non-standard.
  3. Propagation — Lineage-aware classification propagates tags downstream. If source.email is tagged PII, and derived.contact_hash is computed from it, the derived column inherits a derived_from_pii tag automatically.

Ownership and Accountability

Every dataset in the catalog needs an owner. Ownership is not about blame — it is about having a person who can answer questions, approve access requests, and prioritize quality improvements. The data mesh model formalizes this through domain-level data product ownership, where each domain team owns the metadata, quality, and documentation of their published datasets.

Effective catalogs automate ownership assignment by parsing git history (who last modified the pipeline that produces this table?), query logs (who queries this table most frequently?), and organizational structure (which team owns the source system?).

Scaling to Thousands of Sources

At large scale, catalog performance and operational complexity become primary concerns. Search latency must stay under 200ms for the catalog to feel responsive. Lineage graph traversal for impact analysis must complete in seconds, not minutes. Metadata ingestion must handle thousands of sources without overwhelming the catalog's backend.

Ingestion Parallelism and Rate Limiting

Running hundreds of crawlers simultaneously can overwhelm both source systems and the catalog API. A well-designed ingestion orchestrator staggers crawlers by priority tier, rate-limits API calls per source, and implements circuit breakers that pause ingestion when the catalog's indexing queue grows too deep. DataHub's ingestion framework supports this through configurable concurrency limits and backpressure signals from the GMS (Generalized Metadata Service) layer.

Search Index Optimization

Elasticsearch is the search backend for all major catalogs, and its configuration directly impacts discovery quality. Key optimizations include:

  • Custom analyzers — Technical terms like user_id, clickstream_events, and etl_daily need snake_case-aware tokenization that splits on underscores while preserving the full term for exact matches.
  • Boosted fields — Table names and descriptions should rank higher than column names in search results. Business glossary term matches should outrank technical schema matches.
  • Usage-weighted ranking — A table queried 500 times per day should appear above an identically named table queried once per month. Integrating query log statistics into search ranking dramatically improves result relevance.
  • Faceted navigation — Platform, domain, owner, tags, and freshness facets let users narrow results efficiently without crafting complex search queries.

Measuring Catalog Adoption and Impact

Deploying a catalog is meaningless if nobody uses it. Tracking adoption metrics guides investment and surfaces gaps in coverage. A catalog integrated with data observability tooling can correlate catalog usage with data quality improvements and analyst productivity.

Key Metrics to Track

Effective catalog programs monitor several categories of metrics:

  • Coverage — Percentage of datasets in the catalog with descriptions, owners, and tags. Target: 80% documented within 6 months of source onboarding.
  • Freshness — How recently was each dataset's metadata updated? Stale metadata erodes trust. Target: technical metadata refreshed within 24 hours of schema changes.
  • Search success rate — Percentage of searches that result in a user navigating to a dataset detail page. A rate below 40% indicates poor metadata quality or search tuning.
  • Time to data — How long does it take a new analyst to find and start using a dataset? Before-and-after measurements demonstrate catalog ROI to leadership.
  • Lineage completeness — Percentage of datasets with traced lineage back to source systems. Gaps in lineage indicate unmonitored pipelines or unsupported transformation types.

Catalog adoption follows a predictable curve. Early adopters are the data engineering team that built it. The second wave is analytics engineers who use it for impact analysis when modifying dbt models. Broad adoption across business analysts and data scientists comes last and requires active evangelism, training sessions, and integration with the tools those users already use — embedding catalog links in BI dashboards, adding catalog search to Slack bots, and surfacing catalog metadata in notebook environments.

The catalog is not finished when metadata is loaded. It is finished when every person who works with data reaches for the catalog first, because it consistently gives them faster, more reliable answers than asking a colleague or searching Slack history. That is the standard worth engineering toward.