Data Quality Frameworks: Great Expectations, Soda and dbt Tests
Why Data Quality Frameworks Matter
Bad data costs more than missing data. A null value in a revenue column triggers an obvious error. A subtly wrong exchange rate multiplied across millions of transactions creates a report that looks correct but drives wrong decisions. Data quality frameworks exist to catch both cases before they reach stakeholders.
Three frameworks dominate production data stacks today: Great Expectations, Soda Core, and dbt tests. Each takes a fundamentally different approach to the same problem. Great Expectations treats validation as a first-class engineering discipline with extensive configuration. Soda Core prioritizes simplicity and monitoring. dbt tests embed quality checks directly into the transformation layer. Choosing between them—or combining them—depends on where your pipelines break most often.
The decision interacts directly with your data contracts strategy. Contracts define what producers promise. Quality frameworks verify that those promises hold in production.
Great Expectations: Programmable Validation
Great Expectations models data quality as a collection of expectations—individual assertions about a dataset. An expectation might state that a column contains no nulls, that values fall within a range, or that a table has between 900,000 and 1,100,000 rows. Expectations compose into suites, and suites attach to datasources through checkpoints.
The framework's strength lies in its execution model. Rather than pulling data into Python, Great Expectations generates native queries against the execution engine. When connected to BigQuery, an expectation translates to SQL that runs inside BigQuery's distributed engine. Connected to Spark, expectations become Spark operations. This pushdown design means validation scales with your compute, not with your Python process.
# Great Expectations checkpoint configuration
checkpoint:
name: orders_daily
validations:
- batch_request:
datasource_name: warehouse
data_asset_name: orders
expectation_suite_name: orders_quality
action_list:
- name: store_result
action:
class_name: StoreValidationResultAction
- name: slack_alert
action:
class_name: SlackNotificationAction
notify_on: failure
Data Docs—auto-generated HTML documentation—provide a browsable record of every validation run. This matters more than most teams initially expect. When a stakeholder questions a number, you can point them to a timestamped validation result showing every check that passed before the data was promoted to production. This audit trail aligns with data governance requirements that enterprise organizations face.
The trade-off is complexity. Great Expectations requires configuration for datasources, expectation suites, checkpoints, and stores. A minimal production setup involves at least four YAML files and a Python initialization script. Teams with fewer than five data engineers often find the operational overhead hard to justify when dbt tests cover 80% of their validation needs.
Soda Core: Configuration-First Quality
Soda takes the opposite approach to complexity. Where Great Expectations requires Python code and multiple configuration layers, Soda defines checks in a single YAML file. A scan runs all checks against a datasource and produces pass/fail results with optional alerting.
# Soda check definition - checks.yml
checks for orders:
- row_count > 0
- missing_count(customer_id) = 0
- invalid_percent(email) < 5%:
valid format: email
- duplicate_count(order_id) = 0
- freshness(created_at) < 2h
- schema:
fail:
when required column missing:
[order_id, customer_id, total_amount, created_at]
when wrong type:
order_id: integer
total_amount: decimal
The freshness check deserves special attention. It verifies that the most recent row in a table is within a specified time window. This single check catches an entire category of pipeline failures—silent stops where no error occurs but data simply stops arriving. Teams running Airflow or Prefect pipelines benefit from freshness checks as a second line of defense when DAG monitoring misses partial failures.
Soda Cloud extends the open-source Core with a managed UI, anomaly detection, and incident tracking. The anomaly detection uses historical distributions to flag statistical outliers without explicit threshold configuration. When your orders table usually receives between 40,000 and 55,000 rows per day and suddenly gets 12,000, Soda Cloud raises an alert even if no hard threshold was set.
The limitation is extensibility. Custom checks in Soda require writing Python plugins following a specific interface. Great Expectations makes custom expectations straightforward with decorators and a well-documented API. For teams that need domain-specific validations beyond standard statistical checks, this extensibility gap matters.
dbt Tests: Quality at the Transformation Layer
dbt tests validate data as part of the transformation process. Four generic tests ship with every dbt project: unique, not_null, accepted_values, and relationships. These cover the most common data quality assertions and require only YAML configuration in schema files.
# dbt schema.yml
models:
- name: orders
columns:
- name: order_id
tests:
- unique
- not_null
- name: status
tests:
- accepted_values:
values: ['pending', 'shipped', 'delivered', 'cancelled']
- name: customer_id
tests:
- relationships:
to: ref('customers')
field: customer_id
The dbt-expectations package extends this with 50+ tests inspired by Great Expectations. Statistical tests for distribution shape, row count ranges, and column pair correlations bring sophisticated validation into the dbt workflow. The dbt-utils package adds expression-based tests and cross-database assertions.
Custom SQL tests handle business logic that generic tests cannot express. A test is simply a SQL query that returns rows which violate the assertion. Zero rows means the test passes. This design makes it possible to encode arbitrarily complex business rules.
-- tests/assert_order_total_matches_line_items.sql
SELECT o.order_id, o.total_amount, SUM(li.unit_price * li.quantity) AS calc_total
FROM {{ ref('orders') }} o
JOIN {{ ref('line_items') }} li ON o.order_id = li.order_id
GROUP BY o.order_id, o.total_amount
HAVING ABS(o.total_amount - SUM(li.unit_price * li.quantity)) > 0.01
The tight integration with the dbt transformation workflow is both the strength and the limitation. Tests run in the same context as models, using the same warehouse connection and the same scheduling infrastructure. But dbt tests can only validate data inside the warehouse. Source system validation, API response validation, and pre-load checks fall outside dbt's scope.
Framework Comparison Matrix
| Capability | Great Expectations | Soda Core | dbt Tests |
|---|---|---|---|
| Configuration language | Python + YAML | YAML only | YAML + SQL |
| Built-in expectations/checks | 300+ | 25+ | 4 (50+ with packages) |
| Source system validation | Yes | Yes | No |
| Query pushdown | Yes (SQL/Spark) | Yes (SQL) | Yes (SQL) |
| Schema drift detection | Via expectations | Built-in | Via dbt-utils |
| Freshness monitoring | Custom expectation | Built-in | source freshness |
| Anomaly detection | Custom/profiling | Soda Cloud | dbt-expectations |
| Data documentation | Data Docs (HTML) | Soda Cloud UI | dbt Docs |
| CI/CD integration | Checkpoint CLI | soda scan CLI | dbt test CLI |
| Learning curve | Steep | Moderate | Low (for dbt users) |
| Best for | Complex validation | Monitoring | Transformation QA |
Production Deployment Patterns
Pattern 1: dbt-First Quality Stack
Teams with a dbt-centric stack often start with built-in tests and gradually add dbt-expectations. This covers transformation-layer validation completely and integrates natively with incremental model workflows. Source freshness checks in dbt handle the pre-transformation monitoring gap. This pattern works well for teams with fewer than 20 models and predictable source systems.
Pattern 2: Soda for Monitoring, dbt for Validation
Separating monitoring from validation gives each tool its natural role. Soda runs on a schedule independent of dbt, checking source systems for freshness, schema changes, and anomalies. dbt tests run as part of the transformation DAG, validating business logic and referential integrity. Soda catches problems before dbt starts. dbt catches problems during transformation. Together they provide end-to-end coverage that is particularly effective for pipelines ingesting data from CDC sources via Debezium.
Pattern 3: Great Expectations for Data Contracts
Organizations with formal data contracts between producer and consumer teams need the programmatic power of Great Expectations. Expectation suites serve as machine-readable contract definitions. When a producer team changes a schema or data distribution, the consumer's expectation suite fails before the change propagates downstream. This pattern pairs well with data mesh architectures where domain teams own their data products independently.
Integration with Pipeline Orchestrators
All three frameworks integrate with Airflow DAGs through operators or bash commands. Great Expectations provides a dedicated Airflow operator that runs checkpoints and reports results as task metadata. Soda exposes a CLI (soda scan) that returns non-zero exit codes on failure, making it compatible with any orchestrator's bash operator. dbt tests run as part of dbt build or dbt test, both of which fail on test failures.
The orchestration pattern matters. Running quality checks as separate DAG tasks enables conditional branching—quarantining failed data rather than halting the entire pipeline. This quarantine-and-continue pattern keeps downstream models fresh with validated data while flagging problematic batches for review. Teams working with event-driven architectures benefit from this approach since event streams cannot simply pause and retry.
Migration Between Frameworks
Moving from one framework to another is less painful than most teams expect. The core mapping is straightforward: a Great Expectations expect_column_values_to_not_be_null maps directly to dbt's not_null test and Soda's missing_count = 0. The semantic differences are small. The operational differences—configuration format, deployment model, alerting integration—are where migration effort concentrates.
Start by inventorying existing checks. Map each to its equivalent in the target framework. Implement the mapped checks and run both systems in parallel for two weeks. Compare results daily. Discrepancies reveal edge cases in how each framework handles data types, null semantics, and boundary conditions.
This parallel-run approach also helps teams validate Great Expectations production configurations before committing to a full migration. The cost is double compute for validation queries during the parallel period.
Choosing the Right Framework
For teams starting from zero, dbt tests offer the lowest barrier to entry. If your data already flows through dbt, adding tests requires no new infrastructure, no additional dependencies, and no new deployment pipeline. Start here and grow into Soda or Great Expectations when the limitations become concrete rather than theoretical.
For monitoring-first teams that need visibility into source systems before data enters the warehouse, Soda provides the fastest path to coverage. Its YAML-based configuration scales linearly with the number of data sources, and the freshness monitoring alone prevents an entire class of silent failures.
For organizations with formal data contracts, regulated data, or complex multi-team data platforms, Great Expectations provides the programmability and audit trail that simpler tools lack. The investment in configuration pays dividends when stakeholders require traceable proof that data met quality standards at every step of the pipeline.
Whatever you choose, the decision should be informed by your data lineage infrastructure and the observability capabilities already in place. Quality frameworks do not operate in isolation—they are one layer in a stack that includes lineage tracking, anomaly detection, and schema drift management.
Key Takeaways
Data quality is not a single-tool problem. The three frameworks cover different phases of the data lifecycle, and the best production deployments combine them based on where failures actually occur rather than theoretical coverage goals. Start with the tool closest to your existing stack, prove its value on three to five critical tables, and expand coverage based on incidents rather than checklists.
The framework you choose matters less than the consistency with which you apply it. A simple dbt test suite covering every model in your project catches more issues than a sophisticated Great Expectations deployment covering only the models someone remembered to configure. Coverage and consistency beat sophistication every time.