Create, validate, test, and manage data contracts using the Open Data Contract Specification (ODCS) and the datacontract CLI...
This skill helps you work with data contracts following the Open Data Contract Specification (ODCS). You can execute the datacontract CLI directly to create, lint, test, and export data contracts.
datacontract init [--template PLATFORM] - Initialize a new data contractdatacontract lint {description}-CONTRACT.yaml - Validate against ODCS specdatacontract test {description}-CONTRACT.yaml - Run data quality testsdatacontract export {description}-CONTRACT.yaml --format FORMAT - Export to other formatsdatacontract lint {description}-CONTRACT.yaml - Check for best practicesdatacontract breaking {description}-CONTRACT.yaml CONTRACT_v2.yaml - Check for breaking changesTemplates available: snowflake, bigquery, redshift, databricks, postgres, s3, local
Available formats: dbt, dbt-sources, dbt-staging-sql, odcs, jsonschema, sql, sqlalchemy, avro, protobuf, great-expectations, terraform, rdf
Step 1: Gather Requirements Ask the user for:
Step 2: Initialize Contract
datacontract init --template <platform>
This creates a basic contract template for the specified platform.
Step 3: Customize the Contract Edit the generated YAML to include:
Step 4: Lint
datacontract lint {description}-contract.yaml
Always validate after creation or any modifications.
Step 5: Iterate Based on Validation If validation fails:
When validation fails, follow this pattern:
datacontract test {description}-contract.yaml
Tests validate that actual data meets the quality rules defined in the contract. This requires:
If tests fail, help the user understand:
Every data contract must include:
dataContractSpecification - Version of ODCS (e.g., "0.9.3")id - Unique identifier for the contractinfo - Metadata (title, version, description, owner, contact)servers - Data platform connection detailsschema - The data model definitionThe schema section defines your data model:
schema:
type: dbt | table | view | ...
specification: dbt | bigquery | snowflake | ...
<table_name>:
type: table
columns:
<column_name>:
type: <data_type>
required: true|false
description: "..."
unique: true|false
primary: true|false
Define quality expectations:
quality:
type: SodaCL | great-expectations | ...
specification:
checks for <table_name>:
- freshness(<column>) < 24h
- missing_count(<column>) = 0
- duplicate_count(<column>) = 0
- values in (<column>) must be in [list]
How recently data was updated:
freshness(updated_at) < 1h - Data updated within last hourfreshness(load_date) < 1d - Daily updatesEnsuring required data exists:
missing_count(customer_id) = 0 - No nulls in required fieldmissing_percent(email) < 5% - Allow some missing valuesPreventing duplicates:
duplicate_count(order_id) = 0 - Primary keys must be uniqueduplicate_count(user_id, timestamp) = 0 - Composite uniquenessData meets expected patterns:
values in (status) must be in ['pending', 'completed', 'cancelled']invalid_percent(email) < 1% - Email format validationmin(price) >= 0 - Range constraintsmax_length(postal_code) = 5 - Length constraintsGenerate platform-specific artifacts:
datacontract export {description}-contract.yaml --format dbt
datacontract export {description}-contract.yaml --format sql
datacontract export {description}-contract.yaml --format great-expectations
Common use cases:
User: "Create a data contract for our customer orders table in Snowflake"
You should:
datacontract init --template snowflakedatacontract lint {description}-contract.yamlWhen you need more details:
Remember: The datacontract CLI is your primary tool. Execute commands directly, parse output carefully, and guide the user through the process with clear explanations.