Blog

What Data Governance Got Right That AI Governance Proposals Are Missing

featured_image
AI Governance Business

What Data Governance Got Right That AI Governance Proposals Are Missing

How a decade of data governance discipline can close the accountability gaps in today’s AI regulation, and what it would take to actually apply it.

Introduction


Most AI governance debates behave as though the problem of governing complex, high-stakes information systems is brand new. It isn’t. For the better part of two decades, data governance teams inside banks, hospitals, and regulated enterprises have been solving a structurally similar problem: how do you make a system that touches sensitive information traceableaccountable, and auditable, without freezing the organization’s ability to actually use that information?

The uncomfortable truth is that many current AI governance proposals, including flagship regulation like the EU AI Act and voluntary frameworks like the NIST AI Risk Management Framework, gesture at the outcomes data governance already knows how to produce (traceability, accountability, auditability) without adopting the mechanisms that reliably produce them. This piece walks through three of those mechanisms, shows where their AI-governance equivalents currently fall short, and proposes what closing that gap would actually require.

The pattern above is consistent across three of the most mature disciplines in enterprise data governance. Each one solved a version of the accountability problem that AI governance is now re-litigating from scratch.

1. Data Lineage Solved Traceability. AI Governance Is Still Debating Whether Training Data Needs a Paper Trail.

Data lineage, the practice of recording where a piece of data came from, what transformed it, and where it flowed, is table stakes in any mature data governance program. Regulators like GDPR examiners and financial auditors simply do not accept “we don’t know where that number came from” as an answer.

AI training data has no equivalent standard. Foundation models are frequently trained on datasets whose composition, licensing status, and consent provenance are undocumented or unverifiable after the fact. The research community has proposed fixes, most notably Datasheets for Datasets, introduced by Timnit Gebru and colleagues, which called for every dataset to ship with a structured account of its motivation, composition, and collection process, modeled explicitly on the datasheets that accompany electronic components. Margaret Mitchell and colleagues extended the same logic to trained models with Model Cards, documenting intended use, evaluation conditions, and performance across demographic subgroups.

These proposals are excellent, and they remain almost entirely voluntary. Unlike data lineage in a regulated enterprise, where a missing audit trail is a compliance finding, a missing datasheet or model card carries no consequence beyond reputational pressure. Data governance’s real innovation wasn’t the concept of lineage. It was making lineage mandatory and enforced through audit regimes with teeth.

2. Access Control Solved Who Can Touch What. AI Governance Has Barely Started Asking Who Can Touch a Deployed Model.

Role-based access control and the principle of least privilege are foundational to any data governance framework: not everyone in an organization gets access to every dataset, and access is logged, time-bound, and reviewable.

Compare that to how loosely controlled access to model weights, fine-tuning pipelines, and inference endpoints often is once a model is deployed. Which internal teams can fine-tune a production model? Who can adjust a system prompt that changes safety behavior in a customer-facing tool? In many organizations, the honest answer is “we don’t have a clean record of that,” a sentence that would fail a data governance audit outright.

3. Audit Trails Solved Post-Hoc Accountability. AI Governance Still Struggles to Explain Why a Model Behaved a Certain Way, Six Months Later.

Immutable audit trails let a data governance team reconstruct, months after the fact, exactly what happened to a piece of data and who touched it. AI systems, by contrast, frequently drift in behavior after deployment, through updates, fine-tuning, prompt changes, or shifts in the underlying serving infrastructure, without a corresponding log tying a specific behavior back to a specific model version and context.

The NIST AI Risk Management Framework explicitly names this as a structural risk. NIST’s guidance flags growing concern over copyright and data-provenance uncertainty, since training datasets often contain material with unclear licensing or origins, and the framework more broadly treats documentation of data provenance and decision rationale as central to accountability. But the AI RMF, like most of its peers, is voluntary. It tells organizations what good governance looks like without data governance’s second ingredient: an enforcement mechanism that makes non-compliance costly.


The Regulatory Case Study: What the EU AI Act Gets Right, and Where It’s Already Slipping

The EU AI Act is the most ambitious attempt yet to convert AI governance principles into binding law, and its structure borrows more from data governance than most commentary acknowledges. High-risk AI systems are required to implement documented risk management, “robust data governance measures,” full technical documentation, automatic logging, and human oversight before they can reach the market, a regime with real structural similarity to how regulated data governance programs operate.

But the Act’s own timeline is a live illustration of the gap between writing accountability requirements and operationalizing them. The high-risk obligations under Annex III were originally due to take effect on 2 August 2026. In May 2026, EU legislators agreed to defer those obligations, via the “Digital Omnibus” package, specifically because implementation had fallen behind schedule. National competent authorities had not been designated in most member states, and the harmonised standards and compliance tools needed to actually run conformity assessments were not yet finalized.

As of this writing, the revised dates are 2 December 2027 for stand-alone Annex III systems and 2 August 2028 for AI embedded in regulated products under Annex I, a deferral of well over a year for the regime’s most consequential provisions. Importantly, the transparency obligations under Article 50 and the AI literacy duty under Article 4 remain on their original schedule, so the delay is real but narrower than headlines suggest.

This is precisely the kind of implementation gap that a data-governance-trained eye recognizes immediately: the policy language was strong; the operational infrastructure to enforce it (standards bodies, notified bodies, conformity assessment capacity) was not ready. Data governance programs learn this lesson early, usually the hard way: a control that exists only on paper, without a funded enforcement function behind it, is not a control.

What This Comparison Suggests as a Concrete Next Step

If AI governance is going to close this gap rather than keep rediscovering it, the field needs something functionally equivalent to a data lineage record, but for model behavior. Call it a model behavior lineage record. At minimum, it would need to:

ComponentData Governance PrecedentAI Governance Application
Origin recordData lineage / source system loggingTraining data provenance tied to licensing and consent status
Change logChange data capture, versioned schemasEvery fine-tune, system prompt change, or safety-relevant update logged against model version
Access logRole-based access control audit logsWho can modify weights, prompts, or safety configurations, and when
Classification tierData sensitivity classification (public/internal/restricted)Risk-tier classification matching NIST RMF / EU AI Act risk categories
Enforcement mechanismRegulatory audit with real penalties (e.g., GDPR fines)Independent conformity assessment with binding, funded enforcement capacity

The first four rows are largely solved problems; they’re what data governance teams already build. The fifth row is the one AI governance keeps deferring, quite literally, as the EU’s own 2026 timeline shows.

The Takeaway

AI governance doesn’t need to invent the concept of accountability from first principles. It needs to import the operational discipline that data governance has already spent a decade refining, and then do the harder, less glamorous work of funding and staffing the enforcement infrastructure that makes accountability real rather than aspirational. That is a data governance problem wearing an AI governance label, and it’s exactly the kind of problem that professionals with a data governance background are positioned to help solve.


Sources

  1. Gebru, T., et al. Datasheets for Datasetsresearchgate.net/publication/324055506
  2. Mitchell, M., et al. (2019). Model Cards for Model Reporting, FAT* ’19. arxiv.org/pdf/1810.03993
  3. NIST. AI Risk Management Framework (AI RMF 1.0), and related generative AI risk guidance. digital.nemko.com/regulations/nist-rmf
  4. DLA Piper GENIE. The Digital AI Omnibus: Proposed deferral of high-risk AI obligations under the AI Actknowledge.dlapiper.com
  5. Gibson Dunn. EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines and Other Key Changesgibsondunn.com
  6. artificialintelligenceact.eu. Article 6: Classification Rules for High-Risk AI Systemsartificialintelligenceact.eu/article/6

This article is part of an ongoing series exploring the intersection of data governance and AI governance as I transition my focus toward AI safety and policy work. Read more in the AI Governance & Safety section.

Leave your thought here

Your email address will not be published. Required fields are marked *

Select the fields to be shown. Others will be hidden. Drag and drop to rearrange the order.
  • Image
  • SKU
  • Rating
  • Price
  • Stock
  • Availability
  • Add to cart
  • Description
  • Content
  • Weight
  • Dimensions
  • Additional information
Click outside to hide the comparison bar
Compare