Data governance is the operating model that decides who owns each data asset, what it means, where it came from, and who may use it. In Azerbaijan it is driven less by internal appetite than by three external forces: Central Bank supervision of financial institutions, the Law on Personal Data, and the practical impossibility of deploying AI on data nobody can vouch for.

That last force is what changed. For a decade, data governance in this market was a compliance exercise that produced a spreadsheet and a steering committee. Then organisations started putting language models on top of their data, discovered that nobody could say with confidence which of four customer tables was authoritative, and governance stopped being optional.

The numbers behind that shift are now unambiguous. McKinsey's State of AI research for 2026 finds 88% of organisations regularly using AI in at least one business function, but only 39% reporting any EBIT impact at enterprise level, and roughly two-thirds not yet scaling AI beyond pilots. Gartner, meanwhile, predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. The gap between adoption and impact is, in most institutions, a data governance gap.

This guide is written for the people who actually have to run such a programme in an Azerbaijani bank, ministry or large enterprise. It assumes you have legacy systems in three languages, a regulator who asks direct questions, and a budget that has to show something within a year.

What data governance actually is — and what it is not

Data governance is an operating model. It is not a tool, a department or a policy document, though it produces all three.

Concretely, a working programme answers four questions about any piece of data in the organisation:

  1. Who owns it? A named person, not a department. Someone who can be asked "is this correct?" and is accountable for the answer.
  2. What does it mean? One definition of active customer that risk, marketing and regulatory reporting all use. Not three.
  3. Where did it come from? The full path from source system to the number in the report, column by column.
  4. Who may use it, and for what? Access rules tied to data sensitivity, not to whichever ticket someone raised in 2019.

If your programme cannot answer those four questions for your top fifty data assets, it does not yet exist, regardless of how many policies have been signed.

What it is not: a data warehouse project, a BI migration, or a catalog purchase. A catalog is essential infrastructure for governance the way an issue tracker is essential infrastructure for engineering — necessary, and worth nothing on its own.

What poor governance costs

The most-cited figure in this field comes from Gartner: poor data quality costs organisations an average of $12.9 million per year. Research from MIT Sloan Management Review with Cork University Business School puts the loss higher in relative terms, at 15–25% of revenue annually for organisations with serious data quality problems.

Those are global averages, and averages travel badly. The version that survives translation to an Azerbaijani boardroom is narrower and harder to argue with:

  • Supervisory rework. Every regulatory submission that has to be reconstructed by hand consumes days of senior analyst time, repeatedly, forever.
  • Duplicated engineering. Four teams building four versions of "active customer" because none of them could find the first one.
  • Blocked AI. A retrieval system pointed at ungoverned data returns confidently wrong answers, and the project is shelved — which is precisely the cancellation pattern Gartner is describing.

Azerbaijan's banking sector held roughly 57.1 billion manats (about $33.6 billion) in assets as of 1 April 2026, up 6.7% year on year. In an institution of that scale, a fraction of a percent of misdirected effort is a larger number than the entire governance programme.

Why Azerbaijan is a genuinely different environment

Most data governance writing assumes a US or Western European enterprise. Several of its assumptions do not hold here.

Metadata is trilingual, and this is not a cosmetic problem

A typical large Azerbaijani organisation runs systems whose metadata is in three languages simultaneously: an Azerbaijani-language core system, a Russian-language legacy application inherited from an earlier era or a regional vendor, and English-language modern infrastructure — cloud services, analytics tooling, anything bought in the last decade.

The result is that the same business concept exists under three unrelated names, none of which is wrong. müştəri, клиент and customer refer to the same entity, and no automated matching will connect them, because they share no characters.

This is the single most underestimated cost in an Azerbaijani governance programme. Organisations budget for it as a translation task and discover it is a modelling task: you are not translating labels, you are deciding that these three tables describe one concept, and that decision needs a human who understands the business in all three languages.

Practical implication: your business glossary must be multilingual from day one, with one canonical concept and language-specific labels attached to it — not three parallel glossaries. Retrofitting this after a thousand terms are catalogued is expensive and demoralising.

Deployment is on-premise, and that constrains tool choice hard

Banks under Central Bank supervision and state institutions handling citizen data generally cannot place that data in a foreign public cloud.

This follows from the Law on Personal Data, which subjects information systems processing personal data to state registration with the responsible executive authority and permits cross-border transfer only where the destination state provides adequate protection, the data subject has consented, or the transfer is necessary to perform a contract — and prohibits it outright where it would threaten national security or public order. The detail is set out in data residency and personal data law in Azerbaijan.

That single constraint eliminates most of the modern data governance market. The well-known catalogs are increasingly cloud-only or cloud-first, with on-premise offered as a legacy tier that receives less engineering attention. When you evaluate tools, "can it be deployed on-premise" is not a feature checkbox — it is the first filter, and it removes more vendors than any other criterion.

The regulator asks about lineage, and screenshots do not satisfy it

Supervisory questions about regulatory reporting tend to take a specific shape: this figure in your submission — show us how it was calculated.

The honest answer in most organisations is a chain of four systems, two manual reconciliation steps and a spreadsheet a specific analyst maintains. Reconstructing it takes days, and the reconstruction is a narrative rather than evidence.

Column-level lineage changes the character of that conversation. Instead of a narrative, you produce a traced path from the reported figure back through every transformation to the source column, generated from the systems themselves rather than from memory. This is the most concrete, most defensible return on a governance programme in a regulated Azerbaijani institution, and it is worth leading with when you build the business case.

The talent market is thin, so the programme must not depend on heroes

There are not many experienced data stewards in Azerbaijan. A programme designed around a handful of exceptional individuals will collapse when two of them leave.

Design for this: make stewardship a defined part of existing business roles rather than a new specialist career, keep the per-steward workload genuinely small, and automate everything that can be automated — profiling, classification, lineage harvesting — so that human attention goes only where judgement is actually required.

The regulatory drivers, in order of practical force

Central Bank of Azerbaijan (CBAR) supervision. For banks, this is the strongest driver. CBAR's Regulation on Information Security Management in Banks sets minimum information security requirements for banks operating in Azerbaijan and is explicitly built on the ISO/IEC 27000-series standards; it came into force on 1 April 2022. CBAR has since adopted a Cybersecurity Strategy for Financial Markets covering 2023–2026 and issued requirements for ensuring information security across supervised entities in financial markets. Its Financial Sector Development Strategy 2024–2026 continues the same direction of travel.

Read those documents as data governance requirements, because that is how they behave in practice: asset ownership, classification by sensitivity, access control tied to classification, and the ability to evidence how a supervisory figure was produced.

The Law on Personal Data. It establishes obligations around processing personal data, including state registration of the information systems that process it — with exemptions, notably for systems covering fewer than 1,000 data subjects — a lawful basis for processing, and restrictions on cross-border transfer. In governance terms: you cannot comply with obligations about personal data if you cannot say which of your columns contain personal data. That is a classification problem, which is a catalog problem.

National AI and digital policy. Azerbaijan's Innovation and Digital Development Agency has approved an Artificial Intelligence Strategy for 2025–2028 that names data governance as a priority area, and the Action Plan for the Acceleration of Digital Development for 2026–2028, approved by presidential decree on 27 February 2026, carries 58 initiatives across digitalisation, AI, the innovation ecosystem and cybersecurity. A Digital Development Council held its first meeting on 29 June 2026 and put the legislative framework for AI on its agenda. For state institutions, this is the direction the ministry will be measuring against.

The EU AI Act, for anyone serving European clients. Increasingly relevant to Azerbaijani companies with EU customers, and it imports data governance obligations directly: high-risk AI systems carry requirements on training data quality, provenance and documentation that are unmeetable without a catalog. Note that the timing changed — the Digital Omnibus, adopted as Regulation (EU) 2026/1744, deferred most high-risk obligations from August 2026 to December 2027 and August 2028, while transparency duties and GPAI enforcement began on 2 August 2026. The deadline moved; the data requirements did not.

Where governance programmes fail here specifically

Five failure patterns, all observed repeatedly in this market.

Starting with policy instead of inventory. A governance policy written before anyone knows what data exists is a document about an imaginary organisation. It will be approved, filed and ignored. Inventory first: what systems exist, what is in them, who touches them. The policy that follows will be about something real.

Cataloguing everything. An organisation with 40,000 tables that sets out to catalog all of them will still be cataloguing in three years, having demonstrated nothing. Catalog the twenty to fifty assets that appear in regulatory reports and executive decisions. Those are the ones where being wrong has consequences, and they are what will fund the next phase.

Treating the catalog as the deliverable. A populated catalog nobody opens is a more expensive spreadsheet. The deliverable is a behaviour change: an analyst who checks the glossary before building a report, an engineer who checks lineage before changing a schema. Measure that, not row counts.

Appointing stewards without authority or time. Stewardship added to a full workload, with no decision rights, produces a name in a field and nothing else. A steward needs perhaps two hours a week, the authority to settle a definition dispute, and a manager who counts it as real work.

Running it as an IT project. IT can implement a catalog. It cannot decide what active customer means. When governance sits entirely inside IT, definitions get made by people without the business context to make them, the business ignores the result, and the programme quietly ends. Governance is a business function with heavy technical support.

A 12-month programme that actually delivers

Sequenced so each phase funds the next by producing something visible.

Months 1–2: Scope and inventory

Pick one domain. Customer data in a bank, citizen registry data in a ministry — one domain, deep, not five domains shallow.

Inventory the systems in it. Not every table: the systems, their owners, roughly what they hold, and how data moves between them. This is interview work as much as technical work, and it produces the first genuinely useful artefact — a map of a territory that previously existed only in fragments across people's heads.

Deliverable: a system inventory and a data flow diagram for one domain.

Months 3–4: Deploy the catalog, harvest automatically

Stand up the catalog. Connect it to the source systems in scope and let it harvest technical metadata — schemas, columns, types, volumes — and lineage automatically.

Do not have humans type metadata that a connector can extract. The value of a catalog product is precisely that it reads systems directly; an organisation that populates it manually has bought a wiki.

Deliverable: technical metadata and system-level lineage for the domain, harvested rather than typed.

Months 5–6: The glossary, and the arguments it starts

Define the fifty to eighty business terms that matter in this domain. Multilingual from the start: one concept, labels in Azerbaijani, Russian and English.

Expect this phase to be slower and more contentious than planned. Definitions are where governance stops being technical: two departments will discover they have been using active customer differently for years, both defensibly, and someone has to decide. That argument is the point. Resolving it is worth more than the catalog.

Deliverable: an approved multilingual glossary, and a decision log recording what was settled and by whom.

Months 7–8: Ownership and stewardship

Assign a named owner to every asset in scope. Named individuals, not departments.

Stand up a small governance forum — six to eight people, meeting fortnightly, with authority to settle disputes. Not a committee of twenty that meets quarterly and escalates everything.

Deliverable: ownership assigned across the domain, a functioning decision forum, an escalation path.

Months 9–10: Data quality where it is measurable

Define quality rules for the assets that feed regulatory reports. Completeness, validity, timeliness, referential integrity — start with rules that can be checked automatically and that someone will actually act on when they fail.

An alert nobody owns is noise. Every rule needs a named recipient and an agreed response.

Deliverable: automated quality checks on the top twenty assets, with routed alerts.

Months 11–12: Prove it, then extend

Take one regulatory report and trace it end to end through the catalog — from the submitted figure back to every source column. Do it in front of the people who will have to answer for that figure.

That demonstration is what funds year two. It converts governance from an abstraction into a capability the institution can see.

Deliverable: a demonstrated end-to-end lineage trace, and a scoped plan for the next domain.

What a catalog does, and what it does not

Worth stating plainly, because the misconception is near-universal in first conversations.

A data catalog is not a database. It does not store your data. It stores metadata about your data — where it lives, what it is called, what it means, who owns it, where it came from, how it flows. Your data stays exactly where it is.

This matters more here than elsewhere, because the first question in an on-premise security review is usually whether the tool copies data out. It does not. A catalog connects, reads structure and statistics, and writes metadata to its own store.

What it gives you: search across systems, business definitions attached to technical columns, automated column-level lineage, sensitivity classification, quality rule execution, and an access-request workflow.

What it does not give you: decisions. It will not tell you which of your four customer tables is authoritative. It will show you that four exist, which is the necessary precondition for someone to decide.

We cover the mechanics in what is a data catalog.

Choosing a platform for this market

Filter in this order. The first two eliminate most of the market before you look at features.

  1. On-premise and air-gapped deployment, as a first-class product. Not a legacy tier. Ask when the on-premise version last shipped feature parity with the cloud release.
  2. Connector coverage for what you actually run. Oracle, MS SQL, SAP, 1C, mainframe extracts, and the specific core banking system in place. A catalog that cannot read your core system is decorative.
  3. Genuine column-level lineage, parsed from SQL and ETL code — not lineage drawn by hand in a UI, which decays the week after it is drawn.
  4. Multilingual glossary support — one concept, multiple language labels. Check this specifically; several products assume one language per deployment.
  5. A usable interface for non-technical stewards. Your stewards are risk officers and product managers. If the tool requires SQL to answer "who owns this", stewardship will not happen.
  6. Local implementation capability. A tool with no one in-region who has deployed it is a tool you will operate alone.

Gartner created the Magic Quadrant for Data and Analytics Governance Platforms in January 2025 and, in the 2026 edition, extended its evaluation into unstructured data, analytics models and data products, weighting active metadata and automation more heavily. Read it as a map of the market, not as a shortlist: the axes it scores on do not include on-premise parity, which is your first filter.

OvalEdge is the platform we implement, and the comparison against the alternatives is set out in OvalEdge vs Collibra vs Alation rather than asserted here.

How to start without a large budget

The most common blocker is that the business case requires evidence and the evidence requires the programme.

Break the loop with a scoped diagnostic: one domain, four to six weeks, producing a system inventory, a sample of harvested lineage for one real report, and a documented gap assessment against your regulatory obligations. That artefact is what a board approves against — not a vendor deck.

If it helps to have that structured externally, our data governance assessment is built to produce exactly those outputs.

Key points

  • Data governance is an operating model, not a tool. The tool is necessary and insufficient.
  • The business case is now AI as much as compliance: 88% of organisations use AI, only 39% see EBIT impact, and Gartner expects over 40% of agentic AI projects to be cancelled by end-2027 — largely for governance reasons.
  • In Azerbaijan the binding constraints are on-premise deployment, trilingual metadata and supervisory demand for lineage — none of which the standard international playbook addresses.
  • CBAR's information security regulation for banks, in force since April 2022 and built on ISO/IEC 27000, behaves as a data governance mandate in practice.
  • Trilingual metadata is a modelling problem, not a translation problem. Design the glossary for it on day one.
  • Scope narrow and prove lineage on a real regulatory report. That demonstration funds everything after it.