"We need to get our data AI-ready" is one of the more common statements in current technology planning, and one of the least specific. It usually means something closer to "our data is a mess and we suspect that will be a problem," which is a reasonable instinct without being an actionable one.
It is worth being concrete, because the gap between AI-ready and not is usually smaller and more mundane than people expect. It rarely involves a platform migration. It usually involves five properties.
One: definitions that exist somewhere other than in someone's head
If two reports calculate revenue differently, the problem is not the reports. It is that "revenue" was never defined as a shared artefact, so each report author made a reasonable local decision.
Humans absorb this. An analyst knows that the finance number excludes intercompany and the sales number does not, and adjusts. An AI system does not know that, and will not tell you it is guessing. Ask it about revenue and it will confidently give you whichever definition it happened to find.
Being AI-ready starts with a written definition per core measure, owned by a named person, implemented once in a shared semantic model rather than reimplemented per report. This is unglamorous work and it is the highest-leverage thing on this list.
Two: ownership per domain
Ownership answers a specific question: when this data is wrong, who decides what right looks like?
Not who maintains the pipeline — that is a technical role. Who has the authority to say that a customer record is a duplicate, that a category has been retired, or that a definition should change. Without that, data quality issues become permanent because there is no one who can close them.
Ownership should sit with the business function that generates the data, not with the team that stores it. That is often an uncomfortable reallocation, and it is the difference between quality issues being resolved and being logged.
Three: identifiers that survive crossing a system boundary
Most useful questions span systems. Connecting a customer in the CRM to the same customer in the finance system to the same customer in the support system is where analytical value comes from — and where most data estates break down.
When identifiers do not match, someone bridges the gap with a matching rule, usually based on name or email, usually maintained in a spreadsheet, usually by one person. It works until it does not, and the failure is silent.
Being AI-ready means the key entities — customer, product, employee, site — have a durable identifier that is consistent across systems, or an explicitly maintained and monitored mapping. This is often the largest piece of work, and it is worth scoping it to the entities that actually appear in the questions you want answered rather than attempting the whole estate.
Four: lineage that a human can follow
Lineage answers "where did this number come from." It matters for AI-facing data for a specific reason: when a generated answer looks wrong, the first question is whether the model made an error or the underlying data is wrong. Without lineage, that investigation is archaeology.
Full automated lineage tooling is a nice-to-have. A documented map from source system through transformation to the measure a user sees is the actual requirement, and can be maintained in a document if that is what the team will keep current.
Five: quality rules that are checked rather than assumed
Every data estate has implicit assumptions. Dates fall within a plausible range. Amounts are non-negative. Every transaction has a valid category. Reference data matches the master list.
These assumptions hold right up until an upstream system changes, and then they fail quietly. Reports keep rendering. Averages shift slightly. Nobody notices for a quarter.
Being AI-ready means the important assumptions are expressed as checks that run on ingestion and alert when they fail. Not every assumption — the ones that would change a decision. Start with the fields that feed your core measures.
Where Microsoft Fabric fits, and where it does not
Fabric is a strong answer to a specific set of problems: consolidating scattered data into one governed foundation, handling volumes that outgrew a Power BI-only approach, and giving several teams a shared platform rather than a set of overlapping extracts.
It is not an answer to the five properties above. A lakehouse with undefined measures, no ownership, and unmatched identifiers is the same problem in a better-organised location. The platform decision and the governance decisions are independent, and the governance ones determine whether the platform pays off.
A practical sequence: fix definitions and ownership for the measures that matter most, establish identifiers for the entities that appear in your priority questions, then decide whether your current platform is the constraint. Frequently it is not.
A realistic first step
If this reads as a large programme, it is worth narrowing it deliberately. Pick one question the business genuinely wants an AI system to answer. A real one, with a decision attached.
Then trace it. Which measures does the answer depend on? Are they defined? Who owns them? Which systems do they cross, and do the identifiers match? What would have to be wrong upstream for the answer to be wrong, and would anyone notice?
That trace usually takes a few days and produces a specific, bounded list of work — rather than a data strategy that takes a quarter to write and does not unblock the thing you wanted to do.
The organisations that get value from AI on their data are usually not the ones with the most complete data estate. They are the ones who picked a narrow question, made the data behind it trustworthy, and then did it again.
- Microsoft Fabric
- Power BI
- Data governance
About the author
Ahmed Salih
Writing for Aqlyst Technologies on AI agents, Microsoft Cloud delivery, data foundations, and digital experience. Biography and role details pending owner approval.
