Core question
What are the signs that our data is not ready?
Short answer
Five signs show up repeatedly. Employees hunt for the right version of files. Different departments define the same thing differently. People build workarounds outside official systems. Reports need manual cleanup every week. And business knowledge lives scattered across PDFs, emails, and shared drives.
The first sign is spreadsheet archaeology. If your employees routinely spend part of their day hunting for the right version of a file, an AI system will run into exactly the same wall, because it has no way to know which version is authoritative either. What these systems need is structured, trusted, accessible information, and the fix starts with establishing one source of truth for the data that actually drives decisions.
The second sign is inconsistent definitions across departments. Sales describes a metric one way, finance another, and operations tracks something slightly different again. A person reconciles those differences instinctively in a meeting. A system cannot, and it will produce confident answers built on whichever definition it happened to find. Standardizing terminology and assigning clear ownership for each definition is unglamorous work that pays for itself the first time a report holds up under scrutiny.
The third sign is workarounds living outside your official systems. Shadow spreadsheets, personal notes, private AI prompts, and disconnected trackers accumulate when the sanctioned system is too hard to use or too hard to trust. Each one holds real business knowledge that no AI system will ever see, and each one leaves with the person who built it.
The fourth and fifth signs are related. If someone has to fix the numbers before every meeting, that is a data problem wearing a reporting problem's clothes, and automating the report will simply automate the error. And if your business knowledge lives in a scattered mix of PDFs, emails, Word documents, and shared drives full of folders named FINAL_v2_REAL, the knowledge exists but nothing can retrieve it reliably.
Two things make this harder to catch than an ordinary data problem. The first is that a capable model does not fail loudly on weak information. It produces a confident answer built on whichever version it happened to find, so the problem becomes less visible rather than more. The second is that conflicting definitions are a governance question rather than a technical one. No system can settle an organizational disagreement about which number is correct. Someone has to decide which source is authoritative, who may update it, and what happens when an inconsistency surfaces. Data governance is part of AI governance, and it is usually the part nobody assigned.
Real progress here usually starts with boring work: cleanup, governance, structure, and ownership rather than demonstrations. That is exactly why it gets skipped, and exactly where the return actually comes from.