Your Data Does Not Have to Be Perfect, But You Need to Know What It Means

[ad_1] A startup founder who builds analytics tools for real estate owners told me recently that there is no longer ...
builderkp

[ad_1]

A startup founder who builds analytics tools for real estate owners told me recently that there is no longer any need to obtain well-formed data for AI. He added that well-formed data is not realistically available anyway, even in fairly new smart buildings. Current AI, he argued, can sort the data out regardless of its original coherence.

He is not alone. Several founders have told me the same thing about their software: feed it any data in whatever format, and it will deliver useful results. In AEC, where we’ve struggled for decades to make data interoperable without achieving it, that notion sounds liberating.

So is it still worth the effort to pursue semantic and structural data quality before applying AI? It depends.

Why clean data is not always necessary

The real estate case is grounded in a specific and real problem. Built properties generate vast amounts of building telemetry, but most of it is unstructured, siloed, and stored in legacy formats that most systems cannot read without significant preparation. Until recently, making that data usable for AI required manual tagging and mapping, a process that could cost tens of thousands of euros per building and take months. At that price, most projects never got started.

New AI capabilities have changed that threshold. Systems can now ingest fragmented, inconsistently labeled data from building management systems and begin extracting useful patterns without manual preparation. The ingestion problem, which was once a genuine project killer, is becoming tractable.

Firms that delay AI workflow development until their data is in order will wait indefinitely. The founder put it more directly: accessibility matters more than quality. If the data cannot be reached and used, its quality is irrelevant.

What AI actually does with your data

There is another dimension to the problem that the “any data will do” argument bypasses. Structural quality refers to how well data is formatted and consistently organized. Semantic quality refers to whether the labels and descriptions actually convey what the data means. These are different problems, and AI exposes the difference in ways that earlier software did not. Large language models do not understand buildings. They understand the language that describes them. That distinction shapes everything about how reliable the output will be.

A model reasoning about a maintenance schedule, a cost estimate, or a structural system is reasoning about the words and numbers used to describe those things, not about the physical reality they represent. If the description is ambiguous, incomplete, or inconsistently labeled, the model will still produce an answer. It just may not be the right one.

This means that semantic coherence, whether data is described clearly and consistently enough to reason about, matters as much as structural cleanliness. A perfectly formatted dataset with cryptic field names may be less useful to an AI system than a messy spreadsheet with precise, descriptive labels. These are two distinct problems, and conflating them leads to poor decisions about where to invest in data quality.

The cost of messy data does not disappear

A researcher at Aalto University working on a construction AI application had to spend significant time manually cutting, pasting, and reconciling data from multiple sources before it was usable. That kind of labor is now largely avoidable. AI can handle much of the mechanical work of data reconciliation, but only if you give it sufficient context about what it is looking at and what it is supposed to do with it.

The human effort does not disappear; its purpose changes. Instead of cleaning data before the process, you now set context at the start and validate output at the end. In exploratory or low-stakes workflows, that trade is often worth making. In quantity takeoffs, lifecycle cost analysis, or any situation where a confident wrong answer has real consequences, the risk is on another level.

AI can work with messy data if you have the knowledge and means to turn it into trustworthy information. That context-setting work is not a one-time investment; it is project-specific because the data landscape and the stakes vary each time.

Data quality as a value chain asset

Until now, data quality has been evaluated from the perspective of a single use case within a single firm. A quantity surveyor’s dataset was good enough if it served the quantity surveyor. The value did not extend beyond that application, and the cost of improving it rarely justified the effort.

Because project participants could not easily use each other’s data, they largely rebuilt it themselves. This inefficiency is so embedded in AEC that most firms no longer notice it.

When AI systems consume data across organizational boundaries, feeding into procurement decisions, operational analytics, portfolio management, or third-party platforms, the firm that produces well-described, coherent data has something others can use. Imagine a contractor, product manufacturer, or building owner providing AI-ready data that dozens of downstream stakeholders can immediately consume without rework.

In a value chain where AI is increasingly connecting the dots, reliable and semantically rich inputs may become tradeable. This is still emerging, and the commercial models for it are not yet clear. But firms that have been building data discipline for years, not just for compliance but for operational clarity, are well-positioned for data-driven collaboration. Industry-wide efforts, such as Finland’s work on standardized product data, point to a future in which this kind of asset has real market value.

What to do now

Improvements in AI capabilities should not be treated as permission to ignore data quality. They should be read as opportunities for a faster start and a better understanding of the data’s potential. Build the workflows, use what you have, and instrument the process so you can see where coherence is critical and where it is not worth the investment.

Better data quality pays off even without AI. When your data describes what it is, consistently and precisely, you spend less time rebuilding what others already know and more time building on it.

The firms that take this seriously now will not just have better AI outputs. They may have an asset that others will eventually want to pay for.

[ad_2]

This article was originally posted at Source link

Leave a Comment