Published Jul 2026

Copy Link Copied to clipboard

As AI use cases scale, weak data foundations quickly show cracks.

In our latest webinar, we explored common data trust gaps, how to operationalise AI-ready data, and how Aperture Data Studio helps organisations improve data where it lives and use it with confidence.

Watch the recording below:

00:00:13:05 – 00:00:15:16
Good morning. Hello, all. Thank you for joining us.

00:00:15:22 – 00:00:53:13
We appreciate you sharing your time with us today. I’m Evan Almass, a senior product marketing manager at Experian. Focused on Aperture Data Studio and partnerships. My background spans data modeling, analytics and commercial strategy across banking, tech and public service. I’m joined by Antina Lee, our Director of Product Management, who brings deep product and innovation experience for both going to be your primary speakers and then later within the session, we’ll be joined by Yao Li our Chief Product Officer with an experienced data quality, and Josh Boxer, Group product manager for Aperture Data Studio.

00:00:53:16 – 00:01:23:19
We all work out of Experian’s London office. So today we’re going to be talking about Operationalising trusted data as the foundation for AI. But I want to say upfront, this is not going to be a broad abstract AI discussion. We’ll talk about past examples of failures and success, a practical framework, and our solution. The question we’ll explore is this as organisations move faster with AI, analytics, automation and decisioning.

00:01:23:25 – 00:01:49:29
How do you make sure the data behind those decisions and processes can actually be trusted? AI doesn’t just use your data, it stress tests it. It puts more pressure on quality, governance, traceability and controls, especially in regulated operational environments. So I’ll start by talking about the challenges we’re hearing from companies like yours and the Operationalising Data Trust framework to address them.

00:01:50:04 – 00:02:18:23
Then I’ll hand it over to Antina, who will bring this into the practical side. What this looks like in Aperture Data Studio and where organisations can start. So staying on the agenda slide. This is how we’re going to structure the session. First we’ll look at why AI is exposing hidden data risk. Many of these issues are not new to organisations, but AI increases the speed, scale, and visibility of the impact of those problems.

00:02:18:25 – 00:02:45:01
Second, we’ll define what AI data actually requires. The key point is that AI ready data is broader than clean fields or traditional data quality. Third, we’ll talk about the shift from data quality to operational trust. In other words, how you move from identifying issues to embedding controls where data is actually used. Fourth, we’ll look at how organisations can Operationalise trusted data across their complex existing environments.

00:02:45:01 – 00:03:09:21
And finally, we’ll talk about where to start. How do you assess readiness, identify gaps and prioritize what to fix first? By the end of the session, you should be able to identify where your data foundation is most exposed. What I really data actually requires of you and your organisation, and where to start so that your efforts are targeted and deliver the most value.

00:03:09:23 – 00:03:35:15
Next slide please. Before we get into the content, it’s worth acknowledging who’s in the room. And once again I just want to welcome all of you. I know your time is very valuable. We have people joining from risk and credit compliance and governance, data and analytics, operations and technology. That includes everyone from credit risk leaders and decisioning teams to data governance leaders, platform owners and transformation leaders.

00:03:35:16 – 00:03:57:23
So while everyone may come from it, come at it from a slightly different angle. Many of you sit at the intersection of data controls and decisions. The common questions are the same. Can we trust the data being used in the decision process or model? And if challenged, can we prove the right controls were applied and the data was fit for purpose?

00:03:57:25 – 00:04:06:16
That is where data trust must become operational. Next slide.

00:04:06:18 – 00:04:35:22
All right. So let’s start with talking about the pressures that organisations are facing. AI is stress testing every foundation in three ways. So starting on the the left hand side of this slide, the first is more data and faster decisions. AI, analytics and automation increasingly draw from more resources, more formats and more unstructured information. It’s no longer just structured fields in a single system.

00:04:35:23 – 00:05:02:04
It could be PDF documents, emails, notes, behavioral signals, and more. Also, models and agents can push data into workflows much faster than traditional review processes were designed for. Manual tracks, workarounds, and after the fact reviews may have worked when the processes move more slowly, but they can’t be relied on when processes are automated and scaled when there is simply much faster.

00:05:02:06 – 00:05:33:06
These factors require automated, scalable processes and systems that determine whether trustworthiness that determines trust, trustworthiness before processes or decisions are implemented. The second pressure is higher scrutiny. Risk, compliance and governance. Teams increasingly need evidence that data and controls can stand up to review. That applies whether the use case is AI or any other business critical process, such as regulatory reporting or credit decisioning.

00:05:33:09 – 00:06:01:23
And then the third pressure on the right hand side here in the purple column is that legacy controls are falling short. Many organisations have invested in data quality, governance, reporting and analytics. The issue is not a lack of controls. Rather, it is that those controls were designed for slower processes, processes and structured data. They may depend on periodic checks, manual review and local knowledge that sits outside of day to day workflows.

00:06:01:23 – 00:06:27:17
And it’s silent. That creates a gap when AI and automation starts using data continuously across fragmented systems. It’s important to note that AI is not creating all of these data problems. These probably sound really familiar to you. In most organisations, these issues have already existed for years. They were there were already inconsistent definitions, manual fixes, fragmented systems, and local policies.

00:06:27:19 – 00:06:50:29
AI simply makes those issues more problematic and harder to hide. That’s why the challenge is not just better data quality. The challenge is Operationalising data trust, making sure trust is embedded in the way that data is created, validated, governed, and used. Next slide please.

00:06:51:01 – 00:07:18:22
So this slide shows five patterns we see repeatedly when organisations try to move from experimentation into production or for manual workflows in a more automated decisioning and processes. Once again, while these gaps are not new, AI increases their speed, scale, and visibility. So we’ll go through the patterns that we see here. The first pattern is that trust is embedded is defined, but not embedded into workflows.

00:07:18:24 – 00:07:43:06
Most organisations have standards. They know what good data should look like and they have policies definitions and governance governance forms. But the problem is that those standards are not always applied consistently where the data is used. The second pattern is that governance is documented but not enforced. According to IDC, 53% of organisations are at the beginning of their governments journey.

00:07:43:13 – 00:08:11:05
A policy may exist. A catalog may exist. Ownership is probably written down somewhere and the people who are responsible now. But if those controls set outside workflows, when people are under pressure, they’ll they’ll move towards the path of path of least resistance and may bypass them. The third pattern is fragmented ecosystems. This is especially relevant for many organisations here today.

00:08:11:07 – 00:08:39:16
Data doesn’t live in one clean central place. It spans legacy platforms, hybrid environments, case management systems, C, CRM and ERP systems, and other finance systems. Plus, organisations have multiple tools to manage all of the data. In fact, Gartner found that organisations have deployed a dozen management systems on average. I actually saw this firsthand when I worked in monoliths development in a bank.

00:08:39:18 – 00:09:02:11
Data was siloed in separate business units, many of which were acquisitions with their own data models. We needed to monitor all transactions for financial crime risk. But it was difficult to do this because we didn’t have a single customer view, among other challenges. The fourth, fourth pattern is that impact is hard to see. Hard to trace when something changes or breaks.

00:09:02:12 – 00:09:30:12
Teams often struggle to answer basic questions quickly. Where did the issue start? What processes and models does it affect and what is the business consequence? The fifth pattern is that quality is manual and reactive. I mean, IDC found that data quality management was the top investment priority for CDOs today. This makes sense because often teams only know about and fix issues after they’ve broken something at small scale.

00:09:30:13 – 00:09:52:24
People can work around these problems, but at AI scale, those workarounds become risk and potentially loss business and penalties. That is why we need to move from isolated data quality activity to a broader operating model for trust. At the bottom of the slide, you’ll see the symptoms and impacts of these issues. Projects slow down or fail. Teams shift to workarounds.

00:09:52:25 – 00:10:11:05
Value is lost and outputs aren’t as expected. These result in increasing regulatory, financial, operational and reputational risks. So we’ll start with a couple examples where things haven’t gone as planned.

00:10:11:07 – 00:10:43:13
Not every issue on this slide or in your organisation is an AI failure. Yet these examples show a pattern. When trust gaps reach critical decisions or regulated processes, they become business risks. The first example is UK insurance and the Prudential Regulation Regulation Authority. This specific issue was a regulatory reporting issue resulting from ineffective controls. A balance sheet miscalculation overstated the firm solvency position to the regulators and market.

00:10:43:21 – 00:11:16:01
The trust gap was around complete, timely and accurate data supported by effective controls to prevent and detect reporting areas, and the resulting failure was costly. They incurred a 10.6 million pounds five. The reason this matters in today’s discussion discussion is simple. If a data quality or control issue can create regulatory exposure in a traditional reporting process, the risk becomes even harder to manage when similar data feeds, automated workflows or AI decisions.

00:11:16:04 – 00:11:47:04
The second example is Air Canada and Chatbot gave incorrect information about discounted bereavement fares. The trust gap in this example in the center column was not simply that the chatbot produced a poor answer. It was also that approved knowledge, policy context and controls were not sufficiently embedded into the customer advice workflow. Air Canada had to pay the damage in tribunal fees, but the greater cost was poor press coverage and reputational impact.

00:11:47:06 – 00:12:11:18
The third example in this purple column is an issue with optimal health care algorithm. In that case, a healthcare risk algorithm use costs as a proxy for clinical being that underestimated the health needs of black patients and reduce access to care management support. This likely caused serious health issues for the patients. The app was around suitability and representation.

00:12:11:21 – 00:12:19:03
Was the data fit for the decision it was being used to support? Doesn’t sound like it.

00:12:19:05 – 00:12:49:27
Across all three examples, the pattern is consistent. The visible problem may be a wrong report, wrong, or a wrong decision, but underneath the issue is that trust was not Operationalised where the data was used. That is the lesson for everyday data. It is not enough for you to look acceptable in a report or pass in your check. It needs to be reliable, relevant, safe, representative and governed in context of the decision or workflow its use.

00:12:49:29 – 00:12:57:06
We’ll shift to talking about a few examples for companies have gotten it right.

00:12:57:09 – 00:13:32:28
Excellent. The previous slide shows what happens when trust gaps become business risks. This slide shows the brighter side what happens when trusted data is effectively Operationalise. These are all examples of where Experian Aperture Data Studio helped a client produce and scale data they could trust. They show what happens when quality and governance are embedded in everyday workflows. In the first example, in the pink left hand column, a European bank use automated credit data validation to move from a process that took around 24 hours to one that took only 90s.

00:13:33:00 – 00:14:03:23
The benefits are not only in time savings, though they’re also an effective controls. Faster validation supports faster and more defensible decision making. In the second example, the Cleveland Police force automated manual data processes and reduce duplicate accounts. The result was not only operational efficiency, but also better citizen record quality for frontline processes. This not only meant cost savings, but likely also life saving response.

00:14:03:25 – 00:14:37:09
In the third example, a UK charity lottery operator reduced undeliverable mail through better contact validation and segmentation that linked data quality directly to cost leakage, customer experience and improved targeting. While the sectors and use cases are different in these examples, the pattern is the same. Trusted inputs, embedded controls, automated processes, and ultimately measurable outcomes. All of it tailored to the specific use case and reusable across future initiatives.

00:14:37:11 – 00:14:43:00
Next slide please.

00:14:43:03 – 00:15:12:15
Next slide please. So what does AI data actually mean? It does not mean replacing data quality and governance. In fact, those are still the foundation. It means asking whether those capabilities are sufficient for AI analytics or decision in use cases. Classic data quality still matters. Completeness, accuracy, consistency, validity, timeliness and uniqueness are all so important. Those are the building blocks, but they’re not the whole answer.

00:15:12:17 – 00:15:35:22
They already did. It also needs context. It needs to be relevant to the use case. It needs to be safe and appropriately controlled. It must be traceable. It needs to be representative enough for the population or process involved. And it needs to be monitored continuously as data, rules and models change. And this is where the early examples are useful in the Optum healthcare case.

00:15:35:23 – 00:16:02:00
The issue is that the data used didn’t fit the context regarding the decision being made. That lack of context resulted in bias and harm. On the brighter side of things, when things went well. The bank’s credit data validation example showed what happens when trusted data is Operationalised by validating and controlling credit data inside the workflow. The organisation achieved trusted data and speed.

00:16:02:03 – 00:16:26:04
So for. I really data. The question shifts from is this field correct. To can we rely on this data in a real decision model or operational workflow? And that brings us to the most important point. Everyday data is not a static state is something. It is not something you check once and then assume it will hold data changes and context changes.

00:16:26:05 – 00:16:52:20
So readiness has to be maintained in context continuously and inside the workflow where data is actually used. That is what we’ll move into next. How organisations can make this repeatable by understanding what can be trusted. Improving what is not ready, and governing trusted trust as data rules and use cases involved. So what you’re seeing now on the screen is the operating model.

00:16:52:21 – 00:17:14:16
Model we’re going to be using for the rest of the discussion understand, improve and governance. This is how we determine whether data is ready for AI or other use cases. Fix it and ensure controls are in place to ensure it’s always ready and properly evidenced. First, the left hand column. The pink column is about understanding the current state of the data.

00:17:14:18 – 00:17:45:17
That means profiling and assessing the data, identifying issues, understanding dependencies, and seeing where those issues might affect downstream reports, processes, and decisions. Second column the blue column. Improve what is not ready. That means standardizing, validating, duplicating and enriching data with leading Experian or third party data sets so that it’s fit for the intended purpose. Finally, governance, trusted execution and proof.

00:17:45:19 – 00:18:10:11
That means enforcing policies, ownership and approvals, making lineage and impact visible, and maintaining evidence that the right controls were applied so that those can be shared with auditors and regulators. This is not a one time exercise. In fact, we could have actually displayed these three different steps in a circle. As data and context changes, we must understand the data.

00:18:10:13 – 00:18:28:17
We must fix it and ensure it’s well controlled. Then we start the process again. So AI radius has to be continuous, not a one off glance. The goal is sustained confidence in models, agents, processes, decisions and business outcomes.

00:18:28:19 – 00:18:56:04
So this brings us to the shift from data quality to operational trust. Operationalised trust answers one critical question can this data be trusted? But that question has to be answered in context. Trusted for what use? Case. By whom? Under what controls? With what evidence? And what happens if something changes? Operationalise. Trust means teams can see data quality issues, data quality across critical data sets.

00:18:56:09 – 00:19:23:01
They can validate continuously and remediate proactively. They can embed controls and operational workflows. And they can connect data quality to business risk cost and outcomes. This is the distinction a data quality dashboard can tell you that something is wrong. Operational trust helps you prevent, prioritize, fix governance, and evidence the issue in the in the workflows where it matters.

00:19:23:04 – 00:19:35:22
Visibility is important, but it’s not sufficient. Warning signs only matter when controls are embedded where debt is actually used, and you can fix the data.

00:19:35:24 – 00:19:57:01
So before we move into Aperture Data Studio, I want to pause for a moment for a self-assessment. These questions that we have on the slide, I would encourage you to ask inside your own organisation. Think about it now as we talk through it and bring it back to the rest of your teams and kind of think through it as well to see where you might be on your maturity.

00:19:57:04 – 00:20:25:28
Under understand, can you assess data quality continuously? Can you monitor readiness continuously? And can you see what would be affected by a data issue or change? Understanding must be continuous and it should enable effective prioritization. It should help answer the question which issues matter most to the business? Under the improve column, the questions are. Can you reduce manual preparation?

00:20:25:29 – 00:20:54:07
Can you standardise processes across teams and can you scale validation, matching and enrichment? Manual effort is limiting progress. This is where many teams feel the pain. They can fix the problems, but too often the fixes depend on specialist effort, manual workarounds, spreadsheets, and more. These questions point to whether the approach is efficient and automated under governance. Do you have clear ownership?

00:20:54:13 – 00:21:17:01
Are controls embedded into workflows or only documented? Can you trace impact? And can you provide evidence to risk compliance and audit teams? If the answer to any of these is not consistently, then that is a sign that trust may be defined but not yet Operationalised. We’ll now shift to a poll.

00:21:17:04 – 00:21:41:01
So before I hand it over to Antena, we’d like to hear where this is all showing up for you. We’re going to run two quick polls. Both are multi select so please choose everything that applies to you. The first question will be where is trust. The data most critical for you today. And the second question is what is the biggest barrier to Operationalising trusted data within your organisation.

00:21:41:03 – 00:21:50:22
All right. I think we’re no longer seeing many of the responses coming in. So I think Connor, can you please share the answers with the with the audience?

00:21:50:24 – 00:22:17:11
Okay, so I’m kind of curious, what do you think of these results? Does does anything surprise you? Well, for the first question, the worst trusted data, most critical pretty, we see a pretty even split. I think that totally makes sense, given the audience in the room where a lot of these are quite important. For example, the last option, the last four I believe, had a pretty even split there.

00:22:17:11 – 00:23:03:14
But definitely, for example, in analytics and automation initiatives was heavily selected, which is very much what we expected given the topic in the audience, where AI, analytics and automation are putting more pressure on data foundations because they’re using data faster, at greater scale, and often with less manual human in the loop review. And I think the other, another one that I want to call out is the customer citizen operational Workflow, I believe, was another one that had quite a bit of selection there, which once again is really important because it shows trusted data is becoming operational, where poor data doesn’t just affect reports and analytics or machine learning models, it’s effect customer journeys, citizen services,

00:23:03:14 – 00:23:28:23
case management, collections and sort of onboarding and data day decisions as well. So definitely in line with what the audience is experiencing here and what we’re expecting as well and what we’re hearing. Yeah, that makes sense. Yeah, I think we’re seeing the the impact of legacy processes and systems. I think we’re also seeing these these challenges are broad based.

00:23:28:26 – 00:23:59:14
With the second poll. Just want to touch on it briefly. I believe the split was I don’t see the results on screen anymore, but I believe it was pretty top heavy where the options A, B and C had quite a bit of people selecting that as the biggest barrier. So fragmented systems teams and data source, I believe, was the option to have the most selection, which is definitely one of the most common barriers that we see where most organisations are not working off of a clean, centralized data environment.

00:23:59:14 – 00:24:26:16
So as you had mentioned earlier, more data is spread across systems teams, different solutions, different applications, and even partners in third parties. So the trust can break down often at the hand of points. Yeah, yeah. And and that makes it important not to have to have another, you know, system or data migration. Right. Having having systems that you can use within your existing environment.

00:24:26:19 – 00:24:56:16
Wonderful. All right. Well what so what we just talked about, it shows that data trust is really a common problem, a common challenge that organisations are seeing. And it’s surfacing in many different ways across the organisations. So once again, that’s why the answer can’t just be another data quality report. Controls need to be Operationalised for data is actually used.

00:24:56:19 – 00:25:09:04
And so that’s that’s where I think, you know, it makes sense to to bring you into the discussion and to talk a little bit more about Aperture Data Studio.

00:25:09:07 – 00:25:40:25
Sure thing. So before I talk about how Aperture Data Studio Operationalised Data Trust, I think it’s worth noting that the barriers and needs for trusted data is not unique to anyone industry, and we’re definitely sealing similar patterns across, be it financial services, insurance, technology, and other data intensive industries. So aperture is designed to meet you where you are, whether you need lie, and understanding the state of your date to health of your data, improving it, or governing it via these six capabilities on screen.

00:25:40:25 – 00:26:05:22
So left to right aperture can help you understand what data issues exist, help improve data through cleansing, validation, and enrichment, and finally help govern how data is owned, controlled, and more importantly, carry the evidence into your processes, your models and agents that depend on that data. So let’s dive into these in a little bit more detail to show you the solution in action, starting with understanding your data.

00:26:05:22 – 00:26:33:22
And before you can confidently use data, you need visibility to the state of it. So aperture can profile data sets to calculate statistics like completeness, uniqueness, consistency. And there really, they’re acting as a good starting point to determine any to spot and determine any underlying data issues. So profiling gives us that factual baseline. But we also need to understand what should be true for the data in the context of your business needs.

00:26:33:23 – 00:27:04:27
Any regulatory requirements, expectations around the operational processes or the AI use case the data supports. And this is where rules validation comes in. Rules validation allows you to translate your business expectations into repeatable checks. So for example whether mandatory fields are populated, whether those values meet policy thresholds, whether there are sensitive data sensitive fields that are tagged, whether data is current enough for the intended use.

00:27:04:29 – 00:27:30:26
So aperture can accelerate rules creation with smart suggestions one. Back please. Back to the last one. Yeah, one I believe. One more. Thank you. One more please. Sorry. The slides skipped ahead a couple. So this one rules validation is what we’re talking about here. You can see if you can play the video please. Here you can see aperture right.

00:27:30:27 – 00:27:57:15
Coming up with a list of suggestions which is based on the characteristics of the data that you can check to apply. And if you want to configure additional rules, we offer several options, including a mapping of functions to columns you can use. You can see in action here. You can also auto generate rules from natural language descriptions which is coming up in a second here.

00:27:57:16 – 00:28:01:22
Just let that run.

00:28:01:25 – 00:28:14:26
So on screen now you can see an example of the user typing in emails do not contain the word tests. You can adjust the result as required using AI.

00:28:14:28 – 00:28:39:04
And so really this flexibility empowers both technical users and business users to create rules quickly and accurately. See. So in a few seconds you’ll see some of the results preview. And the validation results can be output in several different ways, including interactive reports that allows you to drill into failures. So here you can see an example of that.

00:28:39:07 – 00:29:05:02
And the results can also be posted out to a dashboard. Any external software solutions you use or trigger email alerts and other automations. They can also be easily save as a reusable workflow as a rule set for consistent reference going forward, so you don’t have to recreate every time. So in addition to this, if you already have a defined AI use case in mind, Exide, please.

00:29:05:04 – 00:29:26:14
This is also where our AI ready data assessment can help. The data assessment looks at relevant data sets through the lens of your use case and ask is the data reliable, safe to use, relevant and representative enough for your AI use case? And we’re actually going to touch on this a little bit more. So I’m not going to dwell too much on it.

00:29:26:16 – 00:29:51:09
And the other thing is, as Evan mentioned, since data changes, your source systems change, policy change, models change, and business needs change. Aperture can help monitor data over time, the technologies and continuously check whether data remains fit for use. So after understanding what what’s of your data can be trusted where the gaps and opportunities are in the data, the next question really is how do you improve it?

00:29:51:09 – 00:30:17:15
And this is where aperture can help transform fragmented data into trusted, compliant information very quickly. So first we use machine learning to auto tag columns. Here’s an example of how it’s done. Instantly recognizing and contact data account names date currencies. And then from here it can service smart suggestions for cleansing and transformation activity based on the unique characteristics of your data.

00:30:17:17 – 00:30:36:04
So here it’s suggesting, for example, how zeros and Noles can be standardized. And that’s highlighted here how white spaces should be dealt with. And you can essentially select which ones you want to apply based on these suggestions.

00:30:36:07 – 00:31:02:26
And if you need here’s a preview of the results. And then of course if you need custom transformations. Once again, aperture can turn natural language into executable functions. So you can see an example of that. We’re applying cleansing to values in the email column. It will show instant previews at the bottom here. And you can further tweak with AI once again to iterate the function.

00:31:02:28 – 00:31:08:23
And then of course, once again you can save everything you’ve done as a workflow.

00:31:08:26 – 00:31:47:22
Aperture can also integrate with third party reference data to enforce standards. So here we have an example of how exchange rates are leverage within the workflow to apply consistent currency conversion calculations where the source of the truth lies externally. In addition, aperture can also be coordinated for data cleansing and enrichment activities. So in particular, here’s an example of how aperture can validate address emails phone numbers, which can help lead to if you’re trying to get to a single customer view, or with the right contact data to enable you to have a trusted data foundation for your contact information.

00:31:47:25 – 00:32:17:09
So another major part of improving data is resolving duplication and fragmentation. Next slide please. We talked about how many organisations have multiple records for the same customer, household supplier or business entity. And the records may sit across different systems right. Different solutions, different formats. They may contain conflicting values which can create challenges. So if the same customer, for example, appears three times in three different systems, which versions should should we trust?

00:32:17:15 – 00:32:41:27
Which records should feed the model, which addresses should be used. So our matching linking emerging capabilities can help you create a more unified and reliable view. The matching engine supports rule based approach and has been engineered for both high volume bulk matching and real time matching, which matters a lot for use cases like customer 360, fraud detection, data consolidation, entity resolution.

00:32:41:29 – 00:33:22:05
But matching is not always black and white, so some matches can be ambiguous. Aperture provides a cluster review capability, which allows the users to inspect and resolve potential matches, and it provides that balance between automation with human oversight. So the match and deduplication can be particularly relevant for AI data reduce because duplicates and fragmented entities can distort model training and a genetic workflows if the same entity is overrepresented, for example, or if the records that are related are not connected properly, the alchemist can be compromised.

00:33:22:08 – 00:33:39:11
Next, a quick feature spotlight on GI. We’ve seen Genoa come up earlier a few times, but to showing another example of how it can accelerate time to insights here a user can. Next slide please with the Jenny video here.

00:33:39:14 – 00:34:01:11
There’s a little bit of lag. But so here’s an example of how I can accelerate time to insights. So here the user is describing what they want to do with the data set. So in this case they’re trying to ask this software to calculate average sales by country. And we want to order results and do some modifications. Add a new calculated column aperture returns.

00:34:01:11 – 00:34:23:15
The output shows the action steps are taken on your right on the right here so you can understand how the result was created. And you can make revisions as required. And once you’re the user is happy with the output, the actions can be saved as a workflow and reuse when new data comes in. Okay, so we’ve talked about a lot about data improvement.

00:34:23:19 – 00:34:47:03
The third pillar is around governance. So next slide please. Yeah. Thank you. As organisations scale AI they will increasingly need to prove not only that a model or agent exists, but understand really what data. It depends on who owns that data. What did that data meets. Require thresholds controls that apply and evidence that exists if the output is questioned.

00:34:47:04 – 00:35:12:04
So aperture helps organisations catalog their data models, agents and related assets. And you’re able to assign owners policies approval workflows since. So essentially data quality and governance with an aperture are not treated as separate activities. And they become part of the same control framework. And when changes occur, aperture can support approval workflows to ensure that changes are documented and controlled.

00:35:12:04 – 00:35:39:28
And then finally, organisations may also need to understand how their data process, system models, and business outcomes relate to each other. So our governance provides that clarity to help teams understand dependencies and how changes in one thing can impact another. So I know we covered a lot of capability very quickly from profiling, validation rules, data cleansing transformations, contact validation, matching, deduplication, and governance.

00:35:40:01 – 00:36:00:17
So let’s bring it back to where you might be able to start. If you already have a specific AI use case in mind, a good starting point can be an AI ready data assessment delivered by Experian. This is designed for teams that want clarity on is the data behind an AI use case ready, or the risks and what should be prioritized to fix?

00:36:00:17 – 00:36:26:20
First, the assessment looks across the four dimensions as how we define is AI ready data, which is data that is reliable, safe, relevant and representative. So for example, for reliable, we can look at some of the capabilities mentioned before. What a data is accurate, complete, consistent, unique, valid and timely enough to support a use case. So this includes things like profiling rules validation.

00:36:26:20 – 00:37:06:16
It also includes our contact validation which is checking things like weather address emails and phone data are not only populated, but accurate and usable relevant. We can ask whether the data has the right meaning, the context, the fields definitions for the intended AI use. Safety use. We often look at sensitive data, the presence of IT, restriction attributes, permissions, and external and internal policy rules that should be validated and controlled or validated and applied before I used representative, we can take a look at whether the data gives AI a representative view of the populations and or entities it is meant to support.

00:37:06:16 – 00:37:28:25
So this could include looking at duplication, coverage and distribution gaps. So to make this a little bit more tangible, we’re showing here just a couple of examples of AI use cases and the types of readiness areas that we can we can assess. I do want to call it though that when when we say AI right, they mean they may mean different things to different people.

00:37:28:27 – 00:37:52:02
Some of you may be more focused on predictive models, right? Such as credit decisioning models, fraud detection models, pricing models, others. When they talk about AI, they may be referring to a assistants that do research, summarize documents or drafts responses. And of course, let’s not forget AI agents that retrieve data, call tools, and take actions and make decisions in the workflow.

00:37:52:02 – 00:38:26:08
And really, here is is key to understand that the types of AI right, that use case changes the types of data that’s involved. Predictive models, they often depend very heavily on structured data such as customer data, decision history, transaction history, claims data, genuine agents. They often depend on a broader mix of data, which includes unstructured data such as policies, websites, PDFs, knowledge articles, case notes, contracts alongside structured customer or operational data.

00:38:26:08 – 00:38:49:25
So I do want to call out unstructured data because we have partnered enabled unstructured data capabilities that allows us to support unstructured data in our solutions. So the four dimensions of data stay consistent. So you see the four four areas reliable, relevant, safe and representative. But we may apply them differently depending on a use case. For example let’s take a look at the second one.

00:38:49:25 – 00:39:14:13
There are customer and entity intelligence use case. We may focus on whether customer records can be connected across systems. Are there duplicates our definitions consistent? Are the contact cells valid? Is sensitive data present where they shouldn’t be and controlled? Essentially, does the AI use case have a complete enough view of the customer or entity to support better decisions and outcomes?

00:39:14:13 – 00:39:37:20
So to summarize, our assessment will need to start with use case. Right. We can take a look at we can support both structured and unstructured data that it depends on. And then assess what needs to be true for that data to be ready for that use. Case next slide please. This is just to really take it home with a very specific example around public sector.

00:39:37:21 – 00:40:06:17
The use case here is AI Assistant Citizen service in case management. And the core questions to ask is if records, eligibility data, case history, service interactions and identity data support say safe AI assisted decisions in an assessment. For example, for reliable data, we would look at citizen records, data profiling, address validation, data quality rules. Metrics could include the address validation and pass rates.

00:40:06:17 – 00:40:46:17
For example, if a significant share of citizen address fail validation, a recommendation may be to opt to fix that, but also to standardize and validate address data at the point of capture before it feeds eligibility score for relevant data. An example area to review is the consistency of field definitions across case management, benefits and identity systems. So if eligibility status is defined differently across systems, that recommendation may be to align those definitions in a log before the field is used in an AI assisted process for represented data, we may look at duplicate records.

00:40:46:17 – 00:41:17:05
If duplicate case records exist, the recommendation would be to resolve those duplicates into a golden record. So the model sees one accurate record per citizen. So as you can see, the assessment gives clarity of where the data stands today for use case and what needs attention. And it is a starting point that becomes the basis for ongoing readiness by embedding the right rules, monitoring governance and controls so the data stays fit for use as the data and the use case evolves.

00:41:17:08 – 00:41:40:07
And here we’re going to bring everything together with four key takeaways. First RDNA is a data trust problem. Models, agents and automations are only as reliable as the data foundations behind them. Second, having good quality is a good start, but AI readiness needs to go further. You also need the context, governance, and a way to monitor that the data remains fit for use over time.

00:41:40:09 – 00:42:00:09
Third, data trust has to be Operationalised. The rules, ownership, monitoring and evidence needs to be built into the way data is created, change and used. And finally, start with clarity. You don’t need to solve every data problem at once. For every data set, you can start with a high value AI use case. Identify the data set it depends on.

00:42:00:09 – 00:42:27:03
Assess what is ready, risky and what should be fixed first. So to close off, I hope this session before we move to. I hope we have shown you how Experiment Data Quality and Aperture Data Studio can support your journey to Operationalise data trust and helping you understand how ready your data is, improve the areas in detention, and put the right governance in place so you can have greater confidence in your AI models, agents, processes and business outcomes.

00:42:27:05 – 00:42:34:05
I’m going to now hand the floor back to Connor who will open it up for Q&A.

00:42:34:08 – 00:42:55:08
Thank you very much. That was really insightful session. I really enjoyed that. So we have had some questions into the chat. I’m just going to encourage as well the rest of the speakers. If you could all put your webcams on. I’ll start reading through the questions. And if you could all kind of just have at it when you feel like taking one.

00:42:55:13 – 00:43:00:11
Let me start. Let’s look at look. Bear with me.

00:43:00:14 – 00:43:14:13
So first question how do we know whether a data set is actually ready for AI rather than just passing standard data quality checks?

00:43:14:15 – 00:43:37:19
I will take that question. So. Data quality checks, they often tell you whether data is like complete valid, consistent. Those are essential. But they’re not the whole AI readiness question. So for AI ready data we also need to ask is the data for the AI use case. Is it relevant to the model the agent or the workflow. Right.

00:43:37:23 – 00:44:00:11
Sort of the four dimensions. Is the data safe to use? Does the data represent the population, the process or the entity the AI is meant to support? So essentially a readiness builds on traditional data quality. But it goes further. Right. It looks at those four dimensions reliable, relevant, safe to use and representative in the context of the alchemy you want to enable.

00:44:00:13 – 00:44:21:13
It’s very context specific and use case specific. Whereas traditional data quality may be quite factual. Right? And more around like is my data clean. And so every data is brings in that context and looks at a lot more than that.

00:44:21:15 – 00:44:37:21
Brilliant. Thanks, Tina. So right, moving on to the second question. So how does this help risk compliance or audit teams evidence that data controls are working?

00:44:37:23 – 00:45:09:05
I can take this one. So great. Great question. Thank you for this risk. Compliance and audit teams need more than a point in time statement. The data. The data is good. They need evidence. Aperture Data Studio helps by making data quality checks, rules, ownership, approvals, lineage and audit trails more visible and repeatable. That means teams can show what rules were applied, where issues were found, how they were handled, who owns the data, and what controls are in place.

00:45:09:08 – 00:45:34:01
Our partnerships also help even provide a better understanding of data lineage and data provenance, which makes this even easier. The value is beyond improving the data. It is being able to evidence that the right controls are working in the processes, reports, models or workflows that depend on that data.

00:45:34:03 – 00:45:48:23
And then let’s have a quick look. I think this is kind of related how to ensure GDR compliance while using AI on data.

00:45:48:25 – 00:46:15:17
I’ll try that. So it’s quite a broad question, but my first thought is that AI models and agents should only have access to the correct data in quotes. Could map that by business processes. If a customer or prospect has opted out of being contacted, which is kind of the main part of GDPR, then you make sure that only customers who haven’t opted out are the ones available to an agent.

00:46:15:20 – 00:46:39:19
And then, what? That doesn’t. You might be thinking with that question, and when I read it, GDPR doesn’t really cover this yet. Or maybe there’ll be a further kind of policy, GDPR related as a consumer myself or whatever. Can I opt out of my data being used to train a model or just be aware? Just ask the question of where is my data being used?

00:46:39:20 – 00:47:03:17
Am I being used to train a model somewhere? I don’t mean that really exists as a policy yet, and it’s not something that people are capturing, but I can see it coming in the very near future. So people would probably want to think about that as well. But yeah, mapping data to business processes and making sure the policies that you already have in your organisations are being applied correctly to the data is what you want to capture for that kind of compliance.

00:47:03:19 – 00:47:10:21
If anyone else does need to add to that.

00:47:10:23 – 00:47:31:15
I think I think that’s a great response. Just add to it as well. I think one of the key pieces here is also kind of understanding that that data lineage and data provenance, where the data come from, how is it used, how has it been transformed, and then how is it being, how is it showing up in different reports, models or processes.

00:47:31:17 – 00:47:58:13
And that’s that’s a key part of what we can do within the catalog. We can one with with partnerships. We can actually get to that level of understanding field level transformations. And then that can surface within the catalog and we can apply the policies to it. And then also see if that’s within conformance. So a couple different pieces there.

00:47:58:15 – 00:48:05:16
Thanks both. All right. Let me dive back into the chat with me.

00:48:05:19 – 00:48:15:12
So what is the difference between monitoring data quality and Operationalising data trust.

00:48:15:14 – 00:48:38:03
Yeah I’m happy to start on this one. So monitoring data quality. It tells you when something is wrong or changing. That is one part of the picture. It really touches on the understand pillar that we discussed earlier. So Operationalising trusted data means taking that a couple steps further. That brings in all three of the different pillars we talked about earlier.

00:48:38:05 – 00:49:22:03
Understanding, improving and governing. It means embedding checks, rules, ownership and evidence in the way that data is actually created. Change for use monitoring might tell us the field is failing a rule. Operationalise trust helps you understand why it matters who needs to act and what what downstream processes could be affected. It also implements the facts and proves the issue has been controlled or resolved so that when future stakeholders, auditors, regulators look at your controls, they see that they are properly in place.

00:49:22:06 – 00:49:50:16
Cool. Thank you. I think we’ve got a time for a couple more of you guys are happy. So, so, look, can organisations start with one use case, or does this require an enterprise wide data transformation? I can take that one. So absolutely you can start with one use case. Actually that is the recommended place to start because trying to fix every data issue across the enterprise or business unit teams can be overwhelming.

00:49:50:16 – 00:50:14:19
So that practical approach is to pick a high value use case for a credit model or a customer service agent, and ensure that trusted data foundation for the data that the use case depends on is is of quality. So yeah, definitely start with a small piece and go from there.

00:50:14:21 – 00:50:41:15
And then just quickly what types of data should be assessed. First the AI readiness. I can take that one. And so I’ll start with that. So like I said you want to start with data that enables a high value or high risk use case. So that can be kind of mentioned earlier structured data and data. But for structured data for example it could be customer transaction product or operational data.

00:50:41:21 – 00:51:14:19
It can also involve unstructured data like policies, contracts, case notes, knowledge articles. And so for structured data we can assess that data quality, relevance, duplicated records and coverage of representation. For unstructured data, the focus is often around the relevancy and the safety use aspects of the data. So for example, looking into your documents topics, recency, the right to access PII presence, any sensitive sensitivity that shouldn’t be allowed to be used.

00:51:14:21 – 00:51:39:03
And and of course whether the context right. The endless data based on the context, whether they’re fit for your AI use case. Wonderful. And then just a quick question. I think you’ve answered some of it within that last question. Actually, it’s about the pragmatics of how it works. So how does the AI readiness assessment work? Yeah. Glad you’re thinking ahead to our question.

00:51:39:05 – 00:51:58:07
So it often depends. We often start with the scoping session. Right. Because it really depends on like I said the context in the use case. So we start with the discussion around what are you looking to do with this data. Right. What is the use case. And then we will often have to in terms of like sending data.

00:51:58:08 – 00:52:22:04
I saw this question coming in in terms of sending data, yes, we will need access to the data sets to that you’ll that will be reviewing. And then from there on we usually have a small discovery session. Right. To really understand the context behind the data and what you’re intending to achieve. And before we come out with a assessment report and a playback session to you.

00:52:22:11 – 00:52:38:21
Typically the process depends on right the scope and the complexity of the data sets. And you know how significant that is, how many data sets and the use case specifically. But typically the the average time is about two weeks.

00:52:38:23 – 00:53:11:26
Can we go to the next slide. So there’s a contact information if people do want to follow up. And I want to add that for our existing clients we are able to guide them through running some of these assessments with understandable restrictions around sensitive data, data transfer requirements, etc.. So it’s not always the case that we can work with you on what makes more sense for you.

00:53:11:26 – 00:53:48:27
And then the second thing is, as you can imagine, experience sits on a treasure trove of very sensitive information. So the retention, the way that we handle the data is definitely up to the highest standards. And there’s even individual use contracts that would be signed for each of these, each of these evaluations. So if you have any concerns about that, our are requirement to use the data for that initial assessment.

00:53:48:27 – 00:54:14:17
Definitely reach out. And this email address here can route you to any of the relevant experts that might be able to help you with the follow up questions. That’s brilliant. Thank you. Yeah. And yeah, I think with one minute left we have had one more question. Does anyone want to have a quick hazard. We all out with the last question.

00:54:14:20 – 00:54:44:13
Wonderful. Let me have a quick look. So where do organisations typically underestimate the challenge of becoming AI ready. And what is the first practical step they should take to improve the quality and trustworthiness of their data? Yes, I can take that one very quickly. We’ve been speaking with over 60 different clients as the market research for this, and I think across the board, every industry, every region, we here over and over that people are finding out data.

00:54:44:15 – 00:55:13:29
That’s both the data that used for training, the reference data that your AI agents or chatbots uses in the Rag, and also metadata. The governance side is just not ready. So they’re seeing they literally have to turn models off because it’s not ready for the production. And I think people are realizing, going back to strengthening that data foundation is what’s needed.

00:55:13:29 – 00:55:37:03
And the first particle step back to what and Tina suggested really pick one use case that you want to activate, and we’re happy to help you look at where your data set were, some quick, pragmatic steps that you can take to get your data to make sure that use case come through. So start small and we’re happy to help here.

00:55:37:05 – 00:55:41:23
Thank you all for such active participation too.

00:55:41:26 – 00:56:01:15
Absolutely, absolutely. Just to that, thanks to our speakers and to the audience. Thanks for hanging on. There will be a survey pop up when you close down the webinar. Please do respond to that. It really helps us to shape future activity, and questions around slides and recordings will be on touch in the coming days and weeks with all registrants.

00:56:01:15 – 00:56:05:16
So thank you very much everyone and we hope you have a great rest of the day.

00:56:05:19 – 00:56:18:09

How we can help

Built for environments where processes and decisions carry real consequences, Aperture Data Studio turns fragmented data into trusted data for AI and business critical use cases. It embeds data quality, governance, and enrichment directly into workflows across complex, federated technology stacks so trust is enforced in execution, not assumed.

Find out more

Build trusted data that drives real business outcomes. Our experts can help you get started.

Speak to an expert