Webinar:

Keeping sensitive SAP data lean and secure.

.

 
Play video
Good morning, everyone. Welcome to our SAUG webinar. I hope everyone is well. Thanks to all our members who have joined us in good numbers today. Thanks to Epius Labs, of course, for their ongoing support of SAG and for making this webinar possible. Today's topic is on keeping your SAP or sensitive SAP data lean and secure. Now before we start, just a few details. This webinar is being recorded. The recording and a copy of the presentation or links to those will be made available on the SAEG website shortly, and we'll notify everyone when it's ready. Everyone is on mute, but you can ask questions at any time by using the question function on the right of your screen, and we'll go through questions at the end. There'll also be contact details on the slide, so if you'd like to, you know, put forward any follow on questions or you can even contact us, if you think of anything after the session. And also, we always welcome feedback. So if today's topic makes you think of other topic areas you'd like to see, please let us know also in the, questions section. Now, I'd like to introduce our presenters, person who's well known to everyone at SAUG. It's Daniel Parker, who's solutions director, EPU Slabs. And before we start, I just wanted to say this is actually very timely, session because, some of you might have been on the on our asset management, session yesterday, and there was talk about, using, the new, S4HANA as, asset management, in a nonproductive environment. And a lot of these people who had, critical infrastructure or critical assets. So, I think this, today's topic is very timely. So, thank you, Daniel. Looking forward to it. Excellent. Thank you, Michael. Welcome everyone. Yeah. So as mentioned today, we're going to have a talk about, sort of data, data privacy, and, anonymizing securing data and also helping to, to lean some of that, sensitive data in your SAP systems. I'll just start first a little bit about, FUSE Labs. So if you're not aware of us, we're an SAP focused company, part of what we call Group Elephants, which is a group of a larger group of organizations, worldwide. So focusing around, you know, SAP, type delivery and, SAP use labs in particular focus around R and D, and we roll a lot of our revenue back into ongoing R and D. So building solutions to help our customers manage their SAP systems, whether that's data copy, you know, data scrambling type concepts that we'll talk about today. And yeah, it's all about, building those solution sets out, helping people manage their landscapes, and look after their SAP systems, easier. Okay. So yeah, today we're going to talk a bit about data privacy, and sort of where it's going in Australia, and just some thoughts in terms of what we see elsewhere in the world within FUSE Labs, and how we have helped, you know, build solutions and build products to, for our customers globally to get ahead in terms of their, privacy and legislation requirements in other regions. So as you would be aware, if you've been anywhere near a newspaper in the last sort of twelve, eighteen months, there's been some quite high profile events across Australia around data privacy and data breach. So, Optus, Medibank, Latitude are three, three of the high profile ones. So quite large breaches, large loss of data, large loss of identifiable information with large impacts. So we see those being reported and see those in place because of a notifiable, data breach scheme that, was put in place back in February twenty eighteen. So this is where certain organizations of a particular size and a particular impact of the data loss for the breach need to report that to the OAIC, Office of Australian Information Commissioner. So there's a few gaps and a few holes in terms of that breach scheme, so not everything is reported, but you do still see a good spectrum of data and a good view of what's happening in terms of, privacy issues, data issues, breaches, and leaks. The OAIC. If you haven't come across them before, just do a quick Google search and look them up. They, collate this breach data and they release reports a couple of times a year. So the most recent one, March, twenty three, is covering the back half of last year. So what we can see from that is the notifications in general were up, you know, just under forty percent of those affect more than a hundred people. So each of the notified breaches, covering, issues or problems for fairly large volumes of people. So then not everything is just the scale of a of an Optus or a Latitude, but everything has some kind of impact. What's interesting though is that malicious or criminal act is the source of seventy percent of those breaches, which is a forty percent increase from the previous recording period. So the numbers are going up. If you look at some of the longer trends they have on their website as well, they are generally going up as we're seeing more things reported, which implies there's more things, occurring. And most of the time, contact information is the most common type of PII that's being lost. So, names, addresses, telephone, email, things like driver's licenses or, or visa or, passport information, which is certainly what was wrapped up around, Optus and Medibank and Latitude. So the general environment that we're seeing from what we do get reported in Australia is a general upward trend and more of an angle towards the listless or criminal act as the source of data. So what still makes up a regional component is just, you know, general system error or user error, where maybe something is emailed to the wrong person or a file was sent to the wrong location. So that still makes up a recent component, because it is interesting that, these other activities are driving, a lot of the volume. Now what we're beginning to see more recently from a government point of view and a direction is a little bit more stronger language and a bit more stronger driving around this. So we don't have, you know, privacy laws that are pushing to the limits of what we see in California and in Europe. But the government is beginning to talk that way. So these are hack the hackers. So this is an initiative where there's a team or a department which will, proactively go out to known hacking organizations and try and disrupt them and take them down. So, you know, on the front foot and also the cyber, group with this aim of helping to make Australia the most secure talking more and more in Australia from a federal level is to be more stringent, more proactive, in terms of, breaches and privacy, of data. So what might this mean to your organization or how you manage or look after data? So the story that we've seen over the years within AppEase labs across our global teams is, more and more stringent, legislation within different countries to look after and manage personally identifiable, personally identifiable and sensitive data. So that same trend should follow in Australia. We see in New Zealand, they've stepped up in recent years some of their laws, and they're more stringent around how data should be managed in production. Thailand are bringing in something which is not unlike Europe's GDPR in terms of management of production data. So there's a growing swell in terms of the future direction, there's more stringent legislation to manage and look after, the PII you have in your systems, as opposed to less and slacker. So the big question then, in terms of where we initially start helping organizations that are trying to get in front of this or trying to align themselves with legislation is what identifiable information do I have? So you may have been running an SAP system for fifteen odd years. That's not uncommon in Australia. Started as maybe a four six something or other ECC. Now looking at S4HANA, and you've brought that data all the way through. So you're gonna have quite old and aged information in terms of transactional line items, as well as, you know, customer vendor business partner type information. So, so what do you have? What is the footprint of that? What constitutes PII in terms of your organization? So when we start these kinds of conversations, the obvious point is around, you know, the key building blocks of data SAPs, so things like employee, vendor, business partner, and customer. So these are the most common areas that people think about when they're thinking of their PII and their scope of, sensitive data. So you may be running HR payroll or maybe just a HR mini, so you've got your local employee information that could be linked to a business partner, your employees may be vendors. So there you have three sets of objects where you've got employee name, address, telephone number, email, those kind of concepts, as well as bank account information. If you're running an external facing system, you've got customers or your end customer, so they could be individuals. So they'll have, you know, name, addresses, telephone number, bank details, credit cards, that kind of information again. So it's very quick to build up a small bundle of, of information that is, sensitive, and you need to think about how you manage that or how you control that. And these are the sort of the key areas or the key building blocks that, things like GDPR in Europe, new legislation in the US, in South Africa, Thailand. These are the areas they're focusing in on. So do you have a right to have that aged customer data? How do you manage it? How do you look after it? Now the conundrum in an SAP system is it's quite easy to say, well, I want to anonymize names, but then underneath that you'll have dozens and dozens of tables and it becomes very, very complex. So what is a simple concept in terms of, okay, I don't wanna scramble my employee names, scramble my employee addresses, turns into an exercise of identifying hundreds of tables, hundreds of fields, that becomes quite complex. And then the conundrum there is that these things are all connected. So if I'm an employee, Daniel Parker, I want to rename it to John Doe, I might be linked to a business partner, to a customer, to a vendor. So if I manage these in isolation within the SAP system, I can end up with jumbled and a mess of data. So it's understanding the connections and the links and the relationships between all of these, as well as identifying it is, is a key part of the puzzle that we typically help customers on. So the way we do that in, FPS Labs with what we call our data sync manager solution, which has, data privacy components built in that I'll talk about today. We have a concept of a business object definition. So these are the customer object, the business partner, and the zero customer from a BW concept. So these business objects we've defined and mapped and built in our solution, they tell us the structure of the tables underneath each of those. So you don't need to know what tables are part of customer or part of business partner. We've designed that and built that into an object. And then linking between those is the red line, which we call an integrity map. So when we start to look at these data objects and start to think about, well, I need to manage the name, I need to manage the email, or so forth, I have the integrity map to link them together. So I have a series of integrity maps defining names, phone number, email, and these link together with the business objects. So if I need to perform an operation to anonymize a customer or a business partner, I can do that consistently across these objects and even between different systems because of the way we've built this up and modeled this. And this is extendable as well. So part of that discovery and finding what you've got in your system is also about working out, what z tables, what z fields you have. So any part of this conversation I'll step into soon in terms of managing test data or managing production data always starts with some kind of discovery and a workshop process. So this is where we bring in technology that allows us to see table field, table structures, and map them back to our definitions and help understand where you've got data and what it means in terms of what's going on. So what that looks like practically is we run an analysis program, it looks at the data dictionary and the structure of the system, we're not looking at real data, we're pulling out tables and fields, and data elements, and then mapping those saying these contain data, I don't see the data, and they, are potentially PII. So from that we can build a, a report, And the whole idea is to give you scope and idea in terms of what data you have and what you might need to do with it. So collaborative workshop followed by some in-depth system analysis, running an analysis program, getting metadata about the system, running some workshops to understand, what your challenges are. Maybe it's just test data. Maybe it's, removal of a legacy company code. Maybe it's removal of legacy customer or business partner type data, depending on what you're trying to solve. And then we do further analysis and prepare a final report. So the idea is that we're working off, what we've leveraged from running multiple of these types of projects around the world and building a solution set up and then, building out for you a report that helps you understand your, you know, what data you have, your general privacy requirements for that data, how it links, how it flows, the definitions of those objects, and how they're aligned, and also a breakdown of tables and fields with that cement information in the SAP environment. So this is a standalone piece of work is very useful, we find, across the globe to help customers understand exactly where they are. So what's their start point? So rather than starting from that idea of I need to anonymize employee data, we're starting from here is the types of data I have in my system, and now these are different ways I can apply privacy or reduction to it. So that steps to the next point, which is in terms of the product sets that we have to actually achieve these things. So I'll talk about two basic use cases, but what I'm looking at here today is our test data masking solution and also our production redaction solution. So the test data masking is the most common use component, within this region, but we're seeing the data redaction solution become more of an interest in New Zealand with the current changes they've got underway as people begin to think about and discover how they need to adapt to that legislation. It's also very strong in Europe, and we've got a case study we'll talk about in terms of how that works. So these, these two angles and ways of managing the data give us, we're both part of DataSync Manager, both part of our data privacy solution set, that they allow us to do test data masking and also production data compliance. So the test data masking is fairly common. They're fairly obvious. We've built a test system. We're going to anonymize the data, so it's safe to use, for testing needs, but it's not a replication of my production data. So it's no longer a source for extractive data that could be part of the breach with real information. On the production data compliance side, this is about removing or leaning down the volume of data I've got. So maybe I'm no longer wanting to keep particular company code, particular subset of customers or vendors, so I can redact or de identify that information. So let's first of all have a look at the test data side of things, so data scrambling solution, masking sensitive data in non production systems, So that consistently masking. So if we think back to that image before with the business object definitions, so we wanna consistently scramble data between, you know, in different manners. There's a bunch of out of the box content, which I'll show in a moment on a quick demo. And that is always then extended or changed based on a workshop and that scoping session to find out exactly what PII information you have. And it's all customizable, so you can build it into these ed tables and ed fields. What we'll show is masking a client, but there's different ways to do it. You can bring and scramble the data as part of the copy, and there's no need to manage the data externally, so we're always running it in the system, so we don't have to pull the data out, manage it, and put it back in. It's running locally in the SAP system, so you then end up multiple copies of the data as part of that process. So let's, next slide is the demo. So yeah. So how do we do that? Randomization of synthetic values. So when we mask, we're saying, okay, I want to change the name. I'm going to pick a random name from a list that I've provided, so those lists are editable. Same for address, I want to change the street address, I'll pick a random street address from the list. And then we substitute those in for the real values, and we do it in such a way that you don't have pattern breaking, so it helps with, anonymization. So the name Fred doesn't always become Barry. There's variations. Sometimes John, sometimes it's Barry. So you can't do a one to one map back and, and, and effectively unscramble the data. So that's a key point in terms of how it works. Now we'll have a quick look at the demo. So I've just jumped into a an example system here. What I'll do is I'll run a an in place client mask just to give you an idea for what it does and how it works. But essentially we made up of four, four building blocks. I'll execute a execute we define a policy that meets the requirements for the scrambling you want to do. Policies are made up of rules, so rules are groupings of things, so I'll have all the information for like customer name or customer bank account in a rule, and then integrity maps is the next layer down. The integrity map is the tables and the fields, so we define a BUT triple zero and the name field in an integrity map. And then lastly is the transformation. The transformation is how we change the data. So all of these things, we can't contain pre delivered content, which we'll see when we jump onto here. There we go, that one worked. Perfect, so now we're executing what would be an in place scrambling run, so this may be a system I've already refreshed and I'm going to run the scrambling rules on it. So I can see the policy I've built, just a little example one, and this is made up of a bunch of pre delivered rules and content that I have here. So I've got things like address, banking detail, name, search terms, customer tax information. So this is these are all delivered pieces of content. So this is part of the reference of library set of information we have. Underneath these, these are, you know, the table and field definitions and the transformations. So, essentially what we have is sort of a low code, no code solution. So you should be able to put together the building blocks of what we have and then extend things. So create your own integrity maps, create your own rules, and bring things together. The transformations, so the bits that are doing the work, there's a large library of those. So we have transformations for changing an address, changing a name, changing bank accounts, and there's a lot of functionality built into those, so it's generally not common to have to build transformations. You're usually defining your tables and fields and defining how you want to adjust the data and bringing that together. Once you have that policy, you'd simply just execute it, then we can run the job, specify how many processes to use, and then we can kick it off and then that runs in the system and applies the scrambling, to the solution. There's one I ran earlier, just in the interest of time. And we can see here what we do is we work through each of the business objects, that are part of the, of the policy. So these, again, back to those things where we've mapped out and defined what it is to be a vendor or payments or customer, and then we look at the key. So this is a very, very small system just to help with demo run times. So there's nine vendors and ninety four accounting documents. So these are the items where I've analyzed and tried to determine if we're going to do scrambling. We create a calculation, so this is what we think we're going to apply, and then we apply an update. And we can turn on auditing as well, generally not done for production usage, but useful in this situation. So here I can see what my old value and my new values are as I apply the scrambling in this system. So that's just helpful in terms of analysis and troubleshooting in a way to see what's going on. And then what we're doing is we're changing the field in the database. So this isn't, you know, display time or run time, we're actually editing the value in the database, so then the new values, the anonymized values, are what's available going forward. So that's how we're applying scrambling to test systems. So we haven't leaned out or reduced the data in any way, which simply is just saying in my test environment, I need to anonymize all that PII, so I'm reducing the total footprint of PII in my landscape, limiting it just to production. The next, quick success story in terms of that, so we've worked in recent years with a utilities company, based in Australia. They had a more customized, need, so this is the, you know, the far end of what's possible. Usually we work more back at standard content, but they were already running a scrambling process across their non SAP systems, and we needed to build a solution to match that. So we actually created some transformations, which followed the same algorithm that we're doing to anonymize characters in their other, systems. And then that kept us aligned, and then we could scramble in about six hours forty eight and thirty nine million keys across a nest four and a CRM system. So that's a large set of basically business partner and customer records because they're utilities based. And they were also using it for data copy. So as they were doing a client copy with our client sync solution, they're applying the scrambling as well. So this meant before the data was put into their test systems, it was already scrambled and already anonymized. But if they built a system without that client copy process, they could also scramble, in place. So that was quite a good solution for them. It gave them protection of the PII, but kept them aligned with their existing process across the rest of their larger than SAP footprint. Okay, so that's, you know, a thing about in terms of anonymizing and protecting test data, so just thinking that our end goal is we want to minimize our footprint and our volume of production data, So back in my non production systems, I want to randomize and anonymize, but still keep it usable. So that takes care of some of that footprint. Now on the production system side, this is where we're seeing, you know, direction and traction, especially out of Europe and more so in California in terms of redacting or dropping the volume of PII in a production system. So the solution sets behind this are a little bit different, a little bit more targeted. And they, they come in three components. So we have what's called, don't disclose, redact, and retain. Each of these have a sort of a set, need or set thing that they fulfill. Just disclose is probably the most closely aligned with GDPR. There's a concept of a, subject access request where it's myself as a customer of your retail company, I can, I can ask, okay, what information do you have about me? What PII do you store? And this is where the disclose can be used to present an encrypted PDF report back in a simple manner. They give you a sort of a footprint report of what access they have, what data you may have on them. So that's a requirement of GDPR. So we don't see that as much in other legislations, but it is useful to know that exists as a way to sort of fill that, that access request, type requirement. The key component here though, is data redact. So this is how we can, remove personally identifiable data from the SAP system without removing the complete record. So the conundrum we found as, things like GDPR came to the fore and also in California, is organizations needed to reduce the volume of production identifiable data, but they couldn't necessarily use standard SAP archiving, because the archiving is simply taking the data out, putting it somewhere else, and also the records may not always be accessible to data archiving. So they were looking for a way to de identify it without necessarily removing it. So this is where Redact was built as a solution, so we can, change fields, drop out identifiable information, so maybe drop out phone numbers, drop out emails, change the names, but still keep the shell of the record. So it's still available from a business data, perspective, but it has no longer any identifiable information. So it's leaning that down and reducing the risk. And then lastly, retain is where we do ongoing checks. So building out business rules saying, all customers that haven't transacted for say five years, we're going to, you know, we're going to set them up for a manual, redact run. So you might do that every month, look for anyone that hasn't fulfilled some kind of business criteria, and then they become a, an item potentially going to be, redacted. So if we watch a little bit of an animation here, so the way it flows, and I'll show this in a demo as well, is we start with the disclose, the access request. So it's someone saying, you know, what data do you have on me? And it discloses an index process. So we've looked at the data across the systems, and we've built a view of it. And then we can drop that into a footprint report. So we know that record exists across a variety of systems, and then we can select that and say, okay, so what do we actually have? And then build that out into a list of PII as a simple report, and this is where it can be put in an encrypted and password protected PDF and sent out. So we know for Howland we've got occurrences of your data in five or six spots and how it works. So this fulfills some of those specific requirements for GDPR, which is that data subject access request, and then from that we can feed it back into Redact if the person then says okay, you no longer have a right to have my data. And I'll show you those as a quick walk through in the moment, and then the redaction is actually dropping the information, removing the PII, so it's no longer stored in that production system, but not impacting the data model and the integrity of the data in that solup5hhr7ueck,{""english_name"":""English""
Recorded:  April 2023
 

Do you use live data in test and development environments?

According to Gartner, more than 80% of businesses use sensitive data in non-production systems for development, testing, quality assurance, and pilot projects.  Sensitive data can range from your CEO's salary to information on the prices of your products, discounts, or customer data.

Do you have legacy Personally Identifiable Information (PII) data in your production environments? 

In Australia alone, malicious or criminal attacks accounted for 70% of all notifications reported to the Office of the Australian Information Commissioner (OAIC) from July to December 2022. This resulted in a 41% increase in data breaches, with an influx of reports of attacks from health service providers and the finance industry disclosed in the Notifiable Data Breaches (NDB) report released on 1 March 2023.

Along with the government’s efforts to improve cyber security to make Australia the safest place to connect online, compliance with data privacy laws is more crucial now than ever for organisations to protect their customers' personal data, build and maintain trust, and avoid significant financial and reputational harm. Multiple privacy codes continuously put pressure on organisations in Australia to limit and manage PII in production environments. 

Storing data in archives and manually updating data is not enough. Find out more about a comprehensive and secure approach to identify, manage, redact, and reduce the legacy PII footprint across the entire organisational landscape, including production environments.

In this webinar replay, you will learn more about: 

  • Taking proactive measures to safeguard and ensure the security of your data.
  • Managing compliance and overcoming the difficulties associated with safeguarding confidential information.
  • Customising and anonymising sensitive data.
  • Implementing best practices in security and complying with regulations such as GDPR.

Daniel Parker
Solutions Director at EPI-USE Labs

With more than 20 years of SAP experience, Daniel specialises in data copy automation and data security. With a strong Basis background Daniel has led technical teams around the SAP Lifecycle of Implementations, Upgrades, Conversions & Migrations. He leads an experienced consulting team and delivers a variety of SAP landscape optimisation solutions to organisations in the Asia Pacific region.

EPI-USE Labs needs the contact information you provide to us, using the form embedded in this video, to contact you about our products and services. You may unsubscribe from these communications at anytime. For information on how to unsubscribe, as well as our privacy practices and commitment to protecting your privacy, check out our Privacy Policy. If you would like to receive emails from us, such as invitations, access to webinars and latest SAP insights, please remember to tick the box.