In a world where artificial intelligence plays an increasingly significant role, many organizations are hopping on board the AI train by sending out information into AI-driven engines like Open AI's ChatGPT, Microsoft's Bing, and others. What they may not consider, however, is that this information is then stored by these external providers and used to further train the AI models.
The best way to ensure the protection of sensitive data is by utilizing automated data classification solutions, which effectively scan company data, locate sensitive information, and flag any potential risks.
In my recent chat with Chris Stapenhurst, Product Manager at information management technology provider Veritas, we explored potential risks using novel generative AI engines and how data classification can help prevent them.
Sending Data into AI Engines: A Tricky Business
The widespread use of AI engines for both personal and business use has resulted in a surge of data being shared with AI-based systems. Of course, this has numerous benefits and immense potential to decrease workload, simplify complex tasks, and help open our creative chakra. However, organizations and individuals often overlook the fact that the information they provide may end up stored and utilized by AI providers.
"This lack of awareness raises concerns about data privacy, as sensitive information can unknowingly be used to train AI models," Stapenhurst warns.
"It is, therefore, crucial for organizations to take control of the data they share and ensure its cleanliness before it reaches the cloud."
How Does Data Classification Work?
Data classification is a method of categorizing data based on its content to gain insights into its nature and potential risks.
"Data classification should ideally occur before data is sent to AI-based engines or cloud providers since the enormous volumes of information involved make manual inspection impractical once the data is out of an organization's control." Stapenhurst explains.
Automated classification engines play a crucial role in efficiently scanning and identifying the content of data, including contracts, documents, emails, and policy types.
"These engines offer insights into whether data contains personal information, market abuse references, or off-channel signaling," Stapenhurst adds.
Who Needs Data Classification, and Why?
The answer is simple: Everyone needs data classification. It is essential for almost all organizations, regardless of their size or industry.
"Many organizations primarily focus on capacity management or data management by repository, neglecting the comprehensive understanding of their data," Stapenhurst notes.




