The implications of Artificial Intelligence (AI) are vast and varied, and it’s becoming increasingly clear that there’s a lot we still don’t understand about this technology. This is particularly true when it comes to matters of security and privacy. AI today is like the Internet in the late 80s – it’s a field of fundamental research, latent potential, and academic usage, but not yet ready for public consumption.
“Compromised” Foundation Models
The so-called “open” models are anything but open. While some vendors may boast about their degrees of openness, none provide access to the training data sets or their manifests or lineage. This opacity means that users don’t have the ability to verify or validate the extent of data pollution with respect to intellectual property, copyrights, as well as potentially illegal content.
Moreover, the absence of a manifest for the training data sets eliminates the possibility of verifying or validating the non-existent malicious content. Nefarious actors, including state-sponsored ones, can plant trojan horse content across the web that the models ingest during their training. This can lead to unpredictable and potentially malicious side effects. Once a model is compromised, the only option is to destroy it.
“Porous” Security
Generative AI models present a significant security risk. They are the ultimate security honeypots as all data is ingested into one container. New classes and categories of attack vectors arise in the era of AI. The industry is yet to fully understand the implications of securing these models from cyber threats and how these models can be used as tools by cyberthreat actors.
There are various techniques that threat actors can use to compromise these models. These include malicious prompt injection to poison the index, data poisoning to corrupt the weights, embedding attacks to pull rich data out of the embeddings, and membership inference to determine whether certain data was in the training set. The threat of embedded state-sponsored cyber activity via trojan horses and more is real and present.
“Leaky” Privacy
AI models are helpful due to the data sets they are trained on. However, the indiscriminate ingestion of data at scale creates unprecedented privacy risks for individuals and the public. In the era of AI, privacy has become a societal concern, and regulations that primarily address individual data rights are inadequate.
Static data isn’t the only concern. It’s also crucial that dynamic conversational prompts be treated as intellectual property to be protected and safeguarded. If you are a consumer, you want your prompts that direct creative activity not to be used to train the model or otherwise shared with other consumers of the model. If you are an employee working with a model to deliver business outcomes, your prompts should be confidential, with a secure audit trail.
What happens next?
We are dealing with technology unlike anything we’ve seen before. AI exhibits emergent, latent behavior at scale, and our traditional approaches to security, privacy, and confidentiality are no longer adequate. Industry leaders are throwing caution to the wind, forcing regulators and policymakers to step in. It’s clear that we need to reassess and possibly overhaul our approach to these issues in the era of AI.
