Why I Don't Use AI to Remove PII Before Sending Data to AI

https://hackernoon.imgix.net/images/OLKoObtdw3XwgXasjutZBfeZUoI3-q703740.png

The strange thing about AI-based redaction is that the redaction model still has to see the original data.

There is one job in an AI pipeline that I have become increasingly uncomfortable giving to another AI model.

Removing sensitive data.

Think about the usual setup.

A user pastes a support ticket, medical note, application log, resume, financial document, or customer conversation into your application.

Before sending it to the main LLM, you want to clean it.

So we add another model first.

The first model receives the original text and gets a prompt like:

Remove all personally identifiable information from this text.Replace names, phone numbers, email addresses, account numbers,credit card numbers and other sensitive information with placeholders.

It returns a cleaned version.

Then we send that version to the model that actually performs the task.

At first, this looks like a privacy feature.

Then you draw the data flow.

User...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE