Build a serverless PII redaction pipeline with Amazon Bedrock Data Automation | Amazon Web Services

https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/11/ML-20554-featured-image.png

Organizations that process thousands of scanned documents daily, including medical forms, insurance claims, and financial records, face a recurring compliance need: personally identifiable information (PII) redaction before documents are shared with third parties or processed downstream.

Manual redaction doesn’t scale: It consumes staff hours, introduces human error, and creates compliance exposure. Redaction is also a precision problem, in addition to a detection problem. A single page can contain multiple names, dates, and addresses where only some are sensitive to the use case. Traditional redaction approaches pair optical character recognition (OCR) with pattern matching or custom machine learning (ML) models. However, these approaches have limitations when text is degraded, cannot easily express field-level business logic, and require ML expertise to build and retrain custom models as document formats change.

In this post, we demonstrate how to automate end-to-end PII detection and redaction from documents and images at scale on AWS. We...

Copyright of this story solely belongs to aws.amazon.com. To see the full text click HERE

Read more