Building a Parallel Distributed Audio Watermarking Pipeline on AWS Lambda

https://hackernoon.imgix.net/images/BQ1MAQ1DgPed8YCQAiqJHGK2Ylf1-s2f33rt.jpeg

Invisible watermarking sounds like a signal-processing problem. It is, until someone uploads a two-hour podcast and expects it back before their coffee cools. Then it becomes a distributed-systems problem.

At Adori, I work on OrigID, a platform that embeds an inaudible ownership watermark into audio and video so a creator can later prove where a file came from. The encoder is CPU-heavy digital signal processing, and its cost grows linearly with the length of the file. On one worker, a long episode takes proportionally long. On AWS Lambda, a single invocation can't run longer than 15 minutes at all.

This post walks through the pipeline we built to make processing time roughly independent of file length: split the file, watermark the pieces in parallel on Lambda, and stitch the result back together. The fan-out is the easy half. Most of this post is about the parts that aren't: where you're...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more