Designing Recoverable Background Workflows With Explicit Execution History

https://hackernoon.imgix.net/images/nA3qxOwEzkRboii6qMmGhKNjROp1-7sc3a8b.png

A failed workflow is rarely a blank slate. Knowing what already happened is the starting point for recovery. Illustration by Raideria; conceptual example, not a product screenshot.

A daily report does not arrive. The process exited with an error. Somewhere before that error, it may have downloaded data, written rows, generated a file, or sent a request to another system.

You can restart it. But what, exactly, are you restarting?

That question is where a background job becomes an operational problem. Starting Python is straightforward. Understanding a partially completed workflow—and deciding which work can safely run again—takes more structure.

An illustrative report pipeline: orders and inventory succeed, report generation fails, and publication waits.

I lead Raideria, the company behind Dagychu, a self-hosted platform for running and operating jobs and pipelines. This article explains the execution model behind it through an illustrative reporting workflow. It is not a customer case study...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more