Idempotent Ingress
Inbound webhooks are validated, normalized, and associated with an idempotency key so repeated delivery does not create duplicate work.
Case Study · Automation / Platform Engineering
A public reference implementation for reliable API- and webhook-driven operations workflows, built to demonstrate the difference between “automation that runs” and automation that can be observed, retried, approved, audited, and recovered.
The Problem
Many workflow systems begin as useful scripts or low-code automations and then quietly become business-critical. At that point duplicate events, transient API failures, partial execution, unclear workflow state, missing audit history, and risky actions without review become operational problems.
The goal of this project was to model a more durable pattern: accept work safely, make state explicit, isolate external integrations, recover from transient failures, pause for human approval when needed, and preserve enough evidence to reconstruct what happened later.
Architecture
Inbound webhooks are validated, normalized, and associated with an idempotency key so repeated delivery does not create duplicate work.
Progress moves through persisted states such as received, validated, queued, processing, awaiting approval, succeeded, retry scheduled, failed, and dead letter.
Work is separated from request ingestion so execution can be retried, observed, and scaled independently rather than relying on one synchronous request path.
High-risk actions can enter an explicit approval state instead of relying on informal email or chat approval outside the workflow record.
Provider-specific APIs live behind adapters, reducing coupling between business workflow logic and external service behavior.
Meaningful transitions are recorded, while unrecoverable work is moved into a dead-letter path instead of disappearing into logs.
Reliability Engineering
Transient failures can be retried with bounded backoff, while validation and permanent business-rule failures are treated differently. Retrying everything blindly simply turns failures into noise.
The local demo has been exercised through duplicate and idempotent delivery so repeated external events return existing workflow results instead of creating parallel execution.
The runtime has been tested through human approval, transient retry, dead-letter handling, and audit-history flows rather than only a successful happy path.
The repository includes reliability and recovery documentation, an incident runbook, an example incident scenario, a simulated postmortem, and repeatable incident drills.
Observability
The reference stack exposes workflow metrics through Prometheus and includes a provisioned Grafana dashboard covering workflow totals, success and dead-letter counts, retries, queue depth, approval wait, adapter latency, and success ratio.
The repository also includes alert rules around backlog, queue age, retries, dead letters, and success ratio so observability is treated as part of the operating model rather than a screenshot added at the end.
Production-Minded Delivery
The production-style Docker path uses multi-stage builds, compiled JavaScript, non-root application containers, health checks, restart policies, externalized credentials, read-only application filesystems, and reduced Linux capabilities.
GitHub Actions validates builds, runs tests, builds the container image, and checks the production Compose configuration on relevant pushes and pull requests.
A reference AWS Terraform deployment maps the application into VPC/subnet design, ECS/Fargate services, RDS PostgreSQL, ECR, an Application Load Balancer, CloudWatch, Secrets Manager, and a migration task.
Why I Built It
This project is deliberately public and portfolio-safe. It contains no production credentials, customer data, proprietary business logic, private webhook URLs, or commercial workflow configuration.
What it does expose is the architecture and reasoning: how I think about duplicate work, workflow state, approvals, retry policy, observability, containerization, deployment, incident handling, and the operational consequences of design choices.
Operating Principle
A fast automation that cannot explain its state, recover safely, prevent duplication, or show who approved a risky action eventually becomes another source of incidents.
The platform is a working demonstration of the opposite approach: automation as an operational system with explicit state, guardrails, recovery paths, and evidence.
Explore the Work
Source, architecture, examples, runtime guidance, and project documentation.
Open GitHub →See this project in the context of the broader technical portfolio.
Enter the lab →See how the same operational thinking applies to connected systems and field technology.
Read case study →Operational Origins
Independent workflow work across Make, REST APIs, webhooks, persistent data, AI-assisted processing, and publishing raised practical questions: which step actually completed, what needs review, what can be retried, and what happens when credentials or downstream services stop working?
A successful request does not prove the downstream action completed. Preserve external identifiers and reconcile the result before marking the workflow done.
A timeout may justify a bounded retry. An authorization failure needs credential repair; invalid content needs correction. Use idempotency and reconciliation before repeating an external write.
Consequential publication or business actions need a visible decision boundary. The evidence, approver, and approved content version should travel with execution.
The public platform is a generalized implementation of these patterns, distinct from commercial workflows. It demonstrates design and recovery behavior without claiming measured business outcomes for every integration.
See the shared systems principles →