The problem
A critical data pipeline running roughly 27,000 executions per week kept breaching its SLA, with delays reaching 6 hours. Downstream, that meant people making the morning’s decisions on yesterday’s numbers without knowing it.
Recurring failure at that volume is rarely one bug. It is usually a system that has no way of telling you it is unwell.
What I changed
Three things, in order of how much they mattered:
- Monitoring. Making failures visible immediately rather than discovered downstream by a confused user.
- Error handling. A single bad record degrades one row instead of taking down a run.
- Operational process. Who gets told, what they do, and how a failed run gets recovered without a person reconstructing state by hand.
Where it landed
100% SLA compliance. The delays stopped.
The project took 1st Place at the company level, then went on to take 1st Place at the group level.