Pipelines · 12 Jun 2026
Late data, retries, and the boring parts of a pipeline
Atalaya Digital
Most data projects die in the gap between "the DAG ran green yesterday" and "finance cannot close the month because Tuesday's file showed up on Thursday."
At Atalaya we treat that gap as the actual work. Ingest and orchestration are not a pretty graph. They are a contract: what arrives, when it is late, what we retry, and who gets paged.
What "late" actually means
Late is not a boolean. A plant file that lands 20 minutes after the cut is not the same as a consent feed that never arrived. We model three states:
- On time — process, mark the partition done.
- Late but usable — reprocess the partition, keep the previous result until the new one is valid.
- Missing — do not silently skip. Fail the downstream board or serve a stale watermark, on purpose.
If you only have "success" and "failed", operators invent folklore. Folklore does not survive a holiday week.
Retries are a design choice
Blind retries hide poison. We retry transport and timeouts. We do not retry schema drift or a 400 from a source we own.
A useful retry policy says:
- which task is idempotent
- how long we wait before the next attempt
- when we stop and open a ticket instead of looping until the SLA is already dead
Airflow, Cloud Composer, and Cloud Functions can all do this. The tool is not the point. The point is that a human can read the policy at 2am.
Observe the contract, not the CPU
We alert on freshness, row-count bands, and "this source did not produce a file." We do not alert on "the VM was busy."
If the board is wired to your warehouse, the watermark belongs in the same database. A green DAG and a stale table is a lie. The board should show the lie.
That is the boring part. It is also why a pipeline ships.