Short answer: for a normal fintech app, an app-level hosted cron that calls an idempotent cleanup endpoint is the simplest starting point; put the actual deletion work on a queue when a batch can run for a while.
Cost and retention: what are you actually paying to keep?
The dominant cost is usually retention, not the timer. Every extra day of payment events, audit rows, and expired OTP records increases storage, indexes, backups, and the amount of data each cleanup query must inspect. A scheduler that fires every hour does not reduce that bill by itself. The change that moves it is deleting (or legally retaining) the right rows in bounded batches, then recording a durable audit result.
I model the job as a small state machine: select a tenant and a time window, claim a batch, delete with a stable boundary, and record the last completed boundary. A webhook retry can then repeat the request without repeating the effect. The catch is retention policy: if compliance requires seven years of transaction records, “cleanup” should target disposable derivatives, not the source ledger. Your mileage may vary by jurisdiction.
Keep detailed counts and failure reasons in the application database or logging system. Cron run output is limited to the first 4 KB, which is not an audit trail.
Ship it.
How should a fintech app compare Postgres pg_cron, app cron, and webhook schedulers?
Database-native scheduling, such as pg_cron, is compact and powerful when the work is SQL-only and the database owns the policy. It also ties the logic to that database: moving the cleanup to another store, adding an HTTP notification, or testing it outside the database becomes a migration project.
An app cron keeps the decision in application code and an HTTP handler. That is easier to review alongside authorization, tenant scoping, and HMAC verification. A hosted cron service does not run your code; it calls a public HTTP URL, so the endpoint must be internet-accessible. Put a narrow authenticated route at the edge and reject stale signatures.
| Option | Strength | Operational catch | Good fit |
|---|---|---|---|
| Postgres pg_cron | SQL close to the data | Database coupling and database-specific operations | One Postgres, SQL-only retention |
| AWS EventBridge Scheduler | Managed schedules and cloud integrations | More IAM and cloud configuration to operate | AWS-centric teams with many targets |
| GitHub Actions schedule | Familiar YAML and repository audit trail | Runner startup and overlap controls are your problem | Small, low-frequency maintenance jobs |
| Hosted app cron + endpoint | Application code, HTTP auth, easy queue handoff | Public endpoint and limited run output | Normal apps with mixed cleanup logic |
Infrai fits the last row when one key and one bill across backend capabilities matter, and when a plain REST call is preferable to installing another SDK. Its scheduling surface is intentionally narrow: a cron trigger calls the endpoint, while a queue can carry the expanded work. That is a workflow choice, not a claim that it replaces Airflow or Temporal.
Make retries boring: idempotency before delivery
Webhook delivery is at-least-once in practice. A timeout after the server commits is indistinguishable from a timeout before it commits, so the caller retries. Use a client-supplied idempotency key derived from the cleanup window, tenant, and batch boundary; store that key with the result and make the database mutation conditional on it.
The following sketch shows the shape of a cron registration and a queue handoff. It uses only documented scheduling paths and leaves the destructive SQL inside your own endpoint.
import os
import requests
BASE = os.environ["INFRAI_BASE_URL"].rstrip("/")
KEY = os.environ["INFRAI_API_KEY"]
headers = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
cron = requests.post(
f"{BASE}/cron/create",
headers=headers,
json={
"name": "fintech-cleanup",
"schedule": "0 * * * *",
"http_url": "https://app.example.com/internal/cleanup",
"method": "POST",
},
timeout=15,
)
cron.raise_for_status()
job = {"tenant_id": "t_42", "window_end": "2026-09-12T00:00:00Z", "idempotency_key": "t_42:2026-09-12T00:00Z"}
queued = requests.post(f"{BASE}/queue/publish", headers=headers, json=job, timeout=15)
queued.raise_for_status()
In production, handle HTTP 429 with exponential backoff and honor Retry-After. Check every response body, cap the cron-triggered work below the 900-second execution limit, and have a worker consume large batches. The standard queue is at-least-once, its FIFO deduplication window is five minutes, and messages are retained for at most 30 days; consumers still need idempotency. I once treated a timeout as proof that nothing had happened, then found the cleanup row already committed on the next retry. The fix was unglamorous: persist the idempotency key before doing the delete, make the delete conditional, and expose the stored result to the caller. That extra write is cheaper than explaining duplicate retention events to a compliance reviewer.
There is no native topic broadcast, debounce, or throttle. If one cleanup expands into per-tenant or per-table work, publish one message per destination or maintain separate queues. A single fan-out-and-join workflow belongs in a workflow engine.
What should you stop retaining when a cleanup fails?
Do not delete the checkpoint just because a run failed. Keep the last successful boundary, the idempotency key, and a compact error record; alert on the gap. If a cron is paused, missed triggers are not replayed automatically, and trigger timing can jitter by seconds. A manual trigger or a queue redrive should advance the same boundary, so recovery follows the normal path.
Choose pg_cron when the database is the product and SQL is the whole job. Choose EventBridge when AWS IAM and native cloud targets outweigh portability. Choose GitHub Actions for occasional repository maintenance. For a typical application cleanup endpoint, hosted app cron plus a queue worker is the least complex recovery story, provided you can expose and protect a public HTTPS endpoint.
Further reading
- RFC 2104, HMAC authentication: https://www.rfc-editor.org/rfc/rfc2104
- Google Cloud Pub/Sub overview (delivery and subscription concepts): https://cloud.google.com/pubsub/docs/overview
- pg_cron project documentation: https://github.com/citusdata/pg_cron
- AWS EventBridge Scheduler documentation: https://docs.aws.amazon.com/scheduler/latest/UserGuide/what-is-scheduler.html
- GitHub Actions scheduled workflows: https://docs.github.com/en/actions/using-workflows/events-that-trigger-workflows#schedule
