Below you will find pages that utilize the taxonomy term “Data Pipelines”
A Tiny ETL Binary Competes With curl, jq and SQLite in a Cron Job, So Build It That Small
The real competitor is a shell script in a crontab, and it usually looks like this:
curl -s "https://api.example.com/v1/launches?limit=100" \
| jq -r '.data[] | [.id, .name, .net] | @csv' \
| sqlite3 -csv launches.db ".import /dev/stdin launches"
It works on the day you write it. Then the API answers 429 and curl pipes an error page into jq. Or the API has a second page, and the script never asks for it. When the job dies halfway, the rerun inserts the same rows again or trips over the primary key, depending on how the table was made. A field that starts arriving as a string goes unnoticed until a chart looks wrong. Retries, backoff, pagination, incremental state, idempotent writes and schema drift: that’s the list, and shell scripts get every item on it wrong in predictable ways. (To be fair, curl --retry covers the first two.)
Make for APIs: Rerun Only the Steps Downstream of an Endpoint That Changed
A nightly job pulls a users endpoint and an orders endpoint, normalizes both, joins them on customer ID and renders a report. On most nights the upstream data hasn’t changed since the last run. The job doesn’t know that, so it downloads and parses everything, reruns every transform and writes the same report again. If the API bills per call, or one step takes twenty minutes, you pay for the same answer every night.