Egress
The largest line in almost every streaming bill. Cache policy, origin placement and — where it is justified — moving delivery onto owned capacity entirely.
Availability gets designed, tested and owned. Cost usually gets discovered. Treating spend the way you treat latency — measured, budgeted, reviewed and attributable to a decision in version control — is most of the work.
In most organisations the cloud bill arrives at finance, who cannot evaluate it, and is forwarded to engineering, who did not budget it. It grows quietly because no single person is accountable for the number and because the individual decisions that produced it were each defensible in isolation.
Cost engineering is not a cost-cutting exercise. It is making spend a property that somebody owns, that is measured continuously, and that can be traced back to a specific decision in a specific commit. Once that is true, the reductions tend to be obvious.
The largest line in almost every streaming bill. Cache policy, origin placement and — where it is justified — moving delivery onto owned capacity entirely.
Instances sized against measured concurrency rather than against launch-day nerves, and scaled down for the hours nobody is watching.
A catalogue's access curve is steep. Storage class should follow it instead of leaving 2009 in the hot tier forever.
Reserved capacity and savings plans bought against a baseline that has actually been observed — never against a forecast, and never in a panic.
Tag and map every resource to a service and an owner. Anything unattributable is the first finding, every single time.
Match the bill against real traffic — concurrency, delivered gigabytes, encode minutes — so cost per viewer becomes visible.
Changes applied as Terraform through the normal pipeline, so every reduction is reviewable and reversible.
Budgets and anomaly alerts per service, reviewed on a cadence, so the next drift is caught in days rather than quarters.
Only for the tiers where the numbers say so, and only if they do. Plenty of workloads should stay exactly where they are, and the recommendation follows the measurement.
It depends entirely on what has been engineered already, which is why no percentage is quoted here. A bill review produces a ranked list with a number attached to each item, from your data.
The review is one-off. Keeping the number where the review put it is ongoing, and that is what budgets, anomaly alerts and a regular review are for.
With your concurrency curve and bitrate ladder. Most of the answer is already in the bill, unread.