Cloud spend can accumulate through small configuration choices: a Kubernetes pod without resource requests, a CloudWatch log group retaining data longer than required, or an oversized CI runner. Review usage alongside reliability and performance so teams can identify costs that do not serve the workload.
This checklist groups AWS cost review tasks across compute, storage, data transfer, observability, databases, tagging, and CI/CD. The effect of each change depends on workload, region, configuration, and current pricing.
FinOps Principle: Cost optimization is not a one-time project โ it's an engineering discipline. The teams that win are the ones that make cost visibility a first-class requirement alongside reliability and performance.
๐ฅ๏ธ 1. Compute & Kubernetes
Compute is a major cost category in many environments. Review utilization and request/limit settings before changing capacity.
- โ Set resource requests AND limits on every Kubernetes container. Without requests, the scheduler cannot bin-pack nodes efficiently. Without limits, a noisy neighbour OOMs other pods. Both lead to over-provisioning.
-
โ
Enable the Kubernetes Vertical Pod Autoscaler (VPA) in recommendation mode.
Run it for 7 days and collect its
lowerBound/upperBoundoutput before updating manifests. - โ Migrate stateless workloads to Spot / Preemptible instances. Compare current Spot and On-Demand pricing for the region and workload. Use a mixed instance policy in your ASG or Karpenter node pool where interruption handling is appropriate.
- โ Purchase Savings Plans or Reserved Instances for baseline compute. Commit only to your steady-state floor (p10 of your hourly usage). Use Spot above that.
- โ Enable Cluster Autoscaler or Karpenter. Scale node groups down to zero during off-peak. A dev cluster running overnight with zero pods still costs money.
- โ Schedule non-production environments off outside business hours. Implement a Lambda + EventBridge rule to stop EC2 and scale ECS/EKS node groups to zero at 8 PM and restart at 8 AM.
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: general
spec:
template:
spec:
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"] # prefer Spot, fallback to OD
- key: node.kubernetes.io/instance-type
operator: In
values:
- m5.large
- m5.xlarge
- m6i.large
- m6i.xlarge
- c5.large
- c6i.large
disruption:
consolidationPolicy: WhenUnderutilized
consolidateAfter: 30s # remove underutilized nodes quickly
limits:
cpu: 1000
memory: 2000Gi
๐พ 2. Storage
-
โ
Audit and delete orphaned EBS volumes.
Volumes in
availablestate are detached and still billed. Find them with a single AWS CLI command. - โ Set S3 lifecycle rules on every bucket. Move objects to Intelligent-Tiering or Glacier after 30โ90 days. Delete incomplete multipart uploads (a silent cost many teams miss).
- โ Migrate frequently-accessed snapshots to gp3; delete the rest. EBS snapshots older than 90 days are rarely needed. Build a Lambda that enforces a retention window.
- โ Use gp3 over gp2 for EBS volumes. Compare current gp2 and gp3 pricing and provisioned performance for the workload before migrating.
- โ Compress and deduplicate CloudWatch Logs before shipping to S3. Use subscription filters to stream logs to Kinesis Firehose โ S3 + Athena for long-term querying at a fraction of CloudWatch storage cost.
#!/usr/bin/env bash
# List all detached EBS volumes and their monthly cost estimate
aws ec2 describe-volumes \
--filters Name=status,Values=available \
--query 'Volumes[*].{ID:VolumeId, Size:Size, Type:VolumeType, AZ:AvailabilityZone}' \
--output table
# Check current regional EBS pricing before estimating monthly cost.
# Include volume size, provisioned performance, and applicable charges.
๐ก 3. Data Transfer & Networking
Data transfer is easy to overlook. Review current regional rates and account billing for internet egress, cross-AZ traffic, and service-to-service transfers.
-
โ
Deploy services in the same AZ when they talk to each other heavily.
Use topology-aware routing in Kubernetes (
topologySpreadConstraints) and settrafficDistribution: PreferCloseon Services (K8s 1.31+) to keep traffic within-AZ. - โ Use VPC Endpoints for AWS service traffic. Compare endpoint costs and routing requirements with the current data-transfer charges for the services you use.
- โ Audit your NAT Gateway cost. Check current NAT Gateway processing and hourly charges. If workloads download large artifacts through NAT, consider an in-VPC artifact cache where it fits your operational needs.
- โ Use CloudFront for egress-heavy workloads. CloudFront-to-S3 origin transfer is free. CloudFront egress pricing tiers down significantly at scale vs direct S3 egress.
- โ Release unused Elastic IP addresses. Check current Elastic IP charges and release addresses only after confirming they are not in use.
๐ 4. Observability & Logging Cost Review
Observability infrastructure is one of the fastest-growing cost centres as teams scale. CloudWatch, Datadog, Grafana Cloud, and similar tools bill by volume โ logs, metrics, and spans.
- โ Set CloudWatch Log Group retention policies. Default retention is Never Expire. Set a maximum of 90 days for application logs; 1 year for audit/security logs. Archive to S3 for longer retention.
-
โ
Drop noisy, low-value logs at the source.
Use a FluentBit
greporrewrite_tagfilter to drop repetitive health-check logs (GET /healthz 200) before they reach CloudWatch or Loki, while preserving logs needed for diagnosis. - โ Downsample high-cardinality metrics. Recording 1-second Prometheus scrape intervals for 500 services is expensive. Evaluate which metrics genuinely need sub-minute granularity and scrape the rest at 60s.
- โ Use head-based or tail-based sampling for distributed traces. Choose trace sampling based on service volume, diagnostic needs, and retention costs. Preserve error traces where the tracing system supports it.
- โ Audit custom metric dimensions. In CloudWatch, each unique metric dimension combination is a separate billable metric. A single metric with 5 high-cardinality dimensions can produce millions of billable metric streams.
[FILTER]
Name grep
Match app.*
Exclude log GET /healthz
Exclude log GET /readyz
Exclude log GET /metrics
๐๏ธ 5. Database & Cache
- โ Rightsize RDS and Aurora instances. Use AWS Compute Optimizer or RDS Performance Insights to find instances running at under 10% average CPU. Downgrade instance class or switch to Serverless v2 for variable workloads.
- โ Switch RDS Multi-AZ to a read replica architecture for read-heavy workloads. Multi-AZ doubles your instance cost for HA. For read-heavy applications, a single primary + read replica can be cheaper and provide better read scalability.
- โ Evaluate Aurora Serverless v2 for dev/staging databases. Aurora Serverless v2 scales to 0 ACUs when idle and bills per second. Perfect for environments that aren't used 24/7.
- โ Reduce ElastiCache node sizes and enable data tiering. Compare data-tiering options against the workload's access pattern and current node pricing.
- โ Set DynamoDB table billing mode to PAY_PER_REQUEST for low-traffic tables. Provisioned capacity with auto-scaling often over-provisions during quiet periods. On-demand billing is simpler and cheaper for tables with spiky or low traffic.
- โ Audit RDS automated backup retention windows. A 35-day backup window on a 500 GB database creates significant snapshot storage cost. Reduce to 7โ14 days and use manual snapshots for long-term compliance.
๐ท๏ธ 6. FinOps: Tagging & Visibility (Foundation for everything else)
You cannot optimize what you cannot attribute. Without a consistent tagging strategy, your cost explorer shows a wall of unallocated spend and no team can be held accountable for their resource usage.
-
โ
Enforce a mandatory tag policy via AWS Organizations SCP.
Require
Environment,Team,Service, andCostCentertags on all taggable resources. Reject resource creation that omits them. - โ Enable AWS Cost Allocation Tags and activate them in Cost Explorer. Tags only appear in Cost Explorer 24 hours after activation. Do this now.
-
โ
Set up Cost Anomaly Detection.
AWS Cost Anomaly Detection uses ML to flag unexpected spend spikes. Configure monitors per service and per linked account with an SNS alert to your
#finopsSlack channel. - โ Create per-team AWS Budgets with alert thresholds at 80% and 100%. Budget alerts via SNS โ Slack or email keep engineers aware of spend before it becomes a month-end surprise.
{
"tags": {
"Environment": {
"tag_key": { "@@assign": "Environment" },
"tag_value": {
"@@assign": ["dev", "staging", "production"]
},
"enforced_for": {
"@@assign": ["ec2:instance", "rds:db", "s3:bucket", "eks:cluster"]
}
},
"Team": {
"tag_key": { "@@assign": "Team" },
"enforced_for": {
"@@assign": ["ec2:instance", "rds:db", "lambda:function"]
}
},
"CostCenter": {
"tag_key": { "@@assign": "CostCenter" },
"enforced_for": {
"@@assign": ["ec2:instance", "rds:db"]
}
}
}
}
โ๏ธ 7. CI/CD Pipeline Costs
-
โ
Right-size CI runner instance types.
Most build jobs are I/O-bound, not CPU-bound. A
c5.large(2 vCPU, 4 GB) often runs builds as fast as ac5.4xlargefor typical application code. Benchmark your build times vs instance class. -
โ
Cache dependencies aggressively.
In GitHub Actions, use
actions/cachefornode_modules,.gradle, pip, Maven, and Docker layer caches. Uncached builds that re-download 2 GB of npm packages waste both time and runner minutes. -
โ
Cancel redundant workflow runs on push.
When a developer pushes 3 commits in quick succession, only the latest matters. Use
concurrencygroups in GitHub Actions to auto-cancel in-progress runs. - โ Run lint and unit tests in parallel, not sequentially. Splitting a sequential 20-minute pipeline into parallel jobs of 7 minutes cuts per-commit runner cost significantly and improves developer feedback time.
name: CI
on:
push:
branches: [main, 'feature/**']
# Cancel in-progress runs for the same branch
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Cache Node modules
uses: actions/cache@v4
with:
path: ~/.npm
key: ${{ runner.os }}-node-${{ hashFiles('**/package-lock.json') }}
restore-keys: |
${{ runner.os }}-node-
- run: npm ci
- run: npm test
Quick-Win Priority Matrix
Action โ Effort โ Cost review โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโโโโโโ Release unattached EIPs โ Low โ Confirm usage Delete orphaned EBS volumes โ Low โ Check backup needs Set CloudWatch log retention โ Low โ Compare retention Enable Spot for dev/staging EC2 โ Low โ Compare current rates Filter health-check logs โ Low โ Preserve useful signals gp2 โ gp3 EBS migration โ Med โ Compare volume needs Add S3 lifecycle rules โ Med โ Check retrieval costs VPC Endpoints for S3/DynamoDB โ Med โ Compare transfer costs Kubernetes VPA + Karpenter โ High โ Review utilization Savings Plans / Reserved Instances โ High โ Check steady demand
Conclusion
A structured FinOps practice starts with visibility and account-specific review. Check for unused resources, confirm retention requirements, and compare measured usage with current regional prices before making changes.
Then build the discipline: enforce tags, configure anomaly detection, and review Cost Explorer weekly as part of your team's sprint rituals. Cost optimization is an engineering problem, and like any engineering problem, it rewards systematic thinking over heroic one-time cleanups.