September 19, 2026

Is Spot actually lowering your AWS bill AT ALL?

Surprisingly, the answer could be No Just one of the 5 losses here can reduce savings by HALF! (Hint: its the instance types)

History

Spot used to be a great deal for ~15 years, is it still?

After bidding went away at Re:Invent 2017, you can’t pay more than On-Demand for Spot anymore, and prices also only move up (or down) gradually. Capacity was plentiful for many common instance types, and discounts over On-Demand were high. The 2020’s started with price savings lowering, and less available capacity, leading to increased interruption rates and increased headaches. Savings are much improved as of Q3 2026, but you can’t assume it will stay that way in a cost model.

Summary

So where are we now in terms of pros and cons?

Savings

  • Cost is always less than On Demand, often over 50% off
  • Free if terminated by EC2 within first hour from launch

Losses

  • Cost of churn: Bootstrapping time for interruption replacements is billed, as well as time for graceful shutdowns/draining
  • You’re not using the ideal instance type: Different family, generation, adding an unnecessary suffix, too large/small. This is the biggest hidden cause of losses
  • Could you be using a Savings Plan instead for some of the work?
  • Needing more On-Demand than expected to hit availability goals
  • Increase in related resource usage
  • All of this is more complicated to setup, validate, and maintain; and your time isn’t free

There are many ways to mitigate each of these, but for now let's examine what goes into those savings and losses more carefully.

Spot Savings

This is the more straightforward part. AWS has spare capacity, and sells it at a discount. That discount changes periodically throughout the day and year based on long term availability trends. Let’s assume an average savings of 60% over On Demand as a baseline.

The other small addition to savings often forgotten is that any instance Spot interrupts and terminates in the first hour is free. This is a holdover from when EC2 had hourly billing. The Caveat is that anything gracefully terminated by ASG/K8s/etc from the 2 minute warning coming in does NOT count here. Spot has to terminate it. So this isn’t a huge boost, but helps offset the losses below during more volatile capacity times.

Spot Losses

We’re going to make some minorly pessimistic assumptions here, to get the point across. But everything will be kept within the realm of a realistic workload.

Cost of Churn

Section total losses: 4% of the 60% savings gone: 56% remaining

This is predominantly an issue for applications with long startup or draining time, but will affect every application to some degree. Remember that instances are only billed in the RUNNING state (except hibernating instances in STOPPING).

Startup:

Assume bootstrapping takes 10 minutes, and the instances on average run for 6 hours without Spot (1/36th time lost - 2.8%). If Spot instances live an average of 3 hours, you’ve now doubled that loss.

Draining:

If you’re using Spot, presumably your application doesn’t critically fail without 10 minutes of draining. But that doesn’t mean there’s no penalty. Maybe there’s a 15 minute checkpoint on batch work, partial productivity while draining sessions, etc. Assuming 5 minutes lost, that’s another 1.4% loss, using the Startup section run times.

Less Ideal Instance Types

Section total losses: 35% of the 60% savings gone: 21% remaining

The key to Spot success is Flexibility. Instance type, time, AZ, maybe even Region. The more flexible you can be, the better your availability will be. But that first one, instance type diversification, can really bite you in the budget.

These examples won’t always be what you get, sometimes you’ll get the ideal type. So the lost savings are pro-rated down to be fair.

Size - too small:

Remember the startup churn we talked about above? That applies here too. If you get 20 Large instances instead of 5 2xl instances, you now have 4x the bootstrapping losses.

2% of your savings gone

Additionally, if you have a lot of overhead (large application footprint, Windows, etc), then your overhead is a much larger percentage of your fleet on those smaller instances, and you might need more total instances to run the same work (ex: 22 Large vs 5 2xl’s).

3% of your savings gone

Size - too big:

Did you need 3 Large instances, but instead only XL were available? You now have 2 2XL instances running, 25% more compute than you actually needed.

5% of your savings gone

Wrong family:

Your application will have an ideal instance family. Looking at the core M/C/R families for example, you’ll either lean more CPU heavy, Memory heavy, or in the middle. But to diversify in Spot, people often add multiple families. 1 example:

You need 4 vCPUs and 16GB of RAM for optimal performance. On-Demand prices in us-east-2:

Instance Type vCPUs RAM $/hour % more $
m9g.xl 4 16GB $0.19568 0%
r9g.xl 4 32GB $0.25684 31%
c9g.2xl 8 16GB $0.34656 77%

Your spot metrics may show huge cost savings, but if you’re starting 77% higher, it would probably be cheaper to run the M instance On-Demand vs C as Spot. Being generous:

7% of your savings gone

Unnecessary suffixes:

Same as instance families, your spot % savings will look great, but the baseline number is way higher than your ideal instance type. I’ll even leave out graviton and Flex in this chart, which would have lowered the baseline cost over 10% more:

Instance Type vCPUs RAM $/hour % more $
r8i.xl 4 32GB $0.27784 0%
r8a.xl 4 32GB $0.31952 15%
r8id.xl 4 32GB $0.33264 20%
r8ib.xl 4 32GB $0.4184 51%
r8in.xl 4 32GB $0.4184 51%
r8idb.xl 4 32GB $0.46894 69%
r8idn.xl 4 32GB $0.46894 69%

10% of your savings gone

Older Generations:

This one's a bit harder to pin down, since it needs load testing to see what your actual cost/performance is on a given instance type. But going based on AWS’s numbers, it’s ~20% better per generation. And sometimes we’ll be going 3+ generations back with Spot

8% of your savings gone

Ignoring Savings Plans

Section total losses: 5% of the 60% savings gone: 16% remaining

You know the amount you save with a Savings Plan (SP), and that it won’t change unless the SP becomes heavily underutilized. Spot savings can change quite a bit over those 3 years. And worse, spot availability can change. If you have periods where you need to move large amounts of Spot to On-Demand, the savings math changes quickly.

I’d almost never suggest using Spot for any baseline usage (capacity used over 90% of the time), unless your workload is extremely time flexible (and maybe region flexible) to ensure you’ll almost always be able to run it on Spot. Even then, I’d still probably run some of it as a Savings Plan.

There are 2 losses here that overlap, since both only apply to your baseline load level, which we’ll assume is ~20% of total EC2 usage.

1) Having to run On-Demand (or very expensive non-ideal Spot instances) when you could have been running a SP. Assume 20% of your Spot workload has to be moved to On-Demand for 10% of the year, which could have been a CSP (Compute Savings Plan). A no-upfront 3 year CSP is ~50% cheaper than On-Demand:

2% of your savings gone

2) Since we’ve shown in the other sections the “real” Spot savings will likely be <50% vs On-Demand, it's cheaper to run a CSP for your baseline load the rest of the year too. Taking a conservative value of spot actually being ~35% cheaper than On-Demand, but understanding only maybe 20% of your workload fits in the ‘baseline load’ category, and ignoring the 10% of the year we accounted for right above this:

3% of your savings gone

Needing More On-Demand Instances

Section total losses: 5% of the 60% savings gone: 11% remaining

Are you actually putting enough On-Demand buffer in for your availability goals? As mentioned in the CSP section, are you accounting for times of year when there will be less spot availability and you may need to move some or all of the Spot capacity back to On-Demand? In my experience, people often mis-calculate this with over optimistic estimates from their tests.

5% of your savings gone

Section total losses: 2% of the 60% savings gone: 9% remaining

Are there:

  • large downloads during bootstrapping leading to extra network usage?
  • More public IPs being assigned to multiple small instances
  • What about extra EBS volumes (again, if there are multiple smaller instances vs a few large ones)

The little things can add up:

2% of your savings gone

What is your time worth?

Section total losses: 0% of the 60% savings gone: 9% remaining

Spot inherently will take more effort to get working than On-Demand (well, assuming you care at all about availability ;) ). Are there other things you could be doing to optimize your environment, add new features/revenue, or otherwise use your time better? Opportunity cost is a real thing.

But there’s a big Caveat:

Much of the work to make an application Spot Ready also makes you generally more resilient. It’s not just work for Spot, it's work to have better scalability, availability, and fault tolerance. That type of work often gets put on the back burner until something bad happens. Wrapping it in the sugar coating of Cost Savings With Spot might make it easier to get the boss to prioritize it.

This almost wasn’t in the losses section because it's also a benefit. But I flipped a coin, and here we are. Not adding to the losses since it's technically not a hard cost, but something to keep in mind nonetheless.

Conclusion

Article total losses: 51% of the 60% savings gone: 9% remaining

So, is Spot worth it? For many workloads, especially more stateless ones, or ones which are more flexible; yes, it's definitely still worth it.

But there are many workloads which might have little to no savings at all. The numbers used above are clearly fictitious examples, and you probably wouldn’t have all of the loss types at the same time. But they’re estimates based in reality with a goal to get you to think about what the full picture might look like for your workload. You can easily see how someone could end up losing money with something that looks like a good, diversified configuration on paper.

Action Items

  • Figure out which of the above sections is your biggest loss and focus on optimizing it
  • Reduce bootstrapping time
  • Load test on different instance types to see which are best
  • Test different hardware (ARM vs x86, generations, hyperthreading, etc) to see if some are better for you
  • Use Interruption Metrics in EC2 Capacity Manager to see if some instance types/AZs have way higher churn for you

If you want help architecting your workload to use Spot, or to optimize your existing Spot setup, A Cloud Above Us will be happy to be your guide.

Planning a cloud project? Let's talk it through.

Book a free 30-minute consultation. Tell us where you are headed - migration, cost, architecture or resilience - and we will come back with a straight answer.

Book a free consultation

Send us a message

Book a free consultation

30 minutes, no obligation. Pick a time that suits and we will send a calendar invite.

Loading available times…

Not loading? Open the booking page in a new tab.