FinOpsAugust 29, 20266 min
New post

The cost nobody sees: users, non-production environments, and the contract gap

Your total cloud bill doesn't tell you whether your architecture scales well — cost per user does. A real story about defining a non-production vs. production benchmark, closing the contract gap around QA SLAs, and what to do when a client saturates your test environments without bad intent.

Once a cost owner had been named on the team, the question stopped being how much we were paying for the cloud. It became something more uncomfortable: how much does it actually cost to operate this. And that question branches out fast. How much does it cost per user? Per subscription, per tenant? How much do our non-production environments cost us? The easy answer is always in the billing report: it costs what the report says. But that answer doesn't help you design anything — it justifies an already-made expense, it doesn't help you decide how to grow.

Why does cost per user matter more than total cost?

What does help is looking at cost per user, not total cost. A total cost that rises along with the number of users says nothing about the health of your architecture. What does say something is whether cost per user drops as you grow — and that only happens when you design something that shares the most expensive resource, typically compute, across several clients or tenants, instead of multiplying dedicated infrastructure for each new one that joins.

What proportion should a non-production environment cost relative to production?

Then came an even more specific question: what proportion should a non-production environment cost relative to production? No university anywhere says, with authority, that QA or Staging shouldn't represent more than 10% of production's cost. It's an industry heuristic, not a commandment. But it forces you to ask the right question when the number spikes: why is an environment nobody bills for costing the same, or more, than the one actually serving the client?

You don't need 10% to be your exact number. You need to define what yours is, document it, and monitor it. Without that reference number, a non-production environment can grow silently until it becomes as expensive as the one that actually generates revenue, and no one will notice until the bill stops matching the business.

What happens if your contract says nothing about the SLA of non-production environments?

There arose the sharpest question of all: does the contract with the client explicitly say that non-production environments are not bound to the same SLAs as production? Because if it does, you have an argument. If it doesn't, the client can — legitimately, from their reading of the contract — expect QA to behave like production. The absence of that clause doesn't protect your infrastructure. It only postpones the uncomfortable conversation.

And that conversation arrived with a client whose integration, by design, sent exactly the same volume of orders to QA that it sent to production. Not dozens, not hundreds: thousands of replicated orders falling on an environment that was never sized for that. Nobody did it with bad intent. That's simply how the integration was built, and no one on the other side asked whether it had a real cost on our side.

Non-production overload is almost never an attack. It's almost always an integration design nobody reviewed from a cost standpoint, and that's why it's harder to detect than malicious abuse: it doesn't trigger any security alert. Non-production monitoring can't limit itself to asking whether there's unusual traffic. It has to ask whether the per-tenant volume in QA is consistent with testing, or suspiciously resembles production.

How do you communicate this to a client without them feeling accused?

Identifying the client was the easy part. The hard part was everything that came after: deciding how to tell them without making them feel accused of something they didn't even know they were doing. The difference between "you're abusing the test environment" and "so that QA stays representative, we ask you to limit testing to one or two pilot offices, with a maximum of five or six active users" is, many times, the difference between keeping the business relationship or opening a conflict that escalates to levels nobody wanted to reach.

Is the conversation about your non-production environments still pending?

Does your contract with clients clearly define what's expected of non-production environments, or does that conversation remain pending until someone forces it? In the second part of this story, the other side of the mirror: what happened when it was production, not QA, that had to grow.

Continue to Part 2: the end-of-month Pantitlán station
#FinOps#CloudCostOptimization#SaaS#Fintech#ContractManagement#CloudGovernance#SLA#CloudGovernance#CloudScars
Share:LinkedIn
Quick answerDetail

Does my total cloud bill tell me whether my architecture scales well?

No. A total cost that rises along with the number of users says nothing about the health of the architecture. The real signal is cost per user: it only drops as you grow when you share the most expensive resource (typically compute) across several tenants. Just as important is defining and monitoring a cost benchmark for non-production environments (a common heuristic is that QA/Staging shouldn't exceed 10% of production), and protecting in the contract that those environments aren't bound to the same SLA as production — otherwise, an integration that replicates production volume in QA silently saturates your test environments.

Written and reviewed by Rogelio Barajas González — certified Lead Auditor ISO 27001:2022 and ISO 9001:2015, with direct experience in SOC 1 Type 2 and SOC 2 Type 2. Founder of Barajas Advisory.

Verify his credentials on LinkedIn:linkedin.com/in/rogelio-barajas-gonzalez

Last updated: August 2026

This is one of nine real cases

Cicatrices de Nube — do you want the rest of the stories?

All nine documented cases —FinOps, Release Management, Service Delivery, Compliance, and AI governance— with a self-assessment checklist per chapter and an overall scorecard.

Download the free playbook

Does this resonate?

If you lead operations, technology, or teams at a SaaS company and recognize these situations, let's talk. No strings attached.

Schedule your diagnosis
Usually available

I respond within 2 hours max
Monday to Friday