Three parallel lanes of small glowing nodes, shaped like a web page, a mobile screen, and a CI pipeline, each with a different distribution of check marks weighted toward its riskiest section Image credits: Google Gemini
Engineering and Development

You Don't Have a Testing Problem. You Have a Confidence-Allocation Problem.

Risk-Based Testing Across Web, Mobile, and Infra in Plain Language

In my React Native years, back in the bridge era, the JavaScript side of the app was the part I tested hardest and the part that almost never broke. Reducers, formatters, the small deterministic things you can assert in four lines. All green, more or less permanently.

The app broke somewhere else entirely. A native module would behave one way on iOS and a slightly different way on Android. A version bump would change bridge behavior in some quiet way that nothing written in JavaScript could have noticed. Over time, crash reporting and analytics ended up doing more protective work for us than the whole unit suite did.

Those tests were well written and well maintained. They were also aimed at the cheapest part of the app to aim at.

I’ve since watched the same mismatch show up on stacks that have nothing to do with phones. Coverage accumulates wherever it’s cheapest to add, which is almost never where being wrong would actually cost something. That gap has a name as far as I’m concerned, and the name is confidence allocation.

There’s a formal name for doing this on purpose, which I picked up embarrassingly late: risk-based testing. The International Software Testing Qualifications Board defines it(opens in new tab) in the dry, precise way standards bodies do.

quote

Risk-based testing: testing in which the management, selection, prioritization, and use of testing activities and resources are based on corresponding risk types and risk levels.

ISTQB Glossary

Strip the vocabulary off and you get something an insurance underwriter would recognize. You have a fixed budget of time, compute, and team patience. You price each thing according to what it costs when it goes wrong, then you spend accordingly.

Let’s discuss the two shapes

Any conversation about this eventually turns into a conversation about a diagram, so let’s get the two famous ones out of the way.

The pyramid is many fast low-level tests at the base, progressively fewer and slower ones above, with an end-to-end tip that’s supposed to stay thin. It’s also where credit goes sideways. People attribute the pyramid to Martin Fowler constantly. It’s Ham Vocke’s model(opens in new tab); Fowler’s site is simply where it got published.

The trophy puts the weight in the middle instead: static analysis on the floor, unit tests above that, a thick integration layer, a thin end-to-end cap. That’s Kent C. Dodds’ model(opens in new tab), drawn with JavaScript-heavy frontend work in mind. And the slogan everybody quotes at it, “write tests, not too many, mostly integration,” didn’t originate with Dodds. He points at Guillermo Rauch and a 2016 tweet(opens in new tab) himself, in writing, which is better attribution hygiene than most of our industry manages.

Clean, minimal line-art diagram, two shapes side by side. Left: a wide-based triangle narrowing steadily to a thin tip, labeled bottom-to-top unit / service / UI. Right: a trophy silhouette with a thin static-analysis base, a thin unit-test stem, a thick integration cup, and a thin end-to-end cap. Same line weight and palette on both so they read as a matched pair. No text other than the labels, no logos, no watermarks, flat modern technical-illustration style.

Image credits: Illustration generated using ChatGPT. Prompt is in the alt text.

So: two of the most-cited ideas in software testing, and the credit floats loose on both of them. 🙃

What matters more is that they’re honest diagrams drawn for codebases where risk concentrates in different places, which is exactly why they disagree. Anyone following either one carefully is already doing risk-based testing. They’ve just inherited someone else’s answer about where the risk was.

Why a Shape Is Easier Than a Decision

A shape survives a code review. That’s most of its appeal, if you ask me.

“We follow the pyramid” ends a thread. “This seam doesn’t need an integration test, the blast radius is tiny and the logic’s already covered three lines up” invites a reply, and then another one, and now you’re defending a judgment call at 5pm on a Friday.

Rules are faster than judgment, and everybody is tired. It’s a completely understandable trade to make, and it quietly substitutes a diagram for a strategy.

The question the diagram is standing in for is less comfortable, because it has no default answer: what does it cost if this specific thing is wrong?

A rounding bug in an internal report is an apology and a patch. The identical rounding bug in a billing path is a refund cycle and a dent in how much people trust the product. Same code, same test, wildly different price tag. You can run that comparison on anything you’re about to ship, and the answer doesn’t care whether the code runs in a browser, on a phone, or inside a build system.

Narrowing It Down on a Real Change

In practice this collapses into a short argument with yourself, and it starts with cost rather than with layers.

If the honest answer to “what happens if this breaks” is a support ticket, you’re allowed to do very little, and spending heavily there just burns budget you’ll want somewhere worse. If the honest answer involves a postmortem and a timeline nobody enjoyed writing, you don’t get that option.

Next, what kind of wrong is even plausible. Deterministic, self-contained logic wants a small check you can run a thousand times a day without thinking about it: a pricing rule, a parser, a resource’s conditional branch. If exercising that logic takes a page of setup, the boundary is usually the real problem, and a fancier mocking library will not fix a boundary.

Interaction is a different animal. Nearly every bug I’ve been historically properly embarrassed by lived at a meeting point. A form, its validation, and the request it fires. A screen and the native module underneath it. A resource and the platform API it wraps. Each half looked fine in isolation, and nobody checked the handshake.

Cheapest of all are the checks that run before anything expensive does. Linting, type-checking, schema validation, Rubocop on a cookbook. Unglamorous, constantly skipped, basically free.

Then two questions that remove work instead of adding it:

Is something closer to the code already proving this? Proving it again at a slower layer buys you a second place to triage the same failure a year from now.

And will anyone still believe this check in six months? Plenty of behavior can be verified at a slow, high layer and really shouldn’t be; it’ll flake on timing or network or shared state, and a check nobody trusts has already stopped working as a check. It’s just noise with a red X next to it.

Last one, and the easiest to skip: green doesn’t mean fine. A page can pass every assertion you own and still take long enough to load that people leave, which is roughly the whole subject of Core Web Vitals. The checks that catch what people actually run into tend to live outside the test runner. axe-core wired into CI catches what a keyboard user walks into, Lighthouse CI alongside real user monitoring catches the gap between your laptop and a mid-range phone on a bad connection, and crash reporting catches the ones that only exist at scale.

Off the Browser, the Risk Moves

Both of those diagrams were drawn with browser-delivered products in mind. Most of my work now isn’t one.

Chef takes up most of my week, and very little of what keeps me up at night is whether a given class has a spec file. It’s closer to: does this Ruby patch quietly change behavior that downstream cookbooks depend on? Does this change to the pipeline itself stop actually testing the thing it exists to gate, right before a package goes out to thousands of machines?

RSpec coverage on individual resources is real work and I’m glad it’s there. It is almost never where the expensive surprises come from.

The expensive ones live in the gap between “CI went green” and “this is safe to put on production infrastructure.” They live in the gate itself, which is the one piece of the system whose failure mode is silently permitting everything it was built to stop.

GitHub Actions can scope a workflow to run only when specific paths change(opens in new tab), which means a change in one corner of a monorepo doesn’t drag every unrelated suite along with it. GitHub doesn’t call that risk-based testing and I’m not going to pretend it does. It’s plumbing that happens to make the instinct cheap to act on, and cheap matters more than correct vocabulary.

note

Risk-based testing as a discipline is formally defined, and I’ve cited the definition. The claim that the same reasoning carries across React Native apps and DevSecOps tooling the way it does on browser products is just my own experience across a few stacks, nothing more.

What You’re Actually Buying

The number at the bottom of a coverage report answers a different question than the one you were asking.

What you want to know is whether you can ship this particular change without quietly relying on luck for the parts that would genuinely hurt. The diagrams, the layers, the arguing about ratios, all of it is scaffolding around that single question, and the scaffolding is worth exactly as much as the answer it helps you reach. Risk-based testing, with the vocabulary stripped off, is just taking that question seriously on whatever stack you happen to work on.

A small suite that answers it is doing its job, even if the team next door has ten times the tests. A large one that doesn’t is also telling you something, and piling on more coverage without moving where it points mostly makes it more expensive to maintain the same blind spot.

So go find out where your own risk actually sits..

Cheers!