When Your Own Tools Tell You the Wrong Thing

In one fortnight, three of our own tools reported something false — a record that said work did not exist, a check that could not fail, and a component that had silently lost a capability. None of them errored.

Ganda Tech Services 7 min read
When Your Own Tools Tell You the Wrong Thing

Over one fortnight, three separate tools of ours reported something that was not true.

A tracking record insisted three finished files did not exist. A quality check approved everything it saw, having never been capable of refusing anything. A shared component reported a clean upgrade while silently losing three capabilities.

Not one of them threw an error. All three produced confident, well-formed, wrong output, and each was believed for weeks.

Three in a fortnight is not three bugs. It is a pattern, and the pattern is worth understanding because every business now runs on tools that report.

What the three had in common

Each was answering a narrower question than the one being asked of it.

The tracking record answered “what does this file say?” — which is not “what has been produced”. Correct answer, wrong question.

The quality check answered “does a fingerprint match?” — which is not “is this a duplicate”, because for one whole category of item the fingerprint never existed. Correct logic, absent input.

The upgraded component answered “does this compile and pass its tests?” — which is not “does it still do everything it did before”, because the tests only covered what the new version claimed.

In every case the tool was working. The gap was between the question it actually answers and the question everybody believed it answers, and nothing in the output distinguishes those.

Why nobody catches this

Three reinforcing reasons.

The output looks the same as a correct one. There is no formatting difference between a well-formed true result and a well-formed false one. Every heuristic people use to spot problems — errors, warnings, odd formatting, missing data — is absent.

Confidence accumulates with time. A tool that has reported the same reassuring thing for six weeks feels more trustworthy than one that reported it once. But if the tool is blind, six weeks of output is six weeks of the same absent information, and the growing confidence is entirely unearned.

The reassuring reading is the cheap one. Zero problems found means either nothing is wrong or the check is broken. Investigating costs time; accepting costs nothing today. Everybody accepts, including us.

★ Insight ───────────────────────────────────── The common structure is an assumption that lives outside the code. The record assumed a writer that never ran; the check assumed a file that is never written; the component assumed its tests covered the old behaviour. None of those assumptions is stated anywhere, which is why review does not find them — you cannot review a sentence nobody wrote. ─────────────────────────────────────────────────

The habit we adopted

One rule, and it is cheap: before trusting a tool’s reassurance, make it fail once.

Feed the duplicate check a duplicate. Delete a file and see whether the tracking record notices. Break a capability and confirm the tests go red.

Five minutes each, and it converts an untested assumption into evidence. A control that has never refused anything has produced no information, however long it has been running.

For a business not writing software, the same rule in ordinary terms:

  • A stock report that always balances — deliberately miscount one item and see whether the report shows it.
  • An approval workflow — submit something that should be rejected.
  • An alert — trigger the condition and check the message arrives in a channel somebody reads.
  • A backup — restore it, and time how long that took.
  • An insurance policy — ask what evidence a claim requires, before you need to produce it.

The awkward part

Two of the measurements we published in this period were themselves wrong before they were right.

One audit classified social redirect pages as commercial pages, filling a “pages with no inbound links” list with things that were never meant to have any. Another resolved keyword ownership by whichever site appeared first in a config file, and reported 183 violations against a site for using its own keyword.

Both were caught the same way: the result was surprising, so it was checked before it was reported.

That is the only defence that works consistently. A number that surprises you is either a finding or a bug, and the cost of telling them apart is always lower than the cost of publishing the wrong one — or of reorganising a quarter’s work around it.

What we changed

Three changes came out of the fortnight, and they are deliberately small because large process responses to this kind of finding do not survive contact with a busy quarter.

Every new check ships with one recorded negative case. Not a unit test — an actual run, on the actual path, refusing an actual bad input, with the output kept. It is the proof the control exists.

Any surprising number gets checked before it gets reported. We had two of our own measurements come out wrong this fortnight and both were caught this way. The rule costs an hour occasionally and it has already paid for itself twice.

Anything that holds the same fact in two places gets a reconciliation. Not a dashboard — a comparison that runs on a schedule and says whether the two sides agree. If a reconciliation is impractical, the fallback is that one of the two copies is declared non-authoritative in writing, so at least nobody is choosing between them by accident.

None of that is sophisticated. All of it addresses the same underlying property: a system that cannot disagree with itself cannot warn you.

The question worth asking this month

Pick the report your business trusts most. The one nobody questions, that has been right for years.

Ask what would have to be true for it to be wrong, and whether anything would tell you.

If the honest answer is that nothing would, you have not found a problem — but you have found something that has never been tested, and on our recent record that is where they are.


Ganda Tech Services runs web, cloud, mobile and content operations for a group of Australian brands. Systems, reporting and verification work is handled through Cloud Geeks.

Tags

Digital SolutionsBusiness TechnologyVerification