Insight · CMDB

Why your CMDB stops being trusted

Accuracy decays for four specific reasons, and a cleanup project addresses none of them.

Editor's note for Suresh: this draft is written from general ServiceNow practice, not from your specific engagements. Add your own examples where they fit — a real anecdote from HPE or T-Mobile will do more than anything here. Delete this box before publishing.

Almost every mature ServiceNow instance reaches a point where people stop believing the CMDB. Nobody announces it. It shows up as a change record whose impact analysis everyone ignores, an outage where the affected-services list is wrong, and a spreadsheet on someone's desktop that is quietly more accurate than the platform.

The usual response is a cleanup project. Discovery gets re-run, duplicates get merged, someone spends a quarter on it, and accuracy improves. Eighteen months later the instance is back where it started. That happens because a cleanup treats decay as an event, and decay is not an event. It is the normal behaviour of a data set that has more writers than owners.

Why accuracy decays

Four mechanisms account for most of it, and they compound.

The class model is deeper than anyone maintains

The CMDB CI Class Model is genuinely comprehensive, which is a virtue when you need it and a liability when you don't. Teams frequently extend it early, in the belief that more classes mean more fidelity. Each class then needs a population source, a reconciliation rule, and someone who notices when it stops being filled. Classes without those become half-populated tables that make every query on them untrustworthy.

Discovery coverage gaps go unmeasured

Discovery finds what it can reach. Network segments without credentials, cloud accounts nobody registered, appliances that refuse to be probed, and anything behind a firewall rule written three years ago are all invisible. The problem is not the gap; it is that the gap is not visible as a gap. An empty result and an unreachable target look identical downstream.

Multiple sources write the same attribute

Discovery says one thing, the asset system says another, someone's import set says a third. Identification and Reconciliation rules exist precisely for this, but they only work if somebody has decided which source is authoritative for which attribute. When that decision has not been made explicitly, it gets made implicitly by whichever job ran last.

Nothing is ever retired

Decommissioning is a process nobody owns. Servers get switched off, and their CIs stay Operational for years. Because retirement is invisible, the CMDB grows monotonically, and every count derived from it drifts further from reality.

The fix is not a bigger cleanup

A cleanup with no change to the four mechanisms above buys you accuracy that decays at exactly the rate it did before. The intervention has to change the mechanism.

Start from a decision, not from the data

The question that matters is not "is our CMDB accurate?" It is "which decision do we want to make from this data, and is it accurate enough to make it?" Change impact analysis, outage triage, patch scoping, licence position and vulnerability prioritisation are all different decisions, and they need different data to different tolerances.

Pick one. Make the data that decision depends on genuinely correct. Prove it works. That is a scoped, finishable piece of work with an observable outcome, which is more than most CMDB programmes can say.

Reduce the class model to what you populate

For each class in use, ask: what populates it, how often, and who notices when that stops? Classes that cannot answer all three should be retired or consolidated. A shallower model that is fully populated is worth more than a deep one that is patchy, because a user can trust it without knowing which parts to distrust.

Name an authoritative source per attribute

Not per CI — per attribute. Discovery is usually authoritative for technical attributes it observes directly. The asset or procurement system is usually authoritative for ownership, cost centre and contract. HR is usually authoritative for people. Write it down, configure reconciliation to match, and stop letting the last job to run decide.

Make coverage visible

Track what Discovery should reach against what it does. Publish the gap. An unmonitored subnet that everyone knows about is a manageable risk; the same subnet unknown is a wrong answer waiting to be delivered during an incident.

Build retirement into the process that causes it

Decommissioning has to be triggered by the workflow that actually decommissions things — the change, the cloud teardown, the asset disposal — not by a periodic audit. Anything audit-driven will lag by the audit interval, and audit intervals expand under pressure.

Governance that survives a reorganisation

Most CMDB governance fails the same way: a working group is formed, meets for a while, produces a standards document, and dissolves. The document survives. Nobody reads it.

What lasts is narrower and duller. A named owner per class, not a committee. A small set of health metrics — completeness, correctness, compliance to the model — reported to that owner on a schedule they cannot ignore. A rule that new integrations declare which attributes they write before they are approved. None of this is exciting, and all of it works better than a charter.

How to tell whether it is working

Accuracy metrics are easy to game and hard to argue with. The more honest test is behavioural: do people use the CMDB when it costs them something to be wrong?

If change managers rely on the impact analysis rather than asking around, if incident responders trust the affected-services list during a live outage, and if nobody maintains a private spreadsheet, the data is good enough. If any of those are false, the number on the dashboard is not measuring what matters.

Where we usually start

On an instance where confidence has already been lost, we tend to begin with a scoped assessment of one decision path — most often change impact, because it is high-frequency and the failure is visible. That produces a written finding on what is wrong, which of the four mechanisms is causing it, and what a proportionate fix looks like. It is deliberately small. Rebuilding trust in a data set is not a project you can announce; it is one you have to demonstrate.

Written by Suresh Ganapathy, founder of Tessivant.

Talk to us about your instance →