The score should reason like an analyst
How we wanted threat scoring to work: weigh the sources, admit the uncertainty, and get better with experience — not tally votes or guess.
A threat score is a judgment compressed into a number. The useful question was never whether to have one — it was whether the number reasons the way a careful analyst would. That was the goal we set from the start: not a verdict machine, but a judgment you could follow.
Most scores are made one of two ways, and we didn't want either. The first is a tally: count how many engines flag the indicator and turn it into a percentage. It looks objective, but many of those engines echo the same upstream source — so a dozen "hits" can be one opinion repeated twelve times. The second is a model that emits a number with no reasoning you can retrace. One over-counts. The other can't explain itself. Neither is how a good analyst thinks.
So we started from how a good analyst actually reasons, and asked the score to do the same. An analyst doesn't tally — they ask who is saying it, and how much that source has earned trust. A curated feed naming an active command-and-control server carries more weight than a single scanner's hunch. Two independent sources that agree count for more than one loud voice repeating itself. The score weighs its sources; it doesn't count votes.
It also had to know the difference between good news and no news. "No one has reported this address" is not the same as "this address is clean" — the first is silence, the second is a finding. A careful analyst would never read an empty file as exoneration, and neither should the score. Absence of evidence is not evidence of safety.
We wanted it to be honest about how sure it is. Severity and confidence are two different questions — how bad, and how certain — and collapsing them into one number hides the thing that matters most. A strong lead with thin corroboration is exactly that: worth attention, not yet proven. The score should say "leaning bad, not certain" out loud, instead of faking a precision it doesn't have.
And when the sources do line up — when everything with an opinion points the same way — that agreement is itself a signal. A careful analyst leans in when the room is unanimous. The score leans in too, a little, for the same reason.
The last goal is the one that makes it more than a fixed formula: it should get better the way an analyst does — by finding out what actually happened. Indicators reveal their true nature over time; the calls we got right and the ones we missed eventually become known. A score worth trusting checks its old verdicts against that reality and sharpens. It isn't frozen expert intuition. It learns.
We chose its failure mode deliberately. Faced with doubt, it would rather under-flag than cry wolf, because a flood of false alarms is how teams learn to ignore a tool. But a quiet miss is the more dangerous mistake, so we hunt those relentlessly — the low score that should have been high is the thing we most want to catch.
None of this is a magic number. The goal was a score you can reason with: one that weighs its sources by credibility, admits what it doesn't know, shows how it got there, and improves with experience. A score that thinks like your best analyst — and, unlike a black box, thinks out loud.
- It weighs sources by how much they've earned trust — it doesn't count votes that echo each other.
- Absence of reports isn't a clean bill of health, and the score says how certain it is, not just how severe.
- It checks its past calls against what indicators truly became, and gets sharper over time.