The Only Signal Count That Matters in Fraud Prevention

The fraud prevention industry is trapped in a signal-count arms race.

One vendor says it has a thousand signals. Another says 10,000 data points. The implied message is simple: the biggest number must produce the best decision.

It does not.

A signal count tells you almost nothing until you know what the vendor is including in their count. There is no shared definition, independent benchmark or standard audit, which can make parsing intelligence signals difficult.

Three ways to count a signal

There are a few ways to count a signal, and they produce wildly different outcomes from the same underlying data.

  1. The first is anything usefully exposed somewhere in a product interface. 
  2. The second is anything actually available in a raw data response. 
  3. The third is any component input within a data science or machine learning model, which, in practice, means counting database-layer metadata: the length of a string, the time since a previous transaction, the time zone offset of a timestamp. 

To put this into a different perspective, imagine a photo library counted not by how many pictures you have, but by how many pieces of metadata sit behind each one: file size, GPS coordinates, camera model, color profile, a dozen EXIF fields nobody looks at. You’d get a much bigger number. You’d also be measuring something almost nobody cares about, dressed up to look like the thing they actually do care about.

The question that actually matters

All of these approaches can produce radically different totals from the same information. But they do not answer the main question a fraud leader needs answered: How much of this information can change a real-time risk decision?

That’s what should be the real standard. Whether a risk input can help a team approve, challenge, review or stop a transaction when it counts, not how many fields a vendor can store or how many ways a database can slice a phone number, timestamp or device identifier.

Here’s why this number matters

Interface counts can inflate easily. Add a field to a response payload, whether or not anything downstream ever reads it, and the number goes up. Database-layer counts inflate even faster, because a single meaningful signal, such as a phone carrier or a device fingerprint, decomposes into a dozen metadata fields the moment it lands in a table. Counting the fields separately makes a thin signal set look deep. It also means two vendors with identical underlying data can report vastly different totals, and both are technically telling a version of the truth.

But more data does not automatically mean better signals. A vendor can report an impressive number while counting information that never affects a score, rule, investigation or action. They may describe the size of its data layer, but it does not describe the ability to make smarter risk decisions.

What depth actually looks like

Take an email address. Deliverability answers one question: Can this mailbox receive mail? That same question, however, is worthless against a fabricated identity, because a fabricated identity’s mailbox can work just fine. Add provenance and you can ask how long the address has been observable and whether its construction looks human or generated. Add breach exposure and you can ask whether an address with no visible history is genuinely new or merely unexposed. Add the association, and you can ask which other accounts, domains and prior risk events are connected to it. That’s multiple decisions from one dimension. 

Three questions to ask before you trust a signal count 

Fraud leaders should ask three questions whenever a vendor leads with a signal count:

  1. What exactly are you counting?
  2. Which of those inputs can influence real-time decisioning?
  3. When that context changes, what action can the platform take to help us?

These answers will reveal whether the number represents usable risk intelligence or just an exhaustive set of database metadata.

What we count

SEON tracks 1,100+ fraud signals available for real-time decisioning. It’s deliberately the most conservative number of the three we could have used, because it’s the one tied to actual decisions rather than to how much data sits in a table.

We count the individual fraud inputs our decisioning engine can use because when a fraud ring slips past a set of one-off checks, the question that matters is not how much data sits in a vendor’s database but rather if it’s the right combination of signals across different dimensions that can surface risk before loss occurs.

Take the First Step Toward Transformative Fraud Prevention