One-Line Definition
A risk score is a numerical value—typically between 0 and 100, or 0 and 1,000 depending on the provider—that a payment risk engine assigns to each transaction to express how likely that transaction is to be fraudulent, so the system can decide whether to approve it, send it to manual review, or block it outright.
Real-Life Analogy
Think of a risk score the way you think of a credit score, but compressed into milliseconds and applied to a single purchase instead of a person's lifetime of borrowing. A credit score of 780 tells a lender "this person is very likely to repay." A risk score of 12 tells a payment gateway "this transaction looks almost certainly legitimate—let it through." A score of 87 tells the same gateway "something is off here; a human should look before we ship the goods."
The difference is speed and granularity. A credit score updates monthly and describes a person. A risk score is recalculated for every single transaction and describes one moment: this card, this device, this IP address, this amount, this merchant, right now. A shopper who scores 8 on Monday morning can score 74 on Monday night if they suddenly switch to a new device, ship to a different country, and order three times their usual basket size.
Core Formula
There is no single universal formula, but almost every production risk engine works like a weighted sum of signals, then maps that sum through a calibration curve into a final score:
RawScore = Σ (weight_i × signal_i) + bias RiskScore = Calibration(RawScore) → scaled to 0–100 (or 0–1000)
In practice, the "signals" are hundreds of features: card BIN country, IP geolocation, device fingerprint, email age, velocity (how many orders in the last hour), historical chargeback rate for that BIN, shipping–billing distance, and so on. The weights come from either logistic regression, gradient-boosted trees, or a neural network trained on labeled fraud outcomes.
A simplified illustrative version:
| Signal | Value | Weight | Contribution |
|---|---|---|---|
| IP country ≠ card country | 1 | 18 | 18 |
| New device fingerprint | 1 | 12 | 12 |
| Order value 4× customer average | 1 | 9 | 9 |
| Email age > 2 years | 0 | −7 | 0 |
| Velocity: 5 orders / 60 min | 1 | 15 | 15 |
| **Raw total** | **54** |
That raw total is then calibrated against the merchant's own fraud base rate. A merchant with a 0.3% fraud rate will treat a raw score of 54 very differently from a merchant with a 3% fraud rate, because the same signals imply different absolute probabilities.
Comparison with Related Terms
| Term | What it measures | Output | Typical use |
|---|---|---|---|
| **Risk score** | Probability a single transaction is fraudulent | 0–100 or 0–1000 | Approve / review / decline decision |
| **Fraud probability** | Calibrated likelihood of fraud | 0.00–1.00 | Model output before scaling |
| **Credit score** | Likelihood a person repays debt | 300–850 | Lending, credit limits |
| **Chargeback rate** | Share of past transactions disputed | Percentage | Merchant account health |
| **Velocity check** | Count of events in a time window | Integer | Rule trigger, not a score |
| **3DS risk-based authentication** | Whether to challenge with OTP | Pass / challenge | Step-up authentication |
The key distinction: a risk score is transaction-level and decision-oriented. A credit score is person-level and access-oriented. A chargeback rate is retrospective, while a risk score is predictive at the moment of authorization.
Use Cases
1. Card-not-present authorization. A cross-border merchant selling to the US from a Singapore entity sees an order for $1,200 from a first-time buyer using a Brazilian card and a German IP. The engine returns 82. The merchant routes it to manual review, requests a selfie with ID, and confirms the order 20 minutes later—recovering revenue that a blanket decline would have lost.
2. Marketplace seller screening. A platform scores every new seller onboarding. A seller with a 3-day-old account, a mismatched bank name, and a product catalog copied from an established brand scores 91 and is held for KYC.
3. Subscription renewal protection. A SaaS company scores each recurring charge. A renewal that jumps from $29 to $299, originates from a new IP, and follows a failed payment attempt scores 67, triggering a 3DS challenge instead of an automatic charge.
4. Payout and withdrawal risk. A fintech scores outbound transfers. A user who deposits $500, wins $5,000 in a week, and immediately requests a withdrawal to a brand-new bank account scores 78—high enough to hold for 48 hours and verify identity.
5. Promo abuse and refund fraud. An e-commerce brand scores not just payments but refund requests. A customer with 11 orders and 4 refunds in 90 days, all under $50, scores 71 on the refund model and gets flagged for policy review.
Misconceptions
"A high score always means fraud." No. A score of 85 means "this pattern resembles past fraud." It could be a legitimate VIP customer traveling abroad. That is exactly why thresholds have three zones—approve, review, decline—not two.
"The score is the probability." Only if the model is well calibrated. Many raw model outputs are not probabilities. A score of 70 does not mean 70% chance of fraud. It means "70 on this provider's scale," which maps to a probability only after calibration against your own data.
"One score fits all merchants." False. A digital-goods merchant with instant delivery faces different fraud economics than a furniture store shipping in 10 days. The same score can justify different actions. A 60 might be auto-approve for one and auto-decline for another.
"Higher thresholds are always safer." Raising your decline threshold from 70 to 85 cuts fraud but also cuts good orders. If your average order value is $150 and your margin is 30%, declining 100 good orders to stop 5 fraudulent ones costs you $4,500 in margin to save $750 in fraud. The math rarely favors aggressive thresholds.
"Risk scores are static." They decay. A model trained on 2022 fraud patterns will miss 2024 attack vectors. Production systems retrain weekly or monthly, and scores drift even when the model does not change, because the fraud landscape moves.
Related Terms
- Fraud probability – the calibrated 0–1 output before scaling to a score
- Decision threshold – the cutoff values that route a score to approve, review, or decline
- Velocity rules – count-based triggers often blended into the score
- Device fingerprint – a high-weight signal in most risk models
- 3D Secure (3DS) – step-up authentication often triggered by mid-range scores
- Chargeback – the outcome that labels historical data for model training
- KYC / AML screening – identity and sanctions checks that run alongside risk scoring
- Manual review queue – where mid-range scores land for human judgment
- False positive rate – the share of legitimate orders wrongly flagged, the main cost of aggressive scoring