Article
Three Layers, One Signature: The Science Behind Proving Free and Informed Consent at the Moment of Payment
By Peter Walker, Co-founder & CEO, RTScale
Every fraud team I talk to has the same blind spot, and it isn't a technology gap — it's a question their stack was never built to answer.
Device fingerprinting tells you whether it's the right device. Behavioral biometrics tell you whether it's the right hands on the keyboard. Transaction-graph models tell you whether the payment looks statistically normal. All useful. All necessary. And all of them wave through the transaction that matters most, because in an authorized push payment (APP) scam, an elder-exploitation case, or a coerced transfer, the legitimate account holder is the one pressing approve. The account is right. The device is right. The behavior is right. The person is being scammed or coerced anyway.
The unanswered question is the human one: is the person authorizing this payment in a state of mind consistent with free and informed consent?
That is the question RTScale's State of Mind Signature™ (SoM Sig™) was built to answer. And the reason we can answer it is that human beings, under duress, leak. Not in one channel — in three. This article walks through the science of each layer, and why composing them, rather than relying on any one, is what turns a soft behavioral signal into an audit-grade asset.
Why one channel is never enough
The stakes are not abstract. FinCEN's review of Bank Secrecy Act filings found financial institutions reported over $27 billion in suspicious activity tied to elder financial exploitation across a single year, with a median scam of $33,000. In the UK, the mandatory APP-scam reimbursement regime that took effect in October 2024 has already moved hundreds of millions of pounds of liability directly onto payment service providers. The window between intent and irrevocable settlement has collapsed onto instant rails, and the liability has moved with it. Fraud teams are now being asked to evidence something they have never been able to capture: the payer's state of mind at the instant of authorization.
No single affective signal is strong enough to carry that weight. The research here is unambiguous, and it points in one direction: fusion.
Layer one — the face
Facial affect is the most studied channel in emotion science, and the most immediately capturable. When a payer articulates a short scripted phrase during a Quick Scan — three to five seconds — we get a dense read on the facial action units that decades of work descending from the Facial Action Coding System have mapped to affective states. The face is fast, it is involuntary in part, and it is legible at consumer camera resolution.
But facial expression alone is ambiguous. A furrowed brow is concentration or distress; a still face is calm or suppression. This is precisely why the face is a starting layer, not a verdict — and why our capture is structured and ongoing rather than a single snapshot. The Quick Scan produces an initial signature fast enough to clear a clearly benign payment. When the signal shows variance beyond a calibrated threshold, the system escalates to a Deep Scan: a longer repeated-phrase protocol where cross-repetition consistency — a signal breaking down, an intensity escalating — becomes diagnostic in a way no single frame can be.
Layer two — the voice
Voice is where ambiguity starts to resolve, because prosody is harder to fake than a face and it carries a different cut of the emotional state. The acoustic-emotion literature is consistent: high-arousal states such as fear and anxiety show up in rising and more variable pitch (f0) contours and increased intensity, while vocal stability degrades under strain — measurable as elevated jitter and shimmer and a falling harmonics-to-noise ratio, the acoustic fingerprint of vocal tension. A voice under duress does things a voice at rest does not, and it does them below the speaker's conscious control.
The reason to add voice is not that it is better than the face. It is that it is different from the face, and the two disagree in informative ways. The multimodal emotion-recognition literature has quantified this repeatedly: fusing audio and visual channels outperforms either alone, precisely because each modality disambiguates the other where one is insufficient. On standard benchmarks, adding modalities produces measured accuracy gains over the best single channel — not a rounding error, but the difference between a signal you can act on and one you can't. When the face is ambiguous, the voice often is not, and vice versa.
Layer three — the microexpression
The third layer is the most powerful and the most misunderstood, so I want to be precise about what we claim and what we don't.
The concept traces to Ekman and Friesen's 1969 work on nonverbal leakage: the idea that when someone conceals an emotion, the suppressed feeling can escape as a brief, involuntary facial movement — a flash of the opposite of what is being said. It is a compelling idea, and it is also one the field has rightly pressured. In high-stakes deception studies, clean microexpression leakage appears in only a minority of lies; skilled concealment, masking, and neutral faces are common. Anyone selling microexpressions as a standalone lie detector is overselling the science, and fraud professionals are right to be skeptical.
So we don't use them that way. A microexpression, for us, is not a verdict — it is a trigger and a corroborator. When a keyword-bound micro-movement fires on the recipient's name, or the payer's involuntary signal contradicts their spoken confirmation, that is not a fraud determination. It is a reason to escalate from a three-second scan to a thirty-second one, and a channel-level contribution that either corroborates or fails to corroborate the other two layers. Rare leakage is a problem if it's your only signal. It's an asset when it's the third voice in a corroborating panel
Composition is the product
Here is the core idea, and it is why RTScale's SoM Sig™ is an engineering artifact rather than a psychology claim. No single layer is trusted on its own. The signature is a composite object: separate facial, vocal, and microexpression sub-vectors, a fused cross-modal vector, and the metadata needed to reproduce and explain it. Any variance in the final signature can be decomposed back to channel-level contributions — how much the face drove it, how much the voice, how much a keyword-bound micro-movement.
That composition is also what produces levels of assurance rather than a single pass/fail. More layers present, stronger identity binding, and a more rigorous capture protocol yield a higher-assurance signature. A lightweight presence check sits at one end; a maximum-rigor adjudication capture — all three channels, continuous identity binding, repeated-phrase protocol — sits at the other. A fraud team chooses the tier that matches the risk: routine real-time retail payments run at the workhorse authorization tier, while a first-time payee with urgency markers — the classic APP-scam profile — warrants the maximum-rigor capture. Assurance is dialed to risk, not applied uniformly.
Why this protects accounts — and why it's auditable
Two properties matter to a fraud-prevention team specifically.
First, active confirmation protects the account holder against coercion in a way passive monitoring cannot. Because the payer articulates a scripted phrase and the system reads the involuntary response in real time, a coerced or scammed authorization carries affective evidence of the duress at the moment it happens. You are no longer inferring consent from the fact that a button was pressed. You are capturing the state of mind behind the press.
Second, the output is a verifiable asset. The SoM Sig™ is captured on-device and cryptographically signed as a tamper-evident, non-repudiable pair of token and payload. It is interrogable after the fact: a dispute reviewer, a regulator, or an auditor can decompose exactly why a given authorization was cleared or escalated, down to the channel and the action unit. In a reimbursement regime where the institution must evidence its decision, a signature that can be independently audited is worth more than a score that cannot.
The IP, plainly stated
The State of Mind Signature™ and SoM Sig™ are trademarks of RTScale, built on a foundational patent for detecting deception in audio-video responses. The multi-layer composition, the two-stage capture protocol, the channel-level variance decomposition, and the tiered assurance model described here are RTScale intellectual property. This article describes the science and the architecture; it does not disclose the implementation.
See it on your own transactions
If your team is carrying APP-scam liability, elder-exploitation exposure, or coerced-transfer risk on instant rails, the fastest way to evaluate this is to see a SoM Sig™ generated and audited against a real authorization flow.
Request a demo, and we'll walk your fraud-prevention team through a live capture, the three-layer decomposition, and the audit trail your reviewers and regulators would actually receive.
Sources on the underlying science: Ekman & Friesen, "Nonverbal Leakage and Clues to Deception" (Psychiatry, 1969); the Facial Action Coding System; peer-reviewed literature on prosodic and acoustic markers of arousal and vocal stress; multimodal (audio-visual) emotion-recognition benchmarks demonstrating fusion gains over unimodal models; and critical reviews of microexpression reliability in high-stakes deception. Fraud-context figures from FinCEN's analysis of elder financial exploitation and the UK Payment Systems Regulator's APP-scam reimbursement regime.