CryptoNewsFree - Real Time Crypto News Feed ← Back to Live Stream
CryptoSlate • October 10th 2026, 6:00 PM

Newer AI models missed more payment fraud in Coinbase’s benchmark

New AI Models Missed More Payment Fraud in Coinbase Benchmark

Key Summary

Coinbase reported that newer AI model versions caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service, challenging the assumption that upgrading a model improves existing payment screener. The test replayed 16,140 transactions and found that each newer version had lower recall, precision, and F1 scores, indicating that upgrading may not always lead to better fraud detection.

Please see our real time news feed on our Home Page

New AI Models Missed More Payment Fraud in Coinbase Benchmark

Methodology

Coinbase evaluated Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6. The company replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions. The replay isolated the decision model’s behavior within that setup, rather than comparing redesigned screening systems.

Results

Every newer version had lower recall, a lower combined precision-and-recall score called F1, and lower dollar-weighted recall. Sonnet’s recall fell 22.2 percentage points and its dollar-weighted recall dropped 22.9 points. Opus’s recall declined 0.8 points. Both newer models also had lower precision, meaning a smaller share of transactions they classified as fraud were actually fraudulent.

Limitations

Coinbase’s earlier online experiment compared adding selective LLM review with the existing models and rules alone. That agent-enabled flow recorded 30% fewer fraudulent transactions and 22% less fraud value. The SR-Fraud researchers say the proprietary dataset cannot be released, restricting independent replication and generalization.

Custom Model

Coinbase reported that a post-trained Qwen3.5-9B model exceeded Opus 4.5 across four fraud-detection metrics. F1 improved 9.6 percentage points and dollar-weighted recall rose 35.4 points. The company specialized it using historical fraud outcomes and deterministic rewards balancing fraudulent and legitimate examples.
#Bitcoin#US#Crypto#SEC

Latest Related Headlines