2 min read
Add as a preferred source on Google

Newer AI Models Missed More Payment Fraud in Coinbase Benchmark

A fixed replay of 16,140 Onramp transactions found lower fraud coverage for Opus 5, Sonnet 5 and GPT-5.6 (sol), despite GPT-5.6’s higher precision.

Payment terminal beside blank cards under cool studio light / TokenPost.ai
Payment terminal beside blank cards under cool studio light / TokenPost.ai

In a controlled replay of Coinbase Onramp activity, updated models from three major AI families caught a smaller share of fraudulent payments than the versions they replaced, highlighting potential risks from upgrades in specialized detection systems.

The evaluation covered 16,140 transactions involving 7,293 users, including 813 confirmed fraudulent transactions. Each model was tested against the same transaction group, contextual-review process and risk-to-decision policy.

Across the comparisons, Opus 5, Sonnet 5 and GPT-5.6 (sol) posted lower recall, F1 and dollar-weighted recall than the earlier versions. Recall measures the share of fraudulent transactions detected, while dollar-weighted recall measures the share of fraudulent transaction value identified.

Opus 5’s recall declined by 0.8 percentage points compared with Opus 4.5. It also posted lower precision, F1 and dollar-weighted recall, although the other changes were not quantified.

Sonnet 5’s recall fell by 22.2 percentage points, and its dollar-weighted recall dropped by 22.9 percentage points. Precision also declined.

GPT-5.6 (sol) produced a different trade-off against GPT-5.4. Its precision rose by 11.5 percentage points, while recall fell by 20.7 percentage points and dollar-weighted recall declined by 21.8 percentage points. F1 also fell.

“Newer LLMs are not necessarily better at domain-specific tasks,” Yao Ma and Xuwei Tan wrote. The benchmark showed a regression but did not establish its cause.

Possible explanations included changed decision boundaries, post-training trade-offs and different training-data mixtures. None was assigned as the reason for the results.

The test replayed nine weeks of production data collected before the risk agent was rolled out. It did not compare live customer outcomes or establish a specific amount of customer losses caused by the model changes.

Coinbase Onramp allows users to buy crypto through partner applications, including guest checkout. Its risk system combines a traditional machine-learning model and rules engine with selective large language model review of recent transaction behavior.

A separate evaluation published Oct. 8, 2026, found that a post-trained Qwen3.5-9B model exceeded Opus 4.5 on all four fraud metrics. The model posted gains of 7.5 percentage points in precision, 12.0 percentage points in recall, 9.6 percentage points in F1 and 35.4 percentage points in dollar-weighted recall.

Detection quality and speed were measured separately. Online serving tests recorded P50 end-to-end large language model request latency of 0.683 seconds for the system versus 1.515 seconds for Opus 4.5.

Simon Yoon

Reporter

Simon Yoon reports on blockchain technology for TokenPost. Send corrections or tips to info@tokenpost.com.

Loading…