Coinbase reported on Oct. 7 that a historical replay benchmark for its Onramp service found newer versions of three major AI model families (Opus 5, Sonnet 5, and GPT-5.6) caught fewer fraudulent payments and a smaller share of fraud value compared to their predecessors, despite an unchanged decision policy. Sonnet's recall fell 22.2 percentage points, and Opus's recall declined 0.8 points. While GPT's precision improved by 11.5 percentage points, its recall dropped by 20.7 points, indicating more fraud cases and value escaped detection.