The mistakes that teach
Aggregate metrics say that a detector fails, never why. These are its most confident mistakes, ranked by confidence — a miss at P(AI)=0.51 is noise, one at 0.99 is a lesson.
Trade-offs in this approach
Robustness costs clean accuracy. Wild-simulation augmentation is applied symmetrically to both classes, which trades a fraction of a point of clean accuracy for holding up under JPEG q30 and heavy blur.
FPR control is data-bound. False positives concentrate in real images with AI-like statistics — digital art, high-ISO noise, heavy bokeh. An eval set without those classes cannot reveal them.
These FNs are the generator, not the pipeline. The three panels above are clean stills from DALL·E 3 Advanced, Midjourney 7, and Ideogram 2.0 that the global head under-scored. The heatmaps are the local head on those same forwards — a flat cool field means it missed the image everywhere, not a random overlay.