This project investigates whether strong benchmark performance in visual saliency modelling is accompanied by behavioural fidelity to human gaze. Using the CAT2000 eye-tracking dataset, a failure-oriented evaluation framework is applied to assess state-of-the-art saliency models using both standard benchmark metrics and additional structural measures of spatial dispersion, attention fragmentation, and centre dependence. The analysis identifies systematic mismatches between benchmark success and human-like attention patterns across different image categories, contributing to a deeper understanding of what saliency models learn and where current evaluation practices may be insufficient.