Deep learning-based denoising has enabled substantial reductions in radiotracer dose for PET imaging, yet the absence of full-dose ground truth presents challenges in evaluating image fidelity. This study proposes a supervised classification framework to assess the similarity between AI-denoised and standard-dose PET images without relying on ground truth. A total of 872 features, including radiomic descriptors, quantitative metrics, and visual quality assessments, were extracted at patient and lesion levels. Binary classifiers were trained to predict image similarity across six dose levels. Ground truth labels were derived from a structured visual scoring grid co-developed with expert users and validated against objective similarity metrics. Explainability was integrated using global and local SHapley Additive exPlanations (SHAP), calibrated similarity scores, and textual feature descriptions. The most predictive features included anatomical detail, artefact presence and textural radiomics in the lungs (patient level), and SUV skewness, detection score, and texture strength (lesion level). The combined approach enables interpretable evaluation of AI-denoised PET reconstructions in settings where objective references are lacking, supporting transparency and clinical trust.