ReportGem ReportGem

Academic paper

From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

Authors: Basit Alawode, Moshira Ali Abdalla, Dwarikanath Mahapatra, Muzammal Naseer, and Sajid JavedPublished: 2026-08-04Paper ID: 2608.03508Category: cs.CVLicense: CC BY 4.0

Abstract

Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution Whole Slide Images (WSIs), limiting their generalization across arbitrary resolutions. Gigapixel WSIs inherently contain diagnostic patterns at multiple scales, including cellular morphologies, tissue architectures, and global context, mirroring how expert pathologists examine WSIs. We introduce Multi-Resolution Pyramid Transformer (MRPT), a model that hierarchically aggregates multi-resolution information from cellular to tissue and WSI levels. MRPT employs a biologically meaningful Consecutive Cross-Resolution Attention (CCRA) mechanism to capture scale-independent interactions and enforces multi-resolution semantic consistency by aligning embeddings across resolutions, yielding robust and generalizable WSI representations. Pre-trained in a multi-resolution self-supervised manner on 624M patches, 2.4M regions, and 36K WSIs, MRPT learns rich coarse-to-fine histopathology features. Extensive experiments on 34 diverse datasets show that MRPT surpasses recent foundation models and Multimodal Large Language Models (MLLMs) in cancer subtype classification, tissue phenotyping, and Visual Question Answering (VQA) for WSI understanding.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader