Data

Whole-Slide-Images

All images in this challenge are whole-slide-images (WSIs) generated by a Hamamatsu Nanozoomer XR or a Nanozooomer S360 at 0.23 µm per pixel (40X). Images are provided at several lower resolutions, starting at 10X and decreasing. For each image, an anonymous case id will be provided, along with information on whether the tissue in the image was stained with H&E or immunohistochemistry and if so, which IHC stain. All images originate from FFPE fixated surgical resection specimen from female primary breast cancer patients that were diagnosed in Sweden. Tumor resections, fixations, stainings and scanning were all performed by trained experts in the respective discipline.

The training data set will consist of 750 cases. For each case, one H&E WSIs and up to four IHC WSIs will be provided. Four routine diagnostic IHC stains will be included: ER, PGR, HER2 and KI67. The validation and test data set will consist of 100 and 300 cases respectively. Validation and test cases will only have one H&E WSI and one randomly selected IHC WSI available. Cases in the validation and test set are split off using random stratification based on clinical characteristics that pertain to the IHC stains. 

Annotations

For validation and test cases, we will provide a set of manually generated landmarks for IHC images that are to be registered to the H&E image domain. All landmarks will be placed as uniformly as possible in tissue regions that exist in both images. Challenge participants should then optimize algorithms such that the output of their algorithm is a set of registered landmarks that minimizes the distance to human generated landmarks in the target domain. The landmarks in the target domain will be kept secret from the challenge participants. The landmarks are intended to work as a proxy to assess the distance between corresponding tissue in the registered and target image. 

The annotation protocols will be made publicly available through Github. No annotations will be generated for the training data. Annotations for the validation and test data will be created analogously. Two or more annotators will generate the annotations, with at least two annotators per image pair. For a subset of the test cases, annotations will be generated by all annotators to investigate inter-assessor variability. The purpose of multiple annotators is to reduce the error in correctly placing corresponding landmarks. To reduce this error, one annotator will place landmarks in both images of an image pair, first in the moving and then in the target image. A second annotator will be provided with the moving image annotations of the first annotator and annotate the corresponding landmarks in the target image. All annotators will either have medical training or have worked extensively with WSIs before. 

WSIs

We have received some feedback that the SND download link below leads to corrupted files. There is also the possibility to download files from the Finnish CSC infrastructure:

train_pyramid_1_of_2.zip   461GB   31957b091febb16951cea16546cea24f (md5)
train_pyramid_2_of_2.zip   475GB   6b7f5d0e22f89dcf4afa276cb078b96c (md5)
validation_pyramid_1_of_1.zip   55.4GB   0d97997c3679a470f5628d993ee67bd9 (md5)
test_pyramid_1_of_1.zip   173GB   698392e641bab4acb0e0893a06de9e2a (md5)

Due to a slightly different process to generate the .tiff files that uses a different encoding, the CSC files linked above have larger file sizes. We apologise for the inconvenience and hope to resolve the issue with the SND download as soon as possible. 

The data is also available through the Swedish National Data Service (SND). Please make sure to double-check the sha1sums of the files with the available list, as we have received reports that some files were corrupted during download. We therefore currently recommend the download from CSC. 

Annotations to register

Validation set source points: validation_points_public_1_of_1.csv 

ACROBAT 2022 test set source points: test_points_public_1_of_1.csv