global
Variáveis
Utilitários
ESTILOS PERSONALIZADOS

Open-Source Dataset for the RSNA Screening Mammography Cancer Detection Challenge

Hari Trivedi; Maryam Vazirabad; Felipe C. Kitamura; Yan Chen; Helen Frazer

Radiology: Artificial Intelligence - Volume 8, Number 2 - https://doi.org/10.1148/ryai.250375

Download PDF

Summary

The RSNA Screening Mammography Cancer Detection Challenge dataset contains mammograms with corresponding metadata and pathologic outcomes.

Key points

  • The RSNA Screening Mammography Cancer Detection Challenge dataset includes four-view two-dimensional mammograms with corresponding Breast Imaging Reporting and Data System assessment, breast density, machine identification, and proof of benign (follow-up or biopsy) or malignant outcome (biopsy).
  • The dataset is made freely available to the research community for noncommercial use.

Introduction

Breast cancer represents the most frequently diagnosed oncologic condition worldwide. In 2020, there were an estimated 2.3 million new cases and 685 000 deaths 1. Screening mammography has substantially reduced breast cancer mortality over time, with a 40% decrease in developed nations since the 1980s. However, screening mammography is limited by both false-positive and false-negative results 2,3. These inaccuracies can cause unnecessary stress, lead to additional health care visits, and prompt unnecessary invasive procedures such as biopsies.

Artificial intelligence (AI) holds promise in addressing several challenges associated with breast cancer screening, including enhancing detection rate, decreasing false positives and negatives, and reducing the human effort required for interpreting these examinations 47. Prospective trials in Europe have demonstrated AI as an effective tool in dual-reading environments—screening workflows in which each mammogram is independently interpreted by two radiologists—which could potentially reduce the necessary workforce for screening mammography by half and allow radiologists to focus on downstream tasks, such as diagnostic examinations or biopsies 8,9.

Despite these advances, the development and validation of robust AI models for mammography are hindered by the limited availability of large, diverse, and well-annotated public datasets. For example, the INbreast dataset contains only 410 mammography images from 115 examinations 10, whereas the Curated Breast Imaging Subset of Digital Database for Screening Mammography (or CBIS-DDSM) offers around 3000 mammography studies 11. While these datasets have contributed significantly to the field, their relatively small sample sizes, limited demographic diversity, and varying annotation protocols restrict the generalizability and clinical applicability of AI models trained on such data. To address these limitations, the Radiological Society of North America curated and released a new open-source screening mammography dataset as part of the RSNA Screening Mammography Cancer Detection Challenge 12. This large, multi-institutional dataset includes high-resolution digital mammograms, comprehensive metadata such as Breast Imaging Reporting and Data System (BI-RADS) assessments and breast density, and rigorous follow-up data from two diverse screening populations. This unique combination provides a rich resource for training new models and validating existing models.

By making this dataset publicly available, RSNA aims to accelerate the development and benchmarking of AI algorithms for breast cancer detection and to foster collaboration within the medical imaging research community. This article describes the composition, curation, and potential impact of the RSNA Screening Mammography Cancer Detection Challenge dataset, highlighting its unique features and advantages over existing resources.

Materials and Methods

The development of the dataset for the 2023 RSNA Screening Mammography Breast Cancer Detection Challenge presented unique complexities compared with prior RSNA challenges. Unlike previous competitions that relied on existing datasets, such as the pediatric hand radiographs in 2017 and the National Institutes of Health ChestNet-14 chest radiograph collection in 2018, this dataset was assembled entirely from scratch. All images were labeled with breast density, BI-RADS scores, and pathologic outcomes. Confirmation of negative examinations required longitudinal patient follow-up, as described below. Of note, mammograms pose unique technical challenges due to their large file size and high resolution (up to 12 megapixels), which exceed those of CT, MRI, and other radiographs.

Data Collection and Participants

Screening mammography examinations were collected from two clinical sites: Emory Healthcare (Atlanta, Ga) (hereby referred to as site 1) and BreastScreen Victoria (Victoria, Australia) (hereby referred to as site 2). All data were de-identified in accordance with HIPAA guidelines and institutional review board approval.

Eligible examinations were screening mammograms obtained in asymptomatic female individuals. Only technically adequate screening mammograms that contained all four screening views (craniocaudal and mediolateral oblique views of both breasts) were included. No age criteria were applied; the age range was representative of the screening population at each site. At site 1, according to national screening guidelines, screening typically begins at age 40 years although it can begin earlier in high-risk groups and there is no exclusion for female patients with prior history of breast cancer 13,14. At site 2, screening typically begins at age 50 years, although females aged 40–49 years can undergo screening after a discussion with their primary care physician. Female patients with history of breast cancer in the past 5 years are excluded from screening per national screening guidelines 15.

Selection and Enrichment of Examinations

To ensure sufficient representation of cancer-positive examinations for algorithm development, the dataset was intentionally enriched to 4% cancer prevalence, approximately five times higher than the population prevalence. At site 1, 10 000 participants with 400 cancers were sampled from the site 1 Emory Breast Imaging Dataset (or EMBED), which contains examinations from 2013 to 2020 18. From site 2, 10 000 participants with 400 cancers between 2013 and 2015 were selected. Only one examination per patient was included in the dataset.

Labeling and Classification

Screen-detected cancer examinations were defined as positive and included abnormal screening studies (BI-RADS 0) followed by pathologically proven cancer with biopsy or surgical excision. Cancer labels are assigned at the breast level (left, right, or bilateral).

Negative examinations were categorized into three groups: (a) screen negative: negative (BI-RADS 1) or benign (BI-RADS 2) screening examinations with confirmed 1-year (site 1) or 2-year (site 2) negative follow-up, (b) biopsy-proven benign: abnormal screening examinations (BI-RADS 0) followed by a benign biopsy, and (c) no significant abnormality: abnormal screening examinations (BI-RADS 0) followed by a negative (BI-RADS 1) or benign (BI-RADS 2) diagnostic examination.

Additional Metadata

BI-RADS scores are provided at the breast level for both sites. Race and ethnicity were available for site 1 examinations; for site 2, only the region of birth was available. Scanner manufacturer and breast density (as scored by the interpreting radiologist) are available only for the site 1 examinations. Finally, presence of breast implants is available at the examination level (site 1) and breast level (site 2).

Data Harmonization

Both manual and automated validation methods were used in curation of the dataset to reconcile discrepancies between imaging and clinical metadata. These include removing examinations with missing images, images with invalid or missing Digital Imaging and Communications in Medicine (DICOM) metadata, errant images (eg, diagnostic views in a screening examination), and examinations with inconsistencies between BI-RADS assessments and pathologic outcomes. Examinations that failed such checks were excluded.

Results

Dataset Characteristics

An overview of image and patient characteristics is provided in Table 1 . Initially, the collection comprised 19 794 examinations across the two sites. Following the application data harmonization steps, 376 examinations were removed resulting in a total of 19 418 examinations available for the challenge ( Figure ). Of these, 11 913 examinations were allocated to the training set, 2090 to the public test set, and the remaining 5415 examinations were used in the private test set.

Table 1: Patient and Image Characteristics

CharacteristicSite 1Site 2Overall
Demographics
Age statistics
Count9419999919418
Mean57.459.958.7
SD11.48.29.9
Minimum264040
Maximum898989
25th percentile485451
50th percentile576058
75th percentile666666
Race
African American or Black4120NA
American Indian or Alaskan Native15NA
Asian550NA
White3959NA
Native Hawaiian or other Pacific Islander77NA
Multiple32NA
Unknown, unavailable, or unreported666NA
Ethnicity
Hispanic or Latino448NA
Non-Hispanic or Latino7495NA
Unknown, unavailable, or unreported1476NA
Region of birth
OceaniaNA6830
EuropeNA1878
AsiaNA1069
AfricaNA136
AmericasNA86
Manufacturers
Hologic7783NA
General Electric1248NA
Fuji388NA
Breast density
A939NA
B4012NA
C3956NA
D496NA
Unavailable16
Findings and outcomes
Cancer384400784
Biopsy-proven benign10881001188
No significant abnormality: BR-0 followed by negative (BR-1) or benign (BR-2) on diagnostic189720003897
Negative (BR-1) or benign (BR-2)6050749913549

Note.—Site 1 location is Emory Healthcare (Atlanta, Ga), and site 2 location is BreastScreen Victoria (Victoria, Australia). BR = Breast Imaging Reporting and Data System score, NA = not available.

Image and Metadata Format

A full overview is provided on the RSNA Github 16. Mammograms are provided in DICOM format and are represented as full-field digital mammograms in presentation state, as would be viewed by the interpreting radiologist. All original metadata is included in the DICOM images, except for protected health information–containing elements such as names, dates, accession numbers, patient IDs, and addresses that were either de-identified or removed before data collection. DICOM header information was merged and manually verified or harmonized when appropriate. Only standard craniocaudal and mediolateral oblique views were included. If there were multiple craniocaudal or mediolateral views of the same breast in an examination, one was randomly sampled. The DICOM VOILut function may be used to translate DICOM pixel arrays into an optimal window-level for viewing 17. Metadata are available in .csv format and provided in Table 2 .

Table 2: Metadata

FieldDescription
site_idID code for the source site.
patient_idID code for the patient.
image_idID code for the image.
lateralityWhether the image is of the left or the right breast.
viewThe orientation of the image. The default for a screening examination is to capture two views per breast.
ageThe patient’s age in years.
implantWhether or not the patient had breast implants. Site 1 only provides breast implant information at the patient level, not at the breast level.
densityA rating for how dense the breast tissue is, with A being the least dense and D being the most dense. Extremely dense tissue can make diagnosis more difficult. Only provided for train.
machine_idAn ID code for the imaging device.
cancerWhether or not the breast was positive for cancer. The target value. Only provided for train.
biopsyWhether or not a follow-up biopsy was performed on the breast. Only provided for train.
invasiveIf the breast is positive for cancer, whether or not the cancer proved to be invasive. Only provided for train.
BI-RADS0 if the breast was scored as abnormal, 1 if the breast was scored as negative, and 2 if the breast was scored as benign. Only provided for train.
prediction_idThe ID for the matching submission row. Multiple images will share the same prediction ID. Test only.
difficult_negative_caseTrue if the examination was unusually difficult. Only provided for the training data.

Note.—Available imaging metadata, described at the image level. Because each patient only contains a single examination, patient ID can be considered analogous to an examination ID. Each examination contains four images (views) as denoted by the laterality and view columns. Certain metadata are only available for site 1 patients, as described in Table 1 . Site 1 location is Emory Healthcare (Atlanta, Ga), and site 2 location is BreastScreen Victoria (Victoria, Australia). ID = identifier.

Data Availability

The data are available for download through the AWS Open data program under a custom, noncommercial license: https://registry.opendata.aws/rsna-screening-mammography-breast-cancer-detection/ . The Github for notebooks and metadata can be found here: https://github.com/RSNA/AI-Challenge-Data/wiki/ .

Discussion

We describe the development and release of a robust multi-institution breast cancer screening dataset for the RSNA Screening Mammography Breast Cancer Detection AI Challenge. This dataset is highly enriched to contain 833 cancer examinations, 1188 biopsy-proven benign examinations, and 3897 examinations flagged as abnormal on screening but subsequently determined to be negative (BI-RADS 1) or benign (BI-RADS 2) at diagnostic imaging. As such, this provides a valuable resource for training and validation of breast cancer AI models.

The 2023 RSNA Screening Mammography Breast Cancer Detection AI Challenge (hereafter referred to simply as the Challenge) invited the scientific and industry community to develop image analysis algorithms that can detect cancer in two-dimensional digital mammograms for breast cancer screening. The Challenge was open for submissions from November 2022 through February 2023, and results publicly announced in May 2023. A total of 2146 competitors, forming 1537 teams, participated. Each team submitted their own AI algorithm and was free to use any architecture for AI image analysis. All Challenge algorithms were evaluated on the RSNA Challenge private test set, using the probabilistic F1 score as the primary metric. Some algorithms generated binary predictions, while others output continuous probability scores without a predefined recall threshold. The top-performing algorithm achieved a recall rate, sensitivity, specificity, and positive predictive value of 1.5%, 48.6%, 99.5%, and 64.6%, respectively. An ensemble of the top three algorithms increased sensitivity to 60.7% with a recall rate of 2.4% and specificity of 98.8%. Lower sensitivity was observed for the U.S. dataset than for the Australian dataset (top three ensemble model: 52.0% vs 68.1%; odds ratio, 0.51; P = .02), and greater sensitivity was observed for invasive cancers than for noninvasive cancers (top three ensemble model: 68.0% vs 43.8%; odds ratio, 2.73; P = .001) 12.

Beyond its primary use in developing and benchmarking cancer detection algorithms, this dataset enables additional research opportunities. The availability of breast density and BI-RADS scores supports studies on automated risk stratification and breast density assessment. The dataset’s diversity in patient demographics and imaging equipment allows for investigations into model generalizability and bias, while its scale makes it a valuable resource for pretraining and transfer learning.

The availability of AI models trained on the RSNA Screening Mammography Breast Cancer Detection Challenge dataset has the potential to substantially impact breast cancer detection, particularly in resource-limited and underserved regions. In many low- and middle-income countries, access to skilled radiologists and advanced imaging technologies is limited, resulting in delayed diagnoses and higher breast cancer–related mortality rates. By leveraging AI models trained on diverse, high-quality datasets such as this one, health care providers in these regions could bridge the gap created by workforce shortages and resource constraints.

This dataset contains some limitations. Because the dataset is enriched with cancer and abnormal examinations, performance of models trained using this dataset must be carefully considered. For example, metrics such as positive and negative predictive values cannot be assessed without rebalancing of the dataset. The dataset also excludes interval cancers which is an important but rare subgroup of screening examinations in which a patient develops breast cancer within 1 (United States) or 2 (Australia) years of a negative screening study. Furthermore, this dataset does not include human-provided annotations that precisely localize or delineate specific lesions within the mammograms. Since breast cancers typically occupy a very small area—often less than 1% of the total image pixels—this presents unique challenges for model development. From an architectural and computational perspective, deep learning models must process large, high-resolution images to detect these small abnormalities. If images are downsampled to reduce memory requirements or computational load, there is a significant risk that small lesions may become indistinguishable or entirely lost, potentially reducing the sensitivity of detection algorithms. As a result, specialized approaches or architectures may be needed to effectively identify small lesions in the absence of detailed lesion-level annotations.

In conclusion, the RSNA Screening Mammography Breast Cancer Detection Challenge Dataset provides a large-scale, high-quality resource for developing and validating AI models in breast cancer screening. With diverse, multi-institutional data, including high-resolution mammograms, BI-RADS assessments, and pathologically confirmed outcomes, it addresses key gaps in existing datasets by offering enriched cancer examinations and standardized metadata. This openly available dataset facilitates robust algorithm training, enabling advancements in AI-driven detection while ensuring reproducibility and benchmarking for future research in mammographic image analysis.

Acknowledgments

Special thanks to MD.ai for providing tooling for the data annotation process. Special thanks to Kaggle for providing the competition platform, as well as design and technical support.

Flowchart shows inclusion and exclusion of examinations from site 1 (Emory Healthcare, Atlanta, Ga) and site 2 (BreastSc
Flowchart shows inclusion and exclusion of examinations from site 1 (Emory Healthcare, Atlanta, Ga) and site 2 (BreastScreen Victoria, Victoria, Australia) into the final RSNA Challenge Dataset. Exclusion criteria for examinations include examinations with missing images, images with invalid or missing DICOM metadata, errant images, and examinations with inconsistencies between BI-RADS assessments and pathologic outcomes. BI-RADS = Breast Imaging Reporting and Data System, DICOM = Digital Imaging and Communications in Medicine.

Notas

Challenge Organizing Team: Katherine P. Andriole, PhD, MGH & BWH Center for Clinical Data Science; Robyn Ball, PhD, The Jackson Laboratory; Yan Chen, PhD, University of Nottingham; Helen Frazer, MBBS, BreastScreen Victoria; Tatiana Kelil, MD, University of California–San Francisco; Felipe Kitamura, MD, Universidade Federal de São Paulo; Jackson Kwok, St. Vincent’s Institute of Medical Research and University of Melbourne; Matthew Lungren, MD, Stanford University; Ritse Mann, MD, PhD, Radboud University Medical Center; John Mongan, MD, PhD, University of California–San Francisco; Linda Moy, MD, New York University Grossman School of Medicine; George Partridge, University of Nottingham; Hari Trivedi, MD, Emory University; Xin Wang, PhD, Radboud University Medical Center; Luyan Yao, PhD, University of Nottingham; Tianyu Zhang, Radboud University Medical Center. Data sharing: All data used in this study are publicly available for noncommercial use at https://www.rsna.org/en/education/ai-resources-and-training/ai-image-challenge . The dataset has been de-identified in accordance with HIPAA guidelines and is available for academic research under a data use agreement.