The RSNA Cervical Spine Fracture CT Dataset

Summary
This dataset is composed of cervical spine CT images with annotations related to fractures; it is available at https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/ .
Key points
- This is, to our knowledge, the largest publicly available adult cervical spine fracture CT dataset, with contributions from 12 institutions across nine countries and six continents.
- This dataset includes medical images, segmentations, and expert annotations from a large cohort of radiologists with subspecialist expertise in spine imaging.
- This dataset was used successfully for the Radiological Society of North America 2022 Cervical Spine Fracture Detection competition hosted on the Kaggle machine learning platform. The dataset is made freely available to the research community for noncommercial use.
Supplemental material is available for this article.
Introduction
Cervical spine injuries are a common form of traumatic injury, affecting more than 3 million patients per year in North America 1. Cervical fractures can lead to substantial disability, as 10%–11% of all cervical spinal fractures result in spinal cord injury 2. In the United States alone, more than 1 million patients are evaluated for suspected cervical spine injury annually 3. Most of these cases are assessed in the emergent setting and, in the adult population, almost exclusively by using CT, because of its inherent improved quality and coverage relative to plain film radiography.
Interpretation of cervical spine CT can prove challenging, especially in the older population, as images are commonly confounded by superimposed degenerative disease and osteoporosis, making fracture detection more complex. Given the relatively high incidence of cervical injury in trauma patients and the potential for high morbidity, there is a need for fast and accurate diagnosis. This provides an excellent clinical use case for the assistance of an artificial intelligence (AI) algorithm. Although a few cervical spine fracture algorithms have been developed 4–6, most have limited geographic representation within the training data, restricting model generalizability. Even a model trained on a multi-institution dataset 6 had very limited diagnostic accuracy when used in practice on external datasets 7. Additionally, the lack of publicly available, expertly annotated cervical spine fracture datasets hinders further improvements in model performance using recently developed machine learning algorithms.
The Radiological Society of North America (RSNA) collaborated with the American Society of Neuroradiology (ie, ASNR) and the American Society of Spine Radiology (ie, ASSR) to create the largest publicly available, multi-institutional and multinational expert-labeled dataset of cervical spine fracture CT images for AI research, which was featured in the RSNA 2022 AI Challenge. This dataset is hosted publicly on a machine learning competition platform to help develop machine learning algorithms that can assist in the detection of cervical spine fractures. A summary of how the dataset was constructed can be found in Appendix S1 .
Dataset Description and Usage
The final dataset consisted of images of the cervical spine in Digital Imaging and Communications in Medicine (DICOM) format, two comma-separated values files, and pixel-level segmentation of the cervical spine in Neuroimaging Informatics Technology Initiative (NIfTI) format. The dataset included 3112 CT scans; demographics and frequency of fractures per cervical spine level and the number of studies from each institution are shown in Table 1 . The dataset is composed of 1445 studies positive for fracture (954 men, 491 women; mean age, 56.78 years ± 21.97 [SD]), of which 235 were bounding box annotated. This is supplemented with 1667 studies negative for fracture (1022 men, 645 women; mean age, 50.61 years ± 21.29). Table 2 shows the data distribution used for the Kaggle competition, with 2019 cases in the training set, 304 cases in the public test set, and 789 cases in the private test set.

Table 1: Demographic Distribution and Number of Positive and Negative Studies for Fracture per Institution

Table 2: Demographic and Case Distribution among Training, Public Test, and Private Test Datasets Hosted on Kaggle
Image files were organized into folders according to values stored in the Study Instance UID DICOM attribute, a unique study-level identifier. Individual image files within each folder were named according to their position within the stack of DICOM images via the Instance Number DICOM attribute.
The train.csv file contains study-level ground truth labels for the training set. Study Instance UID was the unique study-level identifier. The patient_overall column indicated if any cervical vertebrae were fractured, while the C1–C7 columns specified each level of the cervical spine. A value of 0 indicates absence and 1 indicates presence of fracture at that level.
The train_bounding_boxes.csv file contains information regarding the fracture bounding boxes for a subset of the training set. Study Instance UID is the unique study-level identifier. The x and y columns specify the upper left-hand corner position of the bounding box, or the point closest to (0, 0). The width and height indicate the bounding box dimensions. The slice_number column indicates the image number within the stack and can be concatenated with “.dcm” to generate the DICOM file name.
The segmentation files were named according to Study Instance UID and represent a subset of the training set. The segmentation labels have values of 1 to 7 for C1 to C7 (seven cervical vertebrae), 8 to 19 for the 12 thoracic vertebrae, and 0 for everything else. All segmented studies have C1 to C7 labels with variable inclusion of thoracic labels. The provided NIfTI files consisted of segmentation in the sagittal plane, while the DICOM files were provided in the axial plane. NIfTI header information was used to determine the appropriate orientation to ensure that the DICOM image and segmentation planes matched, as demonstrated in Figure 2 . Without this information, there was a risk of having the segmentations flipped in the z-axis and/or mirrored in the x-axis.

Discussion
We curated and created expert annotation of a large high-quality cervical spine fracture CT dataset from 12 institutions from six different continents, which, to our knowledge, represents the largest public dataset of cervical spine fractures currently available. Great care was also taken to ensure that data were distributed equally with respect to sex, age, contributing site, and fracture level across the training, validation, and test sets. This additional effort helped to mitigate against unexpected or untoward performance drops between training and external or internal testing. The successful production of this dataset is partially attributed to using unconventional annotation methods by means of prelabeling. Such labels were provided by contributing sites, which allowed for the redistribution of the annotation burden.
Given the size and complexity of the dataset, much time and consideration were devoted to developing an annotation strategy that maximized the use of the annotated data while avoiding overburdening our volunteer annotators. The depth of annotations from patient level to pixel level was considered. Eventually, a hybrid schema was chosen in which images from each patient in the dataset were given a study-level annotation detailing each cervical spine level that was described as fractured in the original radiologist’s report or tagged as no fracture in the control dataset. A smaller subset of patient images (approximately 16% of positive fracture cases) were assigned image-level annotations, including bounding boxes enclosing all the fractured vertebral elements on a given image. This was thought to be the best strategy to optimize the effort of the volunteers to provide “just enough” useful image-level annotations in the dataset along with the large number of additional studies with patient-level annotations. Through trial and error, the most reproducible method for image-level annotations was to have the annotators draw bounding boxes first on key sections where the fracture pattern reached a relative maximum or minimum cross-sectional area. Then the annotators skipped ahead through the image stack to the next relative maximum or minimum section and interpolated the bounding boxes in between these sections.
Establishing a strong overlap between the annotators proved to be challenging. Detailed initial instruction included example bounding boxes, a document outlining the process with image examples, and an instructional video. To help ensure accurate annotation that adhered to the provided instructions, all annotators were provided practice examinations to familiarize themselves with the tools. Performance during the practice phase was evaluated based on the ground truth bounding boxes defined by the committee, and annotators were retrained as needed. The final ground truth bounding box was calculated by taking the largest sum of all individual bounding boxes ( Fig 1 ), which focuses on the sensitivity of fracture detection. An additional subset of cases containing segmentation masks of the vertebrae was also provided so that this could be used to help train the algorithm to detect the fracture level ( Fig 2 ). Thus, this dataset provides multimodal annotation formats of different levels: patient level, vertebra level, bounding box, and segmentation.

The decision to request an abstraction of the radiology report from contributing sites was primarily to explore a different way to reduce the cognitive effort required in the time-consuming annotation process of an entire dataset from scratch. The rationale was that the report generated at the point of care offers the most accurate assessment, as this is when the radiologist is delivering professional services and attention is most concentrated on the task at hand. Experience has shown that volunteer annotators, even under the best of circumstances, are not reviewing examinations under the same level of scrutiny as they might in the clinical environment 12. Our goal was to redistribute and “front-load” the annotation burden and use our volunteer annotators in more of a quality-control activity. This approach, in addition to the requirement of smaller batches to contribute and annotate, offered the best balance without overburdening either the data contributors or the volunteer annotators. The challenge to this method is annotator disagreement with the ground truth report. In response to this, annotators were allowed to dispute the ground truth labels, which were subsequently adjudicated by organizing committee members.
In the future, the current dataset may be optimized by increasing the number and detail of image-level annotations or possibly by adding pixel-level annotations. The value of the dataset is not limited only to cervical spine fracture detection. For example, fractures outside of the cervical spine, including skull base, upper thoracic spine, and posterior rib fractures, were all commonly encountered and could be annotated as well to enhance the value of the dataset.
While the fracture-level distribution in the dataset is imbalanced, potentially affecting algorithm training and performance, the data distribution of fractures is similar to what has been described in real-world scenarios. For example, a multicenter study evaluated blunt traumatic cervical spine fractures at 21 different institutions and found that the most frequently fractured vertebrae were C2, C6, and C7, which together accounted for 63.3% of all cervical spine fractures 13. In the RSNA 2022 Cervical Spine Fracture Detection dataset, these three levels were also the most frequently fractured (with C7 being the most common) and accounted for 62.4% of all cervical spine fractures. A real-world distribution of the data is useful for clinical implementation of a fracture detection algorithm trained on the dataset.
There are several limitations of this dataset. The strict inclusion criteria of axial noncontrast 1-mm-thick section images may limit its application to practices that have different section acquisitions or reformat their CT cervical spine scans from a postcontrast acquisition. Additionally, the dataset treats acute and chronic fractures the same; however, detection of chronic fractures may not be as clinically relevant when evaluating trauma patients. Furthermore, this dataset excluded patients who underwent prior surgery because of the challenges of streak artifacts and altered anatomy. As such, machine learning models trained using this dataset may underperform on postsurgical scans of the cervical spine. Finally, evaluation for cervical spine fractures can be challenging, especially in the setting of severe trauma, and some fractures were visualized that were not accounted for in the radiologist’s report. In these cases, the radiologist’s report was chosen to represent the ground truth because of the limitations of viewing these studies in retrospect on a web-based platform. This method is obviously limited compared with radiologists reading these studies in real time on high-resolution monitors within their picture archiving and communication systems environment, with clinical history and prior imaging studies available to assist in image interpretation.
In summary, the RSNA 2022 Cervical Spine Fracture Detection dataset is, to our knowledge, the largest and most geographically diverse, publicly available expert annotated dataset of cervical spine fracture CT studies. The intent of this dataset is to inspire and enable advances in machine learning research to improve the quality, efficiency, and availability of patient care worldwide. This dataset is made freely available to all researchers for noncommercial use.
Acknowledgments
The authors would like to thank and acknowledge the contributions of Christopher Carr, MA, Sohier Dane, BA, Maggie Demkin, and Michelle Riopel.
Notas
1 The RSNA-ASSR-ASNR Annotators and the Dataset Curation Contributors are listed at the end of this article. Dataset Curation Contributors: Nitamar Abdala, Michael Brassil, Priscila Crivellaro, Allison K. Duh, Fam Ekladious, Eduardo Moreno Júdice de Mattos Farina, Mohamed Gemae, Albert Huang, Omar Islam, Nedim Kruscica, Michael Kushdilian, Robin Lee, Zamir A. Merali, Robert Moreland, Shane Natalwalla, Oleksandra Samorodova, Yekaterina Shpanskaya, Baskaran Sundaram, Suradech Suthiphosuwan, Monica Tafur, Donatella Tampieri, Jefferson Wilson, Christopher D. Witiw, Adil Zia. RSNA-ASSR-ASNR Annotators: Gennaro D’Anna, Allison M. Grayev, Fátima Hierro, Michael D. Hollander, Ichiro Ikuta, Christie M. Lincoln, Lubdha M. Shah, Achint K. Singh, Nathan S. Doyle, Luis G. Colon Flores, Vikas Agarwal, Scott R. Andersen, Katie Bailey, Gagandeep Choudhary, Sammy Chu, Charlotte Y. Chung, Andrea S. Costacurta, Muhammad Danial, Irene Dixe de Oliveira Santo, Venkata Dola, K. Jim Hsieh, Adham Khalil, Neil U. Lall, Laurent Letourneau-Guillon, David Russell Malin, J. Ryan Mason, Mariana Sanchez Montaño, Fanny E. Moron, Jaya Nath, Xuan V. Nguyen, Jacob Ormsby, Mark C. Oswood, Ozkan Ozsarlak, Samuel Rogers, Jeffrey Rudie, Anousheh Sayah, Eric D. Schwartz, Loizos Siakallis, Neil B. Horner, Rogerio Jadjiski de Leão.


