global
Variables
Utilidades
ESTILOS PERSONALIZADOS

Beyond the <i>AJR</i>: Patrolling k-Space to Spot “Data Crimes” Using Public MRI Datasets

Paulo E.A. Kuriki; Felipe Kitamura

American Journal of Roentgenology - Volume 220, Number 2 - https://doi.org/10.2214/ajr.22.28065

Download PDF

Summary

Beyond the AJR: Patrolling k-Space to Spot “Data Crimes” Using Public MRI DatasetsPaulo E. A. Kuriki, MD1 and Felipe C. Kitamura, MD, PhD1,2Audio Available | Share

Summary of the Investigation

MR images are acquired in the Fourier domain, the so-called k-space that is later transformed into the final diagnostic images (imaging domain). Pulse sequences can be programmed to undersample the k-space, resulting in faster acquisitions. Techniques such as compressed sensing (CS), parallel imaging, and dictionary learning have previously been used to accelerate MRI acquisitions. In 2018, Zhu et al. 1 proposed to replace these earlier processing techniques with deep learning (DL) methods that transform sub-sampled k-spaces into their corresponding MR images. This concept has led to a drive to discover the best artificial intelligence model architecture for this purpose. However, as publicly available raw k-space datasets are scarce, these efforts have commonly used public repositories that offer only postprocessed MR images.

Although it is possible to back-synthesize k-space using forward Fourier transformation, those synthetic k-spaces are unlike their corresponding raw versions. MRI vendors apply proprietary processing pipelines, including imaging reconstruction, filtering, padding, and storage of only magnitude data, sometimes with lossy compression. Because of those steps, the synthesized k-space looks different from the raw version (e.g., reduced data entropy), leading to overly optimistic results. In a recent analysis, Shimron et al. 2 show artificial improvements in the performance of well-known MRI reconstruction algorithms of up to 48% and describe this misuse of public data as representing “data crimes.”

Critical Analysis

DL-based reconstruction algorithms are becoming increasingly popular. Yet, Shimron et al. 2 report two forms of algorithmic bias: one related to the effect of zero-padding the raw k-space data, a common MRI reconstruction technique implemented for accelerating image acquisition, and the other related to JPEG lossy compression. They compare algorithm performance of CS, dictionary learning, and DL between real raw data and synthesized k-space, showing how the naive use of postprocessed studies can lead to artifactual improvements that are related solely to the data processing.

Antun et al. 3 also showed that the DL-based MRI reconstruction models can be unstable. Minor changes to collected data in the k-space domain can cause significant differences in the transformed images. This problem can worsen when the acquisition data are undersampled, leading to hallucinations in reconstructed images.

Despite the described biases, comparisons solely of the relative performance of various algorithms using processed datasets remain valid, provided that researchers recognize that the results are overly optimistic of real world performance.

The same challenges in obtaining raw k-space to create data-sets may apply during later model deployment, when only postprocessed images are sent to PACS. Access to raw k-space data requires partnering with vendors to probe MRI scanners' inner workings, which may prove infeasible.

Additional approaches for MRI acceleration are to reduce the number of excitations (NEX) and then apply a DL denoising algorithm on reconstructed images or to decrease the matrix size during acquisition and then apply a superresolution DL model 4. These algorithms are applied directly to regular postprocessed images rather than to k-space data, thereby removing constraints inherent in accessing k-space data and largely eliminating the issues described by Shimron et al. 2. Such approaches also facilitate designing a DL solution that works for many MRI vendors. Although DL-based denoising techniques can be more stable (i.e., no tampering with k-space) and can reduce the risk of hallucinations, the NEX reduction decreases the SNR and makes motion artifacts more visible 5. On the other hand, a faster scan reduces patients' chances of moving.

Takeaway Point

DL reconstruction models trained on publicly available synthetic k-spaces will not perform the same on k-spaces acquired from MRI scanners. Shimron et al. 2 describe publication of these biased results as “data crimes.”

Notas

Commentary on Shimron E, Tamir JI, Wang K, Lustig M. Implicit data crimes: machine learning bias arising from misuse of public data. Proc Natl Acad Sci U S A 2022; 119:e2117203119; https://doi.org/10.1073/pnas.2117203119 . Abstract available at pubmed.ncbi.nlm.nih.gov/35312366/ Provenance and review: Solicited; not externally peer reviewed.