Open-access Racial bias in artificial intelligence for mammography: a challenge to diagnostic justice and equity in Brazil

ABSTRACT

Introduction:  The application of artificial intelligence to mammography has shown promise to improve breast cancer screening, reducing interpretative variability and optimizing early detection. However, variations in performance between different population groups and clinical settings raise concerns about the validity and equity of the systems used.

Objective:  To review the use of artificial intelligence in mammography, focusing on diagnostic performance, representation biases, ethical implications, and regulatory challenges for safe incorporation in Brazil.

Method:  Integrative review between January 2016 and January 2025 in the PubMed, Scopus, and Web of Science databases. Multicenter studies, systematic reviews, meta-analyses, and regulatory and ethical analyses were included. Publications without peer review and articles not directly related to mammography were excluded. After screening and critical reading, thirty-four studies were selected for qualitative synthesis.

Results:  Deep learning systems perform similarly to experienced radiologists and can reduce the number of repeat scans when used as a support for human reading. However, differences in sensitivity and specificity are observed between racial groups and demographic contexts, associated with the underrepresentation of certain populations in the training sets and the absence of external validation in different scenarios. In Brazil, the lack of representative and interoperable databases increases the risk of undetected biases and limits the generalization of results.

Conclusion:  Responsible adoption of artificial intelligence in mammography requires local clinical validation across different regions and population groups, systematic utilization of equity metrics, regulatory transparency, and ongoing human oversight. The combination of ethical governance, independent auditing, and the participation of professionals and patients is essential for technological gains to translate into sustainable and fair clinical benefits.

KEYWORDS:
Mammography; Artificial intelligence; Algorithmic bias; Racial disparities; Equity

VISUAL ABSTRACT

RESUMO

Introdução:  A aplicação de inteligência artificial à mamografia tem-se mostrado promissora para aprimorar o rastreamento do câncer de mama, reduzindo a variabilidade interpretativa e otimizando a detecção precoce. Entretanto, variações de desempenho entre diferentes grupos populacionais e contextos clínicos suscitam preocupações sobre a validade e a equidade dos sistemas utilizados.

Objetivo:  Revisar o uso de inteligência artificial em mamografia, com foco em desempenho diagnóstico, vieses de representação, implicações éticas e desafios regulatórios para a incorporação segura no Brasil.

Método:  Revisão integrativa entre janeiro de 2016 e janeiro de 2025 nas bases PubMed, Scopus e Web of Science. Foram incluídos estudos multicêntricos, revisões sistemáticas, metanálises e análises regulatórias e éticas. Excluíram-se publicações sem revisão por pares e artigos não relacionados diretamente à mamografia. Após triagem e leitura crítica, trinta e quatro estudos foram selecionados para síntese qualitativa.

Resultados:  Os sistemas de aprendizado profundo apresentam desempenho semelhante ao de radiologistas experientes e podem reduzir o número de exames repetidos quando utilizados como apoio à leitura humana. Contudo, observam-se diferenças de sensibilidade e especificidade entre grupos raciais e contextos demográficos, associadas à sub-representação de determinadas populações nos conjuntos de treinamento e à ausência de validação externa em cenários diversos. No Brasil, a carência de bases de dados representativas e interoperáveis aumenta o risco de vieses não detectados e limita a generalização dos resultados.

Conclusão:  A adoção responsável da inteligência artificial na mamografia requer validação clínica local em diferentes regiões e grupos populacionais, utilização sistemática de métricas de equidade, transparência regulatória e supervisão humana contínua. A combinação entre governança ética, auditoria independente e participação de profissionais e pacientes é essencial para que os ganhos tecnológicos se traduzam em benefícios clínicos sustentáveis e justos.

PALAVRAS-CHAVE:
Mamografia; Inteligência artificial; Viés algorítmico; Disparidades raciais; Equidade

RESUMO VISUAL

INTRODUCTION

Breast cancer remains the most common malignant neoplasm among women, accounting for about a quarter of all new diagnoses annually. In 2020, more than 2.3 million cases and approximately 685 thousand deaths were estimated worldwide.1 In Brazil, national guidelines recommend biennial mammography for women aged 50-69 years; however, population coverage is heterogeneous, with lower rates in the North and Northeast regions and among black and brown women, reflecting structural inequalities that contribute to diagnoses in more advanced stages.2-4

In recent years, artificial intelligence (AI) systems applied to mammographic interpretation have demonstrated greater diagnostic accuracy and lower interobserver variability in multicenter studies.5-7 However, external validation results reveal uneven performance between population subgroups, with reduced sensitivity and/or specificity in women from underrepresented groups, especially Hispanic and black, compared to whites.8,9 This disparity highlights limitations in the generalization of algorithms, especially in contexts with ethnic and sociocultural diversity, such as Brazil.

In Brazil, the absence of broad, interoperable national repositories stratified by race/color, combined with the dependence on algorithms developed and validated abroad, accentuates the risk of algorithmic bias and diagnostic inequity.2,10 These factors can directly impact the time to diagnosis, access to timely treatment, and, consequently, clinical outcomes.

Given this scenario, this review synthesizes the national and international literature on the use of AI in mammography, focusing on diagnostic performance, algorithmic biases, and equity, in addition to discussing ethical, regulatory, and operational implications for safe and contextualized implementation in the Brazilian health system.

METHOD

An integrative review of the literature on the application of AI in mammography was conducted, covering publications between January 2016 and January 2025. The searches were carried out in the PubMed/MEDLINE, Scopus and Web of Science databases, using combinations of the descriptors “mammography”, “artificial intelligence”, “algorithmic bias”, “racial disparities” and “equity”. The selection process followed pre-defined inclusion and exclusion criteria, prioritizing studies of high methodological relevance and thematic relevance.

Meta-analyses, systematic reviews, multicenter observational studies, technical validation trials, imaging databases, as well as regulatory and ethical analyses related to the use of AI in mammography imaging were included. Duplicate studies, non-peer-reviewed preprints, and articles whose main focus was not directly linked to mammography were excluded. At the end of the screening process, thirty-seven studies and sources were considered more relevant for the critical analysis and summarized in the Table.

TABLE
Summary of the studies incorporated in this review1-30

The full reading and critical evaluation of these studies allowed the systematic extraction of information on the composition of the datasets used, the performance metrics reported, the presence or absence of stratification by race and ethnicity, and the bias mitigation strategies described. The evidence was analyzed in an integrated manner, seeking to identify performance patterns, methodological limitations, and potential ethical and regulatory implications.

In addition, national data from official reports from the National Cancer Institute (INCA) and the Ministry of Health were incorporated, including indicators of access to screening, staging of diagnosis, and mortality from breast cancer. This integration aimed to contextualize the international findings to the Brazilian epidemiological reality.

RESULTS

The incorporation of AI systems into mammography has evolved consistently, with models based on deep learning, especially convolutional neural networks, achieving performance comparable to that of experienced radiologists in the detection of microcalcifications, architectural distortions, and masses, even in high-volume scenarios.5,11 Multicenter trials indicate that the use of AI as a tool to support human double reading increases the cancer detection rate without substantially increasing false positives or unnecessary biopsies, in addition to reducing the number of recalls and optimizing operational flows, especially in environments with a shortage of radiologists.6,7,12 However, the generalization of these systems depends critically on the diversity and quality of the training and validation sets. Widely used databases, such as DDSM, INbreast, CMMD, and OPTIMAM, have limited racial and geographic composition, often without complete sociodemographic metadata, which restricts external validity and increases the possibility of sampling bias.13-16 In line with these limitations, external validation studies document decreases in sensitivity and/or specificity in underrepresented subgroups.8,9 In Brazil, the literature points to potential operational benefits of AI, such as screening of presumably normal tests and prioritization of suspected cases, but relevant structural barriers persist: absence of interoperable and representative national repositories, dependence on imported solutions, and limited integration between information systems, which compromise broad clinical validation.2

Accumulated evidence demonstrates that the performance of AI algorithms applied to mammography can vary across population groups. This variation relates to the underrepresentation of Black, Brown, Indigenous, and other minority women in training sets and the scarcity of validation in demographically diverse contexts.10,17 In external validation of an ensemble model of deep learning in a diverse cohort (37,317 exams; Athena/UCLA network), Hsu et al. observed that the high performance reported in more homogeneous cohorts was not generalized, with significantly lower sensitivity and specificity in specific subgroups, including Hispanic women and women with a previous history of breast cancer.8 In a convergent manner, algorithms applied to historically underserved populations may have a higher rate of underdiagnosis.9 In addition to sample composition, the domain shift phenomenon reduces the generalization of models trained in homogeneous populations when applied to different epidemiological contexts.18,19 In Brazil, the lack of repositories stratified by race/color and the adoption of algorithms developed abroad without local validation increase the risk of diagnostic error in underrepresented subgroups.2 Mitigating these disparities requires stratified performance assessment and the use of equity-specific metrics, such as equalized odds, nonsense impact ratio, and demographic parity, reported prior to implementation and monitored in post-market surveillance.19,20 Even so, most published studies on AI for mammography do not report results stratified by race or ethnicity, which substantially limits the assessment of equity.20,21

Inequalities in mammographic screening in Brazil constitute a determining context for the safe incorporation of AI. Despite being provided for in the National Oncological Care Policy and guaranteed by the SUS, mammographic coverage remains below the 70% target recommended by the World Health Organization. A population-based ecological study estimated a median of only 21.6% coverage among the health micro-regions (interquartile range from 8.1% to 37.9%)2 and, according to the 2019 National Health Survey (PNS), 58.3% of women aged 50 to 69 years had a mammogram in the previous two years, with lower percentages in the North (43.2%) and Northeast (49.5%) regions.22 Analyses of the PNS of 2013 and 2019 identified brown/black color, lower income and schooling, and residence in the North and Northeast regions among the determinants of not taking the test, evidencing racial and socioeconomic inequities of access.3,23 These disparities are accentuated where there is a lower density of equipment and professionals, in addition to relevant logistical barriers, such as long commutes in rural and riverside areas.2 During the COVID-19 pandemic, the number of screening mammograms in the SUS fell by about 44% in 2020, with reductions ranging from 25% in the North to 48% in the Northeast, accompanied by an increase in the proportion of diagnoses in advanced stages.24,25 In this scenario, the adoption of AI models trained on homogeneous populations, without stratified local validation and subgroup monitoring, may inadvertently widen inequalities in accuracy and access, if mitigation measures such as representative national repositories, interoperability of systems, stratified indicators, and independent audits are not implemented.2.7

The regulation of AI applied to mammographic screening requires a framework that ensures clinical performance, safety, and equity. International guidelines converge on the need for representative pre-market validation and post-implementation surveillance with transparency of results by subgroup.20,21,26 In the United States, the Food and Drug Administration (FDA) classifies such systems as Software as a Medical Device (SaMD) and recommends validation with stratified samples, as well as post-market monitoring and version control.21,26 In the UK, the MHRA adopts principles aligned with the European AI Act, emphasizing traceability, disclosure of metrics and mitigation of inequalities.27,28 The AI Act classifies healthcare applications as high-risk, requiring algorithmic governance, robustness testing, independent audits, and reporting of equity indicators.28 In Brazil, the regulation of medical devices still lacks specific requirements for AI models, with no clear distinction between static and adaptive models and without requiring representative local clinical validation.29 This regulatory gap favors the use of algorithms trained in different demographic contexts, with potential bias described in previous studies.8,9 The equity assessment should incorporate specific metrics, such as equalized odds, disparate impact ratio, and adjusted demographic parity, with confidence intervals, calibration by subgroup, and pre-specified non-inferiority margins.18,19 For the national context, it is recommended that regulatory approval be conditioned to multicenter validation in the country, stratified by race/color, age, breast density, region, and socioeconomic level, including relevant clinical outcomes. Post-market surveillance should provide for independent audits, public stratified performance reports and recalibration mechanisms. A national registry of algorithms, containing information on data provenance, versioning, and equity metrics, would contribute to transparency and governance.30

The ethical and epistemological dimension of AI in mammography transcends technical evaluation. Algorithmic bias reflects inequalities in the processes of data production and curation, which determine who is represented and who remains invisible.10,31 Models trained in homogeneous populations tend to show performance degradation in minority subgroups, infringing the bioethical principle of justice.32,33 Automation bias, in turn, can lead to uncritical acceptance of algorithmic outputs, especially under care pressure, requiring significant human-in-the-loop oversight, transparency about uncertainties, traceability of decisions, and override protocols.34,35 In addition, there is the dimension of epistemic injustice: when certain bodies and clinical signs are underrepresented in the data, they become less recognized as legitimate sources of knowledge.17,36 In the Brazilian context, the dependence on imported solutions and the scarcity of stratified databases increase this mismatch between the model and the reality of care. Technical strategies, such as targeted data augmentation, sample reweighting, subgroup calibration, and threshold adjustments, can reduce disparities, as long as they are accompanied by equity metrics and relevant clinical outcomes.18 The adoption of ethics-by-design, transparent governance, and user and patient participation is essential for performance gains to translate into distributed benefit rather than the amplification of inequalities.30,32,33

Finally, future perspectives point out that the integration of AI into mammographic screening in Brazil has the potential to expand early detection and standardize readings, as long as it is supported by validation in diverse populations and continuous monitoring with stratified performance.18,20 The creation of multicenter and interoperable national repositories, with standardized sociodemographic annotation and detailed clinical and technical variables, is a priority.10,30 Evidence generation should occur in pragmatic trials conducted in the country, with relevant clinical outcomes and pre-specified equity metrics.19,20 In clinical operation, AI should act as a support for medical judgment, with continuous training, override protocols, and performance drift surveillance.34,35 Representative data, stratified multicenter validation, and inclusive governance are essential conditions for AI innovation to contribute to reducing, not widening, inequalities in the early diagnosis of breast cancer.10,20

As mentioned, the search included thirty-seven studies and sources, compiled and synthesized in the Table.

DISCUSSION

The findings of this review indicate that AI applied to mammography represents a substantial advance in breast cancer screening, with the potential to expand early detection and optimize diagnostic flows.5-7,12 However, the consolidation of this promise depends on structural, regulatory, and ethical conditions that ensure consistent and equitable performance.18,20,21 The literature demonstrates that, although deep learning systems can achieve accuracy similar to that of experienced radiologists, the added clinical benefit is highly sensitive to the diversity of training and validation sets.13-16 The lack of demographic representativeness, especially in relation to race, ethnicity, and socioeconomic context, reduces the external validity of algorithms and can produce systematic biases that perpetuate diagnostic inequalities.2,8-10

These disparities do not stem only from technical limitations, but from historical processes of data production and circulation that reflect preexisting social and racial inequalities.10,31 Algorithmic bias, by reproducing patterns of exclusion observed in clinical practice and health systems, acquires ethical and political relevance.37 Differences in sensitivity and specificity documented between different populations illustrate that the generalization of models without local validation can directly compromise the principle of diagnostic justice.8,9 In the Brazilian context, this issue is particularly critical: the country combines wide population heterogeneity with weak data infrastructure, resulting in a high risk of adopting models developed in different epidemiological and technical environments.2,3

The integrated analysis of the studies shows that the simple incorporation of AI into tracking does not guarantee a reduction in inequalities. On the contrary, when applied uncritically, the technology can accentuate asymmetries in access and diagnostic accuracy.2,12,19 This finding reinforces the need for a systemic approach, in which technological innovation, public policy, and ethical governance are articulated. Multicenter validation models, with stratification by subgroups and continuous monitoring, should be a regulatory requirement.20,21,26 International experience shows that robust frameworks, such as those of the FDA, the Medicines and Healthcare products Regulatory Agency, and the European AI Act, favor greater transparency, traceability, and risk mitigation.21,27,28 The absence of equivalent requirements in Brazil represents a significant gap that potentially generates inequities.29

In addition to the technical and regulatory dimensions, this review highlights epistemological aspects that are often neglected. The production of knowledge in medical AI has been marked by geographic and institutional concentration, with a predominance of data from a few centers and homogeneous populations.10,31 This concentration generates what can be described as algorithmic epistemic injustice: certain bodies and contexts remain less visible to machine learning, which restricts the legitimacy of their clinical experiences as a source of knowledge.17,36

In the case of mammography, this translates into models that are potentially less sensitive to radiological patterns associated with certain groups, age groups, or breast densities, which may reinforce historical exclusions.17 The response to this problem is not limited to technical adjustments, but requires participatory governance, ethics of co-authorship, and active involvement of affected communities in the development and validation cycle.33

From an operational point of view, evidence suggests that AI can be particularly useful in health systems with a shortage of radiologists, such as Brazil, working in the screening of presumably normal exams and in the prioritization of suspected cases.2,6,7 However, efficiency should not be pursued at the expense of equity. Reduced recalls and increased productivity, while desirable, translate into population benefit only when accompanied by equivalent performance across subgroups.18,20,21 The implementation of equity metrics, with longitudinal monitoring and public transparency, is a necessary condition to legitimize the clinical and regulatory use of AI.19,20

The ethical debate about the role of human oversight also emerges as central. Automation bias, characterized by the uncritical acceptance of algorithmic outputs, can weaken clinical judgment and obscure limitations of the model.35 Human-in-the-loop strategies, override protocols, and interfaces that make uncertainty and decision justification visible are indispensable components of responsible use.34 In addition, post-implementation surveillance should include automatic detection of data drift and performance, with objective triggers for recalibration and revalidation.18,20

The analysis of structural inequalities in Brazilian mammography screening provides essential context for the incorporation of AI. Low coverage, regional and racial barriers, and dependence on imported models indicate that technology, if poorly implemented, can reproduce and even amplify the inequalities already documented.2-4,24 On the other hand, if accompanied by public policies aimed at data representativeness, interoperable digital infrastructure, and independent auditing mechanisms, AI can work as an instrument for correcting inequities.2,28

Finally, the convergence between data science, regulatory policy, and bioethics points to the need for an ethics-by-design paradigm, in which equity is not just a result to be measured, but a constitutive principle of technological development.30,33,36 The creation of multicenter national repositories, the requirement for stratified local validation, the public transparency of metrics, and the inclusion of patients and professionals in decision-making processes are the pillars of a responsible innovation ecosystem.20,28,30 From this perspective, AI is no longer a mere instrument of diagnostic efficiency and becomes a component of a broader agenda of diagnostic justice and the strengthening of the SUS as a State policy.

In summary, the ethical and equitable integration of AI into mammography requires overcoming the logic of technological adoption centered on average performance and recognizing that the clinical value of innovation depends on its ability to benefit populations that have historically remained on the margins of diagnostic progress in a fair and sustainable way.30,32,33

CONCLUSIONS

The application of AI to mammography demonstrates potential to expand breast cancer detection and reduce recall when employed as a complement to human reading. However, the heterogeneity of performance between subgroups shows that aggregate metrics do not ensure diagnostic equity and may hide clinically relevant losses in underrepresented populations. Equity, therefore, should be treated not as a by-product of efficiency, but as a central criterion of validation and surveillance. In the Brazilian context, responsible incorporation requires local multicenter validation with stratified samples, systematic use of equity metrics, and transparent post-market surveillance by subgroup, in line with international regulatory benchmarks. The availability of representative databases, associated with robust governance and public technical documentation, is an essential condition to avoid the uncritical transfer of trained models in different demographic contexts. In clinical practice, AI should act as an instrument to support medical judgment, and not as a substitute for human decision. Continuous supervision, explicit exposure of uncertainty, and override protocols are indispensable elements of safety and legitimacy. Participatory implementation models and independent auditing strengthen trust and allow you to identify failures not captured by aggregated metrics.

References

  • 1 Deng T, Zi H, Guo XP, Luo LS, Yang YL, Hou JX, et al. Global, regional, and national burden of breast cancer, 1990-2021, and projections to 2050: a systematic analysis of the Global Burden of Disease Study 2021. Thorac Cancer. 2025;16(9):e70052. https://doi.org/10.1111/1759-7714.70052
    » https://doi.org/10.1111/1759-7714.70052
  • 2 Nogueira MC, Fayer VA, Corrêa CSL, Guerra MR, De Stavola B, Dos-Santos-Silva I, et al. Inequities in access to mammographic screening in Brazil. Cad Saúde Pública. 2019;35(6):e00099817. https://doi.org/10.1590/0102-311X00099817
    » https://doi.org/10.1590/0102-311X00099817
  • 3 Silva JL, Albuquerque LZ, Rodrigues MES, Thuler LCS, Melo ACM. Ethnic disparities in breast cancer patterns in Brazil: examining findings from population-based registries. Breast Cancer Res Treat. 2024;206(2):359-367. https://doi.org/10.1007/s10549-024-07314-w
    » https://doi.org/10.1007/s10549-024-07314-w
  • 4 Santos MO, Lima FCS, Martins LFL, Oliveira JFP, Almeida LM, Cancela MC. Estimativa de incidência de câncer no Brasil, 2023-2025. Rev Bras Cancerol. 2023;69(1):e-213700. https://doi.org/10.32635/2176-9745.RBC.2023v69n1.3700
    » https://doi.org/10.32635/2176-9745.RBC.2023v69n1.3700
  • 5 McKinney SM, Sieniek M, Godbole V, Godwin J, Antropova N, Ashrafian H, et al. International evaluation of an AI system for breast cancer screening. Nature. 2020;577(7788):89-94. https://doi.org/10.1038/s41586-019-1799-6
    » https://doi.org/10.1038/s41586-019-1799-6
  • 6 Elhakim MT, Stougaard SW, Graumann O, Larsen LB, Nielsen M, Foldager CB, et al. Breast cancer detection accuracy of AI in an entire screening population: a retrospective, multicentre study. Cancer Imaging. 2023;23(1):127. https://doi.org/10.1186/s40644-023-00643-x
    » https://doi.org/10.1186/s40644-023-00643-x
  • 7 Sharma N, Ng AY, James JJ, Khara G, Ambrózay É, Austin CC, et al. Multi-vendor evaluation of artificial intelligence as an independent reader for double reading in breast cancer screening on 275,900 mammograms. BMC Cancer. 2023;23(1):460. https://doi.org/10.1186/s12885-023-10890-7
    » https://doi.org/10.1186/s12885-023-10890-7
  • 8 Hsu W, Hippe DS, Nakhaei N, Wang PC, Zhu B, Siu N, et al. External validation of an ensemble model for automated mammography interpretation by artificial intelligence. JAMA Netw Open. 2022;5(11):e2242343. https://doi.org/10.1001/jamanetworkopen.2022.42343
    » https://doi.org/10.1001/jamanetworkopen.2022.42343
  • 9 Seyyed-Kalantari L, Zhang H, McDermott MBA, Chen IY, Ghassemi M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med. 2021;27(12):2176-2182. https://doi.org/10.1038/s41591-021-01595-0
    » https://doi.org/10.1038/s41591-021-01595-0
  • 10 Larrazabal AJ, Nieto N, Peterson V, Milone DH, Ferrante E. Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis. Proc Natl Acad Sci U S A. 2020;117(23):12592-12594. https://doi.org/10.1073/pnas.1919012117
    » https://doi.org/10.1073/pnas.1919012117
  • 11 Lundervold AS, Lundervold A. An overview of deep learning in medical imaging focusing on MRI. Z Med Phys. 2019;29(2):102-27. https://doi.org/10.1016/j.zemedi.2018.11.002
    » https://doi.org/10.1016/j.zemedi.2018.11.002
  • 12 Zeng A, Houssami N, Noguchi N, Nickel B, Marinovich ML. Frequency and characteristics of errors by artificial intelligence (AI) in reading screening mammography: a systematic review. Breast Cancer Res Treat. 2024;207(1):1-13. https://doi.org/10.1007/s10549-024-07353-3
    » https://doi.org/10.1007/s10549-024-07353-3
  • 13 Lee RS, Gimenez F, Hoogi A, Miyake KK, Gorovoy M, Rubin DL. A curated mammography data set for use in computer-aided detection and diagnosis research. Sci Data. 2017;4:170177. https://doi.org/10.1038/sdata.2017.177
    » https://doi.org/10.1038/sdata.2017.177
  • 14 Moreira IC, Amaral I, Domingues I, Cardoso A, Cardoso MJ, Cardoso JS. INbreast: toward a full-field digital mammographic database. Acad Radiol. 2012;19(2):236-248. https://doi.org/10.1016/j.acra.2011.09.014
    » https://doi.org/10.1016/j.acra.2011.09.014
  • 15 Cai H, Wang J, Dan T, Li J, Fan Z, Yi W, et al. An online mammography database with biopsy confirmed types for machine diagnosis of breast (CMMD). Sci Data. 2023;10(1):123. http://doi.org/10.1038/s41597-023-02025-1
    » http://doi.org/10.1038/s41597-023-02025-1
  • 16 Halling-Brown MD, Warren LM, Ward D, Lewis E, Mackenzie A, Wallis MG, et al. OPTIMAM mammography image database: a large-scale resource of mammography images and clinical data. Radiol Artif Intell. 2021;3(1):e200103. https://doi.org/10.1148/ryai.2020200103
    » https://doi.org/10.1148/ryai.2020200103
  • 17 Gichoya JW, Banerjee I, Bhimireddy AR, Burns JL, Celi LA, Chen LC, et al. AI recognition of patient race in medical imaging: a modelling study. Lancet Digit Health. 2022;4(6):e406-14. https://doi.org/10.1016/s2589-7500(22)00063-2
    » https://doi.org/10.1016/s2589-7500(22)00063-2
  • 18 Garrucho L, Kushibar K, Jouide S, Diaz O, Igual L, Lekadir K. Domain generalization in deep learning based mass detection in mammography: a large-scale multi-center study. Artif Intell Med. 2022;132:102386. https://doi.org/10.1016/j.artmed.2022.102386
    » https://doi.org/10.1016/j.artmed.2022.102386
  • 19 Gianfrancesco MA, Tamang S, Yazdany J, Schmajuk G. Potential biases in machine learning algorithms using electronic health record data. JAMA Intern Med. 2018;178(11):1544-7. https://doi.org/10.1001/jamainternmed.2018.3763
    » https://doi.org/10.1001/jamainternmed.2018.3763
  • 20 Kolbinger FR, Veldhuizen GP, Zhu J, Truhn D, Kather JN. Reporting guidelines in medical artificial intelligence: a systematic review and meta-analysis. Commun Med. 2024;4(1):71. https://doi.org/10.1038/s43856-024-00492-0
    » https://doi.org/10.1038/s43856-024-00492-0
  • 21 Muralidharan V, Adewale BA, Huang CJ, Nta MT, Ademiju PO, Pathmarajah P, et al. A scoping review of reporting gaps in FDA-approved AI medical devices. NPJ Digit Med. 2024;7(1):273. https://doi.org/10.1038/s41746-024-01270-x
    » https://doi.org/10.1038/s41746-024-01270-x
  • 22 Instituto Brasileiro de Geografia e Estatística (IBGE). Pesquisa Nacional de Saúde 2019: ciclos de vida - Brasil. Rio de Janeiro: IBGE; 2021.
  • 23 Silva DMd, Cavalcante YA, Oliveira BLCAd, Lopes MVdO, Fernandes AFC, Pinheiro AKB, et al. Determinantes sociais de saúde associados à realização de mamografia segundo a Pesquisa Nacional de Saúde de 2013 e 2019. Ciênc Saúde Coletiva. 2025;30(1):e11452023. https://doi.org/10.1590/1413-81232025301.11452023
    » https://doi.org/10.1590/1413-81232025301.11452023
  • 24 Furlam TO, Gomes LM, Machado CJ. COVID-19 e rastreamento do câncer de mama no Brasil: uma análise comparativa dos períodos pré-pandêmico e pandêmico. Ciênc Saúde Coletiva. 2023;28(1):223-30. https://doi.org/10.1590/1413-81232023281.06442022
    » https://doi.org/10.1590/1413-81232023281.06442022
  • 25 Rocha AFBM, Freitas-Junior R, Ferreira GLR, Rodrigues DCN, Rahal RMS. COVID-19 and breast cancer in Brazil. Int J Public Health. 2023;68:1605485. https://doi.org/10.3389/ijph.2023.1605485
    » https://doi.org/10.3389/ijph.2023.1605485
  • 26 Gerke S, Babic B, Evgeniou T, Cohen IG. The need for a system view to regulate artificial intelligence/machine learning-based software as medical device. NPJ Digit Med. 2020;3(1):53. https://doi.org/10.1038/s41746-020-0262-2
    » https://doi.org/10.1038/s41746-020-0262-2
  • 27 Medicines and Healthcare products Regulatory Agency (MHRA). Software and AI as a Medical Device Change Programme - Roadmap. London: MHRA; 17 out. 2022. Disponível em: https://www.gov.uk/government/publications/software-and-ai-as-a-medical-device-change-programme/software-and-ai-as-a-medical-device-change-programme-roadmap
    » https://www.gov.uk/government/publications/software-and-ai-as-a-medical-device-change-programme/software-and-ai-as-a-medical-device-change-programme-roadmap
  • 28 Laux J, Wachter S, Mittelstadt B. Trustworthy artificial intelligence and the European Union AI Act: on the conflation of trustworthiness and acceptability of risk. Regul Gov. 2024;18(1):3-32. https://doi.org/10.1111/rego.12512
    » https://doi.org/10.1111/rego.12512
  • 29 Jambor D. Avaliação regulatória de software como dispositivo médico com foco na inclusão e equidade nas aplicações de inteligência artificial em saúde. Rev Direito Sanit. 2025;25(1):e0017. https://doi.org/10.11606/issn.2316-9044.rdisan.2025.231708
    » https://doi.org/10.11606/issn.2316-9044.rdisan.2025.231708
  • 30 Gebru T, Morgenstern J, Vecchione B, Vaughan JW, Wallach H, Daumé III H, et al. Datasheets for datasets. Commun ACM. 2021;64(12):86-92. http://doi.org/10.1145/3458723
    » http://doi.org/10.1145/3458723
  • 31 Mittelstadt BD, Allo P, Taddeo M, Wachter S, Floridi L. The ethics of algorithms: mapping the debate. Big Data Soc. 2016;3(2):2053951716679679. http://doi.org/10.1177/2053951716679679
    » http://doi.org/10.1177/2053951716679679
  • 32 Chen IY, Joshi S, Ghassemi M. Treating health disparities with artificial intelligence. Nat Med. 2020;26(1):16-7. https://doi.org/10.1038/s41591-019-0649-2
    » https://doi.org/10.1038/s41591-019-0649-2
  • 33 Gerke S, Minssen T, Cohen G. Ethical and legal challenges of artificial intelligence-driven healthcare. In: Bohr A, Memarzadeh K, editors. Artificial Intelligence in Healthcare. London: Academic Press; 2020. p. 295-336. https://doi.org/10.1016/B978-0-12-818438-7.00012-5
    » https://doi.org/10.1016/B978-0-12-818438-7.00012-5
  • 34 Lyell D, Coiera E. Automation bias and verification complexity: a systematic review. J Am Med Inform Assoc. 2017;24(2):423-31. https://doi.org/10.1093/jamia/ocw105
    » https://doi.org/10.1093/jamia/ocw105
  • 35 Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J Am Med Inform Assoc. 2012;19(1):121-7. https://doi.org/10.1136/amiajnl-2011-000089
    » https://doi.org/10.1136/amiajnl-2011-000089
  • 36 De Proost M, Pozzi G. Conversational artificial intelligence and the potential for epistemic injustice. Am J Bioeth. 2023;23(5):51-3. https://doi.org/10.1080/15265161.2023.2191020
    » https://doi.org/10.1080/15265161.2023.2191020
  • 37 Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-53. https://doi.org/10.1126/science.aax2342
    » https://doi.org/10.1126/science.aax2342
  • How to cite this article
    Farias LGM, Siqueira ACC, dos Santos KFA, Zini C, Linder T. Viés racial na inteligência artificial para mamografia: um desafio à justiça e à equidade diagnóstica no Brasil. BioSCIENCE. 2026;84:e00019. https://doi.org/10.55684/2026.84.pt.e00019
  • Central Message
    The application of artificial intelligence to mammography has shown promise to improve breast cancer screening, reducing interpretive variability and optimizing early detection. However, variations in performance between different population groups and clinical settings raise concerns about the validity and equity of used systems.
  • Perspective
    Responsible adoption of artificial intelligence in mammography requires local clinical validation across different regions and population groups, systematic utilization of equity metrics, regulatory transparency, and ongoing human oversight. The combination of ethical governance, independent auditing, and the participation of professionals and patients is essential for technological gains to translate into sustainable and fair clinical benefits.
  • Funding:
    None
  • Data availability:
    Data are available from the corresponding author upon reasonable request.

Edited by

Data availability

Data are available from the corresponding author upon reasonable request.

Publication Dates

  • Publication in this collection
    21 Aug 2026
  • Date of issue
    2026

History

  • Received
    27 Apr 2026
  • Accepted
    25 May 2026
  • Published
    12 June 2026
location_on
Associação Médica do Paraná - AMP Rua Cândido Xavier, 575 - Água Verde. , Cep: 80240-280 , Tel: +55(41)3024-1415 - Curitiba - PR - Brazil
E-mail: bioscience@bioscience.org.br
rss_feed Stay informed of issues for this journal through your RSS reader
Go to top Report error