SUMMARY
OBJECTIVE: Artificial intelligence-driven conversational models are increasingly used for patient education, yet whether information patients find clear and satisfying also meets expert accuracy standards remains unclear. The aim of this study was to compare patient and physician evaluations of ChatGPT-5 and DeepSeek V3.1 responses to questions on Hashimoto's thyroiditis.
METHODS: In a cross-sectional, double-blind, within-participant study, twenty standardized, endocrinologist-developed questions across six domains were submitted to both models in October 2025. Seventeen adults with confirmed Hashimoto's thyroiditis (mean age 47.6±13.5 years; 88.2% female; 64.7% bachelor's degree or higher) rated each blinded response for clarity and satisfaction on 10-point Likert scales and selected a preferred response. Two endocrinologists also rated responses blindly; with only two physicians, their ratings were summarized descriptively. Paired t-tests with Holm correction compared ratings, and preference was analyzed with mixed-effects logistic regression.
RESULTS: Patients rated ChatGPT-5 higher than DeepSeek V3.1 for clarity (9.15±0.70 vs. 8.41±1.14; Holm-adjusted p=0.005; Cohen's d_z=0.87) and satisfaction (8.93±0.94 vs. 8.19±1.37; Holm-adjusted p=0.005; d_z=0.85), and preferred ChatGPT-5 in 68.8% of selections (odds ratio 2.20; 95%CI 1.75–2.77; p<0.001). In exploratory analyses, this advantage was concentrated among higher-education participants and was undetectable in the small lower-education subgroup (n=6). Owing to the small physician sample (n=2) and low inter-rater agreement (κ=0.12; accuracy intraclass correlation coefficient=0.48), physician accuracy ratings (5.90 and 5.70/10) are exploratory and do not support a robust patient–physician comparison.
CONCLUSION: Patients preferred ChatGPT-5 for clarity and satisfaction, though this advantage may not extend to lower-health-literacy populations; high patient satisfaction should not be taken as a proxy for clinical accuracy. Physician-supervised deployment and language-specific validation are recommended for responsible integration of artificial intelligence into patient education.
KEYWORDS:
Artificial intelligence; Patient satisfaction; Health communication; Thyroid diseases; Patient-centered care
INTRODUCTION
Artificial intelligence (AI) technologies have catalyzed transformative changes in healthcare delivery, particularly in patient information systems1,2. AI models, such as ChatGPT and DeepSeek, have emerged as potentially powerful tools for enhancing healthcare accessibility and patient education3,4. However, the reliability, accuracy, and patient acceptance of AI-generated medical information remain subjects of ongoing investigation.
Hashimoto's thyroiditis is the most common cause of hypothyroidism in iodine-sufficient regions, affecting approximately 5% of the population with a pronounced female predominance5,6. Its chronic, lifelong nature requires sustained patient self-management, dietary adherence, and medication compliance attributes that render accessible, high-quality AI-mediated health information particularly valuable in routine ambulatory care. The complexity of its management, including dietary considerations and medication interactions, makes it an appropriate condition for evaluating AI-mediated health information delivery systems7.
Given that patient knowledge and health literacy are critical determinants of how individuals comprehend and evaluate health information, clarity and satisfaction with AI-generated content may be moderated by patients' baseline knowledge and educational background, with direct implications for patient-centered care8,9. A critical and underexamined question in AI-in-medicine research is whether information that patients find clear and satisfying also meets expert standards of clinical accuracy. Patients may prioritize clarity and accessibility, whereas accurate, evidence-based content is essential for safe self-management10,11. Few studies have evaluated AI-generated medical responses from both lay and expert standpoints within a single study design, leaving the relationship between patient-perceived quality and expert-assessed accuracy poorly understood. This study addresses this gap by comparing two AI models on patient-centered dimensions while also obtaining physician accuracy assessments, in the context of patient education for Hashimoto's thyroiditis.
ChatGPT was selected as one of the most widely recognized and extensively evaluated large language models in healthcare contexts12. DeepSeek was selected on the basis of its open-source architecture and rapidly growing global adoption, representing a distinct development philosophy from proprietary models and offering a clinically meaningful benchmark for comparison13. The aim of this study was to evaluate and compare the clarity, satisfaction, and preference ratings of medical responses generated by ChatGPT-5 and DeepSeek V3.1 among patients with Hashimoto's thyroiditis, and to examine whether high patient-perceived quality corresponds to physician-assessed accuracy, in order to inform the responsible design of AI-assisted patient education systems.
METHODS
Study design and setting
This was a cross-sectional, double-blind, within-participant comparative study conducted at an endocrinology outpatient clinic. The protocol was approved by the Institutional Ethics Committee (Decision No: İ05-384-25) and conducted in accordance with the Declaration of Helsinki. Written informed consent was obtained from all participants. No a priori sample size calculation was performed; the study was designed as an exploratory pilot investigation, and post-hoc power was not computed.
Participants
Eligible patients were adults (≥18 years) with a confirmed diagnosis of Hashimoto's thyroiditis based on clinical presentation and laboratory findings (elevated TSH together with positive anti-thyroid peroxidase and/or anti-thyroglobulin antibodies), who were able to provide informed consent and had adequate language proficiency. Patients with cognitive impairment or a concurrent acute illness were excluded. Two experienced endocrinologists participated as physician assessors. Seventeen patients and two physicians were included.
Questions and artificial intelligence response generation
Twenty standardized questions on Hashimoto's thyroiditis were developed collaboratively by two endocrinologists, drawn from those most frequently encountered in routine clinical practice and distributed across six domains: disease pathophysiology (n=5), symptomatology (n=4), diagnostic approaches (n=3), treatment modalities (n=4), lifestyle modifications (n=2), and prognosis (n=2). All questions were formulated in Turkish, a deliberate choice given that the study primarily assessed patient-perceived clarity and satisfaction; querying in participants' native language reproduces the conditions under which patients would realistically interact with these tools. The full list of questions, comprising the original Turkish wording and an English translation, is available from the corresponding author upon reasonable request.
Using identical, standardized, patient-directed prompts, all questions were submitted in October 2025 to ChatGPT-5 (OpenAI) and DeepSeek V3.1 (DeepSeek), each accessed through its official public web interface. These public web interfaces do not expose user-configurable generation parameters (e.g., temperature); the providers' default settings therefore applied. Each question was submitted once and the single generated response was recorded verbatim, with no regeneration, repeated sampling, or averaging. Each model was queried in a separate, independent session to eliminate contextual carryover between interactions.
Outcomes and blinding
Model identities were concealed; responses were labeled Response A and Response B, and this anonymization was maintained throughout. Patients rated each response for clarity and satisfaction on 10-point Likert scales without knowledge of the source model. The two physicians independently rated each response for clarity, satisfaction, and accuracy on the same 10-point scales and selected a preferred response for each question, also under blinded conditions. Thus, the two physicians rated 40 response targets (20 questions×2 models).
The co-primary patient-reported outcomes were the participant-level mean clarity and mean satisfaction scores across the 20 questions. Patient preference (the response selected for each question) was a secondary outcome.
Statistical analysis
All analyses were performed on the blinded response sets (Response A and Response B). Continuous variables are presented as mean±standard deviation (SD) and categorical variables as number and percentage. Descriptive statistics were computed in IBM SPSS Statistics version 25.0; mixed-effects logistic regression, bootstrap confidence interval (CI) estimation, and generalized estimating equation analyses were performed in R version 4.5.1 (R Foundation for Statistical Computing, Vienna, Austria). A two-sided p<0.05 was considered statistically significant.
For the co-primary patient-reported outcomes, each participant rated both responses across the 20 questions, and participant-level mean clarity and satisfaction scores were calculated. Because the primary comparison was within participant, paired-samples t-tests were used to compare scores between Response A and Response B. Normality of the paired differences was assessed using the Shapiro-Wilk test, which is more appropriate than the Kolmogorov-Smirnov test in small samples; because the paired differences were not normally distributed, Wilcoxon signed-rank tests were performed as sensitivity analyses. The Holm procedure was applied to the two co-primary p-values to control the family-wise error rate. Effect sizes were expressed as paired standardized mean differences (Cohen's d_z), with 95%CIs estimated by bootstrap resampling.
For the preference analysis, each participant contributed 20 binary choices, so observations were not independent; the original chi-square approach was therefore not used as the primary inferential analysis. Instead, preference for Response A versus Response B was analyzed using mixed-effects logistic regression with a random intercept for participant. A sensitivity model additionally included a random intercept for question to account for question-level heterogeneity. Results are reported as odds ratios (ORs) with 95%CIs and p-values, and approximate intraclass correlation coefficients (ICCs) were derived from the random-effect variance components. A participant-level sensitivity analysis, using each patient's proportion of Response A selections across the 20 questions, was also performed and tested against the null value of 0.50.
Because only two physicians participated, physician ratings were summarized descriptively only and were not interpreted inferentially; no physician-versus-patient comparison was made. Inter-rater reliability between the two physicians was assessed using Cohen's kappa for the categorical preference and two-way random-effects, single-rater ICCs (absolute agreement) for the clarity, satisfaction, and accuracy ratings.
To address potential confounding by educational attainment, exploratory education-related analyses were performed. Patient-level differences between Response A and Response B were examined by education group, dichotomized as below bachelor's degree versus bachelor's degree or higher, together with an exploratory education-adjusted analysis and a sensitivity analysis restricted to participants below bachelor's level. Preference by education group was examined using a generalized estimating equation (GEE) model accounting for repeated selections. Given the small sample size and the post hoc nature of these education analyseswhich were added in response to peer review rather than prespecifiedall subgroup analyses are reported as exploratory and interpreted cautiously.
RESULTS
Study population
Seventeen patients with Hashimoto's thyroiditis were enrolled. The mean age was 47.6±13.5 years (range 26–66); 15 patients (88.2%) were women and 10 (58.8%) were married. Eleven patients (64.7%) had a bachelor's degree or higher, indicating a sample skewed toward higher education (Table 1).
Patient-reported clarity and satisfaction (co-primary outcomes)
Patients rated Response A (ChatGPT-5) higher than Response B (DeepSeek V3.1) for both co-primary outcomes (Table 2). For clarity, mean scores were 9.15±0.70 versus 8.41±1.14 (mean paired difference 0.74; 95%CI 0.30–1.18; paired t=3.57; unadjusted p=0.0025; Holm-adjusted p=0.0051; Cohen's d_z=0.87; bootstrap 95%CI 0.56–1.33). For satisfaction, mean scores were 8.93±0.94 versus 8.19±1.37 (mean paired difference 0.74; 95%CI 0.29–1.19; paired t=3.51; unadjusted p=0.0029; Holm-adjusted p=0.0051; Cohen's d_z=0.85; bootstrap 95%CI 0.55–1.29).
Shapiro-Wilk testing indicated non-normal paired differences (clarity p=0.0077; satisfaction p=0.0047). Wilcoxon signed-rank sensitivity analyses were consistent with the paired t-tests (clarity p=0.0033; satisfaction p=0.0018), supporting the robustness of the rating advantage in the full sample.
Patient preference (secondary outcome)
Patients preferred Response A in 234 of 340 selections (68.8%). Because each patient contributed 20 selections, preference was analyzed using mixed-effects logistic regression rather than a chi-square test. In the participant random-intercept model, the odds of preferring Response A over Response B were significantly greater than 1 (OR 2.20; 95%CI 1.75–2.77; p<0.001). The participant-level random-effect variance was approximately zero (approximate ICC≈0), indicating that between-patient clustering did not materially explain preference variation. In a sensitivity model additionally including a question-level random intercept, Response A remained preferred (OR 2.42; 95%CI 1.56–3.75; p=0.0001), with most residual variation at the question level (approximate ICC≈0.17). A participant-level analysis was concordant: the mean per-patient proportion preferring Response A was 68.8%, differing significantly from 0.50 (one-sample t-test p=0.00015; Wilcoxon signed-rank p=0.00076) (Table 2).
Physician ratings (descriptive only)
Physician ratings are reported descriptively because only two endocrinologists participated; no inferential patient-versus-physician comparison was made. Both models received high clarity and satisfaction ratings, whereas accuracy was rated substantially lower for both (Table 3). For Response A, mean physician ratings were 9.07±0.80 for clarity, 8.20±0.99 for satisfaction, and 5.90±1.52 for accuracy; for Response B, the corresponding values were 9.00±1.09, 8.15±1.05, and 5.70±1.07. Within the physician ratings, high clarity and satisfaction therefore coexisted with only moderate accuracy. Physician preference selections were balanced (Response A 19 of 40, 47.5%; Response B 21 of 40, 52.5%) and were not analyzed inferentially.
Physician descriptive ratings and inter-rater reliability (two endocrinologists; 40 rated targets per model).
Inter-rater agreement between the two physicians was low. For the categorical preference, Cohen's κ was 0.12 with 55% raw agreement. Single-rater ICCs (two-way random, absolute agreement) were 0.48 for accuracy, 0.20 for clarity, and 0.13 for satisfaction. This limited agreement indicates that the physician ratings cannot support a reliable group-level evaluation and are presented for descriptive context only.
Exploratory education analyses
Education-stratified sensitivity analyses indicated that the patient rating advantage of Response A was concentrated among participants with a bachelor's degree or higher. Among the 11 higher-education participants, the mean A–B difference was 1.13 for both clarity and satisfaction; among the six participants below bachelor's level, the corresponding differences were 0.02 (clarity) and 0.04 (satisfaction) and were not statistically detectable. An exploratory education-adjusted analysis was directionally consistent (higher-education coefficient 1.12 for clarity, p<0.001; 1.09 for satisfaction, p=0.0001). Because only six participants were below bachelor's level, the absence of a detectable difference in this subgroup reflects very low statistical power rather than demonstrated equivalence, and these analyses should be interpreted as exploratory only.
For preference, the proportion selecting Response A was similar across education groups (below bachelor's 78/120, 65.0%; bachelor's or higher 156/220, 70.9%), and education did not significantly predict preference in the clustered (GEE) model (OR 1.31; 95%CI 0.65–2.64; p=0.447); this moderation test is underpowered given the sample size and should not be interpreted as evidence that education does or does not modify preference.
DISCUSSION
In this study, patients with Hashimoto's thyroiditis rated ChatGPT-5 responses as clearer and more satisfying than DeepSeek V3.1 responses, with large effect sizes, and preferred ChatGPT-5 in roughly two-thirds of selections. These patient-centered findings were statistically robust: they survived Holm correction, were reproduced by nonparametric sensitivity analyses, and persisted in mixed-effects models that accounted for the repeated, clustered structure of the preference data. The consistency of the rating advantage across analytic approaches supports a genuine patient preference for ChatGPT-5 in this setting.
A central observation of this study is that high patient-perceived quality did not coincide with high physician-assessed accuracy. Although both models received high patient clarity and satisfaction scores, physicians rated their accuracy as only moderate (ChatGPT-5: M=5.90/10; DeepSeek V3.1: M=5.70/10), indicating that responses patients found clear and satisfying did not fully meet expert standards of medical accuracy. No established threshold defines an acceptable accuracy score for AI patient-education responses; "moderate" is therefore used descriptively, denoting ratings near the midpoint of the 10-point scale and well below the corresponding clarity and satisfaction ratings. This dissociation is important for clinical deployment: although patients found responses clear and satisfying, the moderate physician accuracy ratingseven with the caveat of the small physician samplesuggest that patient satisfaction alone should not be taken as a proxy for clinical reliability14. The concern is compounded by the low agreement between the two physicians on what constituted an accurate response; when experts themselves do not converge on accuracy and neither model is rated highly, relying on patient satisfaction as a surrogate for clinical reliability becomes even more hazardous. Physician oversight therefore remains indispensable when AI tools are incorporated into patient education workflows, regardless of which model patients prefer.
By contrast, the physician data did not yield an interpretable preference pattern. Only two endocrinologists participated, and their inter-rater agreement was low (preference κ=0.12, with 55% raw agreement; accuracy ICC=0.48). With so few raters and such limited agreement, the physician ratings cannot support a reliable group-level evaluation, and this study cannot make a formal patient–physician comparison. Physician results are therefore presented as descriptive context only, and the balanced physician preference distribution should not be interpreted as evidence of a systematic divergence between lay and expert judgment.
The superior clarity of ChatGPT-5 likely reflects its training emphasis on generating accessible, human-readable text, an advantage that is particularly valuable in healthcare, where patient understanding underpins treatment adherence15,16. The strong correlation between clarity and satisfaction scores supports this interpretation.
In the present sample, patient clarity and satisfaction ratings favored the same model to a similar degree, consistent with the interpretation that comprehensibility is an important driver of patient satisfaction with AI-generated health information.
These findings support user-centered design principles in AI healthcare applications. Optimal systems may require adaptive interfaces that tailor responses to their intended audience, with different optimization criteria for patient-facing versus physician-facing applications17,18.
All queries and responses were evaluated in Turkish. Large language models are predominantly trained on English-language corpora, and their performance in other languages may differ in fluency, terminology, and accuracy; this linguistic asymmetry may have influenced both the patient-perceived clarity and the physician-assessed accuracy observed here17. Evaluating responses in the language in which patients would actually use these tools strengthens the ecological validity of the patient-centered findings, but it also means that the results may not transfer directly to English-language or other-language settings.
Strengths
The present study has several methodological and contextual strengths worth acknowledging. First, the double-blind design effectively minimized evaluator bias toward brand recognition, a particularly relevant concern in AI research, in which model reputation may influence perceived quality. Second, the design combined patient-centered evaluation with independent physician accuracy assessment, allowing patient-perceived quality to be interpreted alongside, rather than in isolation from, expert judgments of accuracy. Third, participants were individuals living with a confirmed diagnosis of Hashimoto's thyroiditis rather than healthy volunteers or surrogate evaluators, ensuring that clarity and satisfaction judgments reflected the authentic informational needs and health literacy profiles of the target population. Fourth, the standardized questions were developed collaboratively by two experienced endocrinologists drawing directly from frequently encountered clinical queries, lending the question set a degree of content validity grounded in real-world practice. Fifth, conducting the study in participants' native language ensured that responses were evaluated under conditions that closely reflect how patients would realistically interact with AI tools in clinical settings.
Limitations
Several limitations of this study warrant careful consideration. Most notably, this study can be characterized as a pilot investigation: the patient sample (n=17) and the physician sample (n=2), which precludes inferential statistical analysis of professional assessments, together restrict the generalizability of findings. The physician subsample, in particular, limits conclusions regarding expert evaluation patterns; therefore, accuracy ratings were reported descriptively rather than inferentially. Replication in larger cohorts with expanded physician panels is warranted. The comparative scope of the study was limited to two AI modelsChatGPT-5 and DeepSeek V3.1and findings should not be generalized to the broader landscape of available large language models, which includes other widely used platforms such as Google Gemini and Anthropic Claude. The focus on a single chronic autoimmune conditionHashimoto's thyroiditismay limit applicability across broader clinical domains, as AI model performance may vary substantially by disease complexity, specialty context, and information density. The study was conducted exclusively in Turkish; the potential impact of language on model response quality is addressed above. The rapidly evolving nature of large language model development represents a structural limitation inherent to this research area: model versions are updated frequently, and performance benchmarks established at one point in time may not reflect subsequent iterations. Findings should therefore be interpreted as a cross-sectional snapshot of model performance rather than a definitive comparative assessment. Finally, the patient sample was predominantly highly educated (64.7% with a bachelor's degree or higher), which may have inflated clarity and satisfaction ratings and limits generalizability to populations with lower health literacygroups that may rely most on accessible patient-education materials. This concern is reinforced by prior evidence that educational background shapes how patients comprehend and evaluate health information8 and that cognitive performance varies across populations9. We also did not assess participants' baseline knowledge of Hashimoto's thyroiditis, which may have influenced their ratings. Future studies should enroll more diverse educational strata to test the robustness of these findings across health-literacy levels.
Clinical applications
The findings of this study have practical implications for the integration of AI tools into patient education in clinical settings. Given the time constraints inherent to routine ambulatory care, AI-generated health information may serve as a valuable supplementary resource, enabling patients to access clear, comprehensible explanations of their condition between consultations. However, AI-generated responses should under no circumstances replace individualized clinical assessments; the modest physician accuracy ratings observed for both models in this study serve as a critical reminder that patient satisfaction alone is an insufficient proxy for clinical reliability. Therefore, the appropriate role of AI in this setting is one of physician-supervised supplementation, supporting rather than substituting the patient–clinician relationship.
Future research
Several directions for future investigation emerge from the present findings. Studies enrolling larger patient cohorts and expanded physician panels are essential to confirm and extend the observations reported here, and to achieve the reliable inter-rater agreement needed for valid expert-level comparison. The comparative scope of AI evaluation should be broadened to include other widely available large language modelssuch as Google Gemini, Anthropic Claude, and Meta Llamato provide a more comprehensive benchmark of AI performance in patient education contexts. Bilingual and multilingual study designs would enable the direct assessment of language-specific performance differences, addressing the linguistic asymmetry concern raised in this and prior work19. Future research should also employ more comprehensive physician-oriented assessment instruments that capture dimensions such as clinical accuracy, guideline adherence, and appropriate expression of medical uncertainty, which appeared to influence expert preference behavior in ways not fully captured by the current scoring framework. Beyond comparative evaluation, longitudinal studies examining whether AI-assisted patient education translates into measurable improvements in treatment adherence and disease self-management would provide valuable evidence for the clinical integration of these tools. Finally, real-world usage studiesexploring how patients with chronic conditions naturally interact with AI platforms, how frequently they seek AI-generated health information, and how they reconcile AI responses with clinical advicewould offer important insights into the practical role of AI in ambulatory care.
CONCLUSION
This study showed that patients with Hashimoto's thyroiditis preferred ChatGPT-5 over DeepSeek V3.1 for clarity and satisfaction–an advantage that, in exploratory analyses, was concentrated among higher-education participants–supporting context-specific model selection rather than assuming the universal superiority of any single platform. At the same time, the two-physician panel rated both models' accuracy as only moderate; given the small expert sample and low inter-rater agreement, this assessment is exploratory and does not support a reliable patient–physician comparison. These findings indicate that high patient satisfaction with AI-generated responses cannot be equated with clinical accuracy. As AI tools become increasingly integrated into patient education workflows, physician oversight remains indispensablenot as a formality, but as a substantive safeguard against the gap between perceived and actual response quality. Language-specific validation is needed before such tools are deployed in non-English clinical settings.
ACKNOWLEDGMENTS
The authors thank all participants and healthcare professionals who contributed to this study. Artificial intelligence-assisted tools were used solely for English language editing and fluency improvement during manuscript preparation. All scientific content, data interpretation, and conclusions are the sole responsibility of the authors.
DATA AVAILABILITY STATEMENT
The datasets generated and/or analyzed during the current study are available from the corresponding author upon reasonable request.
REFERENCES
-
1 Davenport T, Kalakota R. The potential for artificial intelligence in healthcare. Future Healthc J. 2019;6(2):94-8. https://doi.org/10.7861/futurehosp.6-2-94
» https://doi.org/10.7861/futurehosp.6-2-94 -
2 Sezgin, E. Artificial intelligence in healthcare: complementing, not replacing, doctors and healthcare providers. Digit Health 2023;9:20552076231186520. https://doi.org/10.1177/20552076231186520
» https://doi.org/10.1177/20552076231186520 -
3 Alam L, Mueller S. Examining the effect of explanation on satisfaction and trust in AI diagnostic systems. BMC Med Inform Decis Mak. 2021;21(1):178. https://doi.org/10.1186/s12911-021-01542-6
» https://doi.org/10.1186/s12911-021-01542-6 -
4 Amann J, Blasimme A, Vayena E, Frey D, Madai VI. Explainability for artificial intelligence in healthcare: a multidisciplinary perspective. BMC Med Inform Decis Mak. 2020;20(1):310. https://doi.org/10.1186/s12911-020-01332-6
» https://doi.org/10.1186/s12911-020-01332-6 -
5 Vanderpump MP. The epidemiology of thyroid disease. Br Med Bull. 2011;99:39-51. https://doi.org/10.1093/bmb/ldr030
» https://doi.org/10.1093/bmb/ldr030 -
6 Hu X, Chen Y, Shen Y, Tian R, Sheng Y, Que H. Global prevalence and epidemiological trends of Hashimoto's thyroiditis in adults: a systematic review and meta-analysis. Front Public Health. 2022;10:1020709. https://doi.org/10.3389/fpubh.2022.1020709
» https://doi.org/10.3389/fpubh.2022.1020709 -
7 Li J, Huang Q, Sun S, Zhou K, Wang X, Pan K, et al. Thyroid antibodies in Hashimoto's thyroiditis patients are positively associated with inflammation and multiple symptoms. Sci Rep. 2024;14(1):27902. https://doi.org/10.1038/s41598-024-78938-7
» https://doi.org/10.1038/s41598-024-78938-7 -
8 Soares Junior JM, Oliveira HMC, Luquetti CM, Zuchelo LTS, Arruda Veiga EC, Raimundo JZ, et al. Adolescents' knowledge of HPV and sexually transmitted infections at public high schools in São Paulo: a cross-sectional study. Clinics (Sao Paulo). 2022;77:100138. https://doi.org/10.1016/j.clinsp.2022.100138
» https://doi.org/10.1016/j.clinsp.2022.100138 -
9 Zangirolami-Raimundo J, Raimundo RD, Silva Noll PRE, Santos WS, Leone C, Baracat EC, et al. Postmenopausal women's cognitive function and performance of virtual reality tasks. Climacteric. 2023;26(5):445-54. https://doi.org/10.1080/13697137.2023.2190511
» https://doi.org/10.1080/13697137.2023.2190511 -
10 Liu X, Keane PA, Denniston AK. Time to regenerate: the doctor in the age of artificial intelligence. J R Soc Med. 2018;111(4):113-6. https://doi.org/10.1177/0141076818762648
» https://doi.org/10.1177/0141076818762648 -
11 Kolanska K, Chabbert-Buffet N, Daraï E, Antoine JM. Artificial intelligence in medicine: a matter of joy or concern? J Gynecol Obstet Hum Reprod. 2021;50(1):101962. https://doi.org/10.1016/j.jogoh.2020.101962
» https://doi.org/10.1016/j.jogoh.2020.101962 -
12 Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183(6):589-96. https://doi.org/10.1001/jamainternmed.2023.1838
» https://doi.org/10.1001/jamainternmed.2023.1838 -
13 Tordjman M, Liu Z, Yuce M, Fauveau V, Mei Y, Hadjadj J, et al. Comparative benchmarking of the DeepSeek large language model on medical tasks and clinical reasoning. Nat Med. 2025;31(8):2550-5. https://doi.org/10.1038/s41591-025-03726-3
» https://doi.org/10.1038/s41591-025-03726-3 -
14 Panch T, Mattie H, Celi LA. The "inconvenient truth" about AI in healthcare. NPJ Digit Med. 2019;2:77. https://doi.org/10.1038/s41746-019-0155-4
» https://doi.org/10.1038/s41746-019-0155-4 -
15 Traylor DO, Kern KV, Anderson EE, Henderson R. Beyond the screen: the impact of generative artificial intelligence (AI) on patient learning and the patient-physician relationship. Cureus. 2025;17(1):e76825. https://doi.org/10.7759/cureus.76825
» https://doi.org/10.7759/cureus.76825 -
16 Pandey VK, Munshi A, Mohanti BK, Bansal K, Rastogi K. Evaluating ChatGPT to test its robustness as an interactive information database of radiation oncology and to assess its responses to common queries from radiotherapy patients: a single institution investigation. Cancer Radiother. 2024;28(3):258-64. https://doi.org/10.1016/j.canrad.2023.11.005
» https://doi.org/10.1016/j.canrad.2023.11.005 -
17 Ye C, Zweck E, Ma Z, Smith J, Katz S. Doctor versus artificial intelligence: patient and physician evaluation of large language model responses to rheumatology patient questions in a cross-sectional study. Arthritis Rheumatol. 2024;76(3):479-84. https://doi.org/10.1002/art.42737
» https://doi.org/10.1002/art.42737 -
18 Aminololama-Shakeri S, López JE. The doctor-patient relationship with artificial intelligence. AJR Am J Roentgenol. 2019;212(2):308-10. https://doi.org/10.2214/AJR.18.20509
» https://doi.org/10.2214/AJR.18.20509 -
19 Tekin S, Oguz SH, Dagdelen S. ChatGPT-4o as a digital health tool for diabetes technology education: insights on reliability, quality, and readability. Endocrine. 2025;90(2):652-9. https://doi.org/10.1007/s12020-025-04400-x
» https://doi.org/10.1007/s12020-025-04400-x
Edited by
-
Scientific Editor:
José Maria Soares Júnior https://orcid.org/0000-0003-0774-9404
