Open-access Problem-solving in pre-service physics teacher education: A didactic analysis of the role and use of metacognitive skills

Abstract

The ability to generate robust analogical reasoning is a cornerstone of physics education, particularly for elucidating abstract and counterintuitive concepts. This study investigates the differential impact of artificial intelligence (AI)-assisted versus traditional face-to-face (FF) problem-solving modalities on the capacity of pre-service physics teachers to generate analogical reasoning for learning thermodynamics. Fourteen participants from physics and mathematics teacher education programs were assessed, revealing distinct, modality-dependent performance patterns. The results indicated that participants utilizing the traditional FF problem-solving modality demonstrated superior success in formulating analogies for concepts of low-to-intermediate difficulty. Conversely, a concordant failure was observed across both modalities on high-complexity items, indicating persistent cognitive barriers concerning entropy and irreversibility, regardless of the problem-solving approach. Significantly, the AI-assisted modality enabled marginal success (8–17%) on analogical tasks of extreme complexity, where the FF modality yielded no success. This evidence suggests that neither traditional problem-solving paradigms nor AI-driven platforms automatically confer the sophisticated analogical reasoning requisite for a deep understanding of thermodynamics. Consequently, this study advocates for an evidence-based hybrid pedagogical model that strategically integrates explicit training in analogical reasoning with FF interaction for foundational concepts and AI-powered support for advanced and complex problem-solving, thereby addressing a critical gap in physics teacher education.

Keywords:
Pre-service physics teachers; artificial intelligence in education; thermodynamics education; analogical reasoning; teacher education

1. Introduction

Problem-solving constitutes a foundational pedagogical practice in physics and mathematics education, serving not merely as an assessment tool but as a vehicle for developing essential scientific competencies including model formulation, quantitative reasoning, and knowledge transfer to novel contexts [1]. Its significance has gained international recognition through prominent assessment frameworks such as the Programme for International Student Assessment (PISA), which since 2015 has emphasized collaborative problem-solving and capacity to address non-routine problems requiring adaptive reasoning [2].

Despite the formative value of problem-solving pedagogy, extensive research has documented persistent challenges faced by physics students – even at university level – in successfully tackling complex problem-solving tasks [3, 4]. These difficulties stem not solely from gaps in conceptual understanding but fundamentally from limited metacognitive regulation of the planning, monitoring, and evaluative processes essential for strategic problem-solving [5, 6]. Students often approach problems through surface-level pattern matching and equation manipulation rather than engaging in deep conceptual analysis and systematic strategic reasoning [7].

Empirical literature consistently demonstrates that metacognitive skill development enhances learning quality and directly correlates with academic success across educational levels and domains [7, 8]. In this context, while frameworks like Self-Regulated Learning (SRL) and constructs such as self-efficacy provide a broad backdrop for understanding student motivation, this study specifically narrows its focus on the metacognitive regulation components – planning, monitoring, and evaluation – as the operational levers for problem-solving in physics.

The formation of competent physics teachers represents a critical component for strengthening educational systems and advancing scientific literacy [9, 10]. However, research reveals that many pre-service teachers hold persistent misconceptions about fundamental physical phenomena, particularly those involving non-intuitive scales or abstract concepts such as thermodynamics [11, 12]. These conceptual difficulties, compounded by low motivation and inadequate self-regulatory strategies, pose significant obstacles to their development as effective educators [13, p. 247]. Beyond mastering physics content, pre-service teachers must develop sophisticated pedagogical content knowledge including facility with analogies, representations, and explanatory strategies that make complex physics accessible to diverse learners [14].

In recent years, the emergence of artificial intelligence (AI) technologies has introduced transformative possibilities for educational innovation. Generative AI tools such as ChatGPT, Claude, and specialized educational platforms have demonstrated capacity to provide personalized explanations, adaptive scaffolding, immediate feedback, and on-demand tutoring at scale [15, 16]. These affordances suggest potential for AI to democratize access to high-quality learning support, particularly in resource-constrained contexts where individualized human tutoring proves infeasible. However, the educational effectiveness of AI tools compared to traditional face-to-face (FF) instruction remains incompletely understood, particularly for complex domains such as physics teacher education where both disciplinary mastery and pedagogical skill development are essential [17, 18].

Existing research comparing AI-supported learning to traditional instruction shows mixed results that vary substantially by implementation characteristics, disciplinary context, and outcome measures. Meta-analyses reveal modest positive effects of AI interventions overall [15], but these effects prove heterogeneous and contingent on factors including integration with human instruction, adaptivity of AI systems, and nature of learning outcomes assessed. Notably, AI tutoring systems perform comparably to human tutors for well-structured procedural knowledge but significantly underperform for conceptual understanding and transfer tasks [17] – precisely the competencies most critical for physics teacher education.

A significant gap in existing literature concerns the comparative effectiveness of AI versus face-to-face instruction specifically within science teacher education contexts. While substantial research examines AI applications for student learning, relatively few studies investigate AI’s role in developing the dual expertise required of teachers: deep disciplinary understanding coupled with pedagogical reasoning about how to teach that content effectively [9]. Physics teacher education presents unique demands including mastery of counter-intuitive concepts, facility with multiple representations, and capacity to construct pedagogical tools such as analogies that bridge abstract physics principles to accessible everyday experiences.

This study addresses this gap by systematically comparing AI-supported learning versus face-to-face instruction in developing thermodynamics conceptual understanding and analogical reasoning among pre-service physics teachers. Thermodynamics provides an ideal domain for this investigation because it involves abstract concepts (entropy, irreversibility) well-documented as challenging even for advanced students, requires integration of macroscopic and microscopic perspectives, and demands sophisticated reasoning about complex systems – cognitive demands that differentially test the capabilities of AI versus human instruction.

The study employs a cross-sectional design comparing two problem-solving modalities through analogical reasoning assessment: AI-assisted and face-to-face problem-solving. While this natural experiment design limits causal inference due to systematic differences between modalities – including students’ prior knowledge – it offers valuable ecological validity by examining learning outcomes under authentic implementation conditions rather than in controlled laboratory environments.

Understanding the differential strengths and limitations of AI-supported versus face-to-face solving problem can inform evidence-based instructional design in science teacher education. Specifically, identifying which competencies develop effectively through each modality – and which prove resistant to both without additional intervention – can guide strategic integration of AI tools while preserving essential elements of human mentorship in teacher preparation. This knowledge proves increasingly urgent as higher education institutions navigate post-pandemic decisions about permanent incorporation of educational technologies developed during emergency remote teaching.

This investigation is guided by three research questions: Q1: To what extent does the efficacy of AI-assisted and traditional face-to-face (FF) problem-solving modalities differ in developing pre-service physics teachers’ ability to generate analogical reasoning for thermodynamic concepts of varying conceptual difficulty?, Q2: How do the performance patterns between the AI-assisted and FF modalities diverge, particularly regarding the generation of analogies for high-complexity concepts such as entropy and irreversibility?, Q3: How do the inherent limitations of each problem-solving modality-as evidenced by concordant failures-explain the persistence of fundamental cognitive barriers to mastering thermodynamics?

These questions guide our empirical investigation and inform interpretation of comparative performance patterns between modalities.

2. Theoretical Framework

2.1. Metacognition and self-regulated learning in physics education

This study adopts a comprehensive conception of metacognition as the capacity to know, monitor, regulate, and evaluate one’s own cognitive processes, encompassing both metacognitive knowledge and regulation [13, 19, p. 21]. Metacognitive knowledge refers to what individuals understand about their cognition (knowing what, how, when, and why), while metacognitive regulation involves active control processes such as planning, monitoring, and evaluation [20].

These capacities are particularly critical in scientific problem-solving, directly influencing how individuals interpret problems, formulate solution plans, and review outcomes [21]. Students with well-developed metacognitive skills approach problems systematically, whereas those with limited development often resort to superficial strategies like pattern matching without strategic planning [7].

Recent large-scale research has established a robust causal link between metacognitive instruction and academic achievement. A comprehensive meta-analysis by [22] found that metacognitive strategy instruction yields substantial and lasting positive effects on academic performance. Crucially, they demonstrated that interventions incorporating explicit instruction in planning, monitoring, and evaluation strategies were significantly more effective than those relying merely on prompts or questioning. This suggests that effective metacognitive development requires systematic instruction and guided practice, rather than simply encouraging students to “think about their thinking.”

In the specific context of physics education, research highlights a critical gap between students’ perceived and actual metacognitive skills [21] found that students’ self-reports of their strategic abilities correlated only moderately with their actual behavior during problem-solving. This “illusion of understanding” underscores the need for assessments that capture enacted metacognition. Their work also demonstrated that actual metacognitive skillfulness was a strong predictor of problem-solving performance, accounting for a significant portion of variance beyond domain knowledge alone. This establishes metacognitive competence as a distinct skillset essential for achievement in physics.

The Self-Regulated Learning (SRL) framework provides a more holistic lens by integrating metacognitive, motivational, and behavioral dimensions [23, p. 13; 24]. SRL is a cyclical process involving forethought (goal-setting, strategic planning), performance (self-control, self-observation), and self-reflection (self-judgment, adaptation). A key component within this framework is self-efficacy – an individual’s belief in their capacity to succeed [25]. Within the broader scope of Self-Regulated Learning (SRL), which describes a cyclical process of goal setting and adaptation [26], p. 299], our investigation specifically isolates the cognitive and metacognitive strategies employed during the problem-solving phase. While motivational beliefs (self-efficacy) are acknowledged as drivers of this process [25, 27], this study concentrates empirically on the observable manifestation of metacognitive skills during the construction of physical explanations and analogies. Research in physics education confirms the pivotal role of self-efficacy [27], using structural equation modeling, found that self-efficacy was a significant predictor of physics achievement and partially mediated the relationship between metacognitive strategy use and performance. This suggests that effective strategies enhance achievement partly by bolstering students’ confidence. Further, longitudinal work by [28] revealed that self-efficacy is not static but fluctuates, creating “virtuous cycles” where strategic effort leads to success, which in turn enhances self-efficacy for future challenges. However, they also discovered that these benefits are highly domain-specific; self-efficacy in mechanics did not transfer to thermodynamics. This finding implies that robust teacher education must cultivate self-regulation and self-efficacy across diverse physics domains.

2.2. Analogical reasoning as a pedagogical tool in physics education

Analogical reasoning – the process of transferring knowledge based on structural similarities between a familiar source and an unfamiliar target – is a fundamental mechanism for learning in science [29]. In physics, analogies are powerful pedagogical tools for connecting abstract principles to concrete experiences [30]. According to structure-mapping theory, successful analogy requires mapping the underlying relational structure, not superficial features. Novices, however, often fail precisely because they fixate on surface similarities, leading to misconceptions [31].

Effective analogy instruction must therefore be explicit. Research by [32] showed that when teachers explicitly highlighted the structural relationships between domains – and, critically, the limitations where the analogy breaks down – student comprehension improved substantially. A comprehensive review by [33] synthesized key principles for effective analogy use, including being systematic, discussing similarities and differences, evaluating limitations, and providing multiple complementary analogies.

Studies implementing these principles, such as the Teaching with Analogies (TWA) model, have demonstrated significant learning gains in thermodynamics compared to traditional instruction [34]. The consistent advantage of structured analogy instruction across multiple physics domains underscores its pedagogical power when taught explicitly rather than used casually [35, p. 11].

However, analogies present significant challenges, especially for complex topics like entropy. This well-documented challenge in teaching entropy and the Second Law of Thermodynamics [36, 37] motivates our specific focus on this high-difficulty concept, as it serves as an ideal test case for assessing the limits of different pedagogical modalities. Traditional approaches often fail to convey its essential statistical-mechanical foundations. For example, innovative research in the Brazilian context has highlighted the necessity of pedagogical approaches that explicitly bridge macroscopic observations with microscopic statistical models to build deep conceptual understanding [36]. This well-documented challenge in teaching entropy motivates our specific focus on this high-difficulty concept, as it serves as an ideal test case for assessing the limits of different pedagogical modalities.

These findings underscore that generating and using analogies is a sophisticated skill, not an intuitive one. For pre-service teachers, this competency is critical but challenging to develop. Recent work by [38] revealed that pre-service science teachers struggled to generate accurate analogies, identify their limitations, and avoid creating mappings that could introduce misconceptions. Importantly, strong disciplinary knowledge did not guarantee superior analogical reasoning skills, indicating that pedagogical competence with analogies requires distinct training beyond content mastery. This directly motivates the present study’s focus on developing this critical competency.

2.3. Artificial intelligence in STEM education: Affordances and limitations

The advancement of generative AI has opened transformative possibilities for education [16]. A landmark meta-analysis by [15] found a modest but significant positive effect of AI interventions on learning outcomes. Critically, their analysis revealed that AI tools were most effective when blended with human instruction, when providing adaptive scaffolding, and when supporting practice and application rather than delivering primary instruction. This suggests AI’s value lies in augmenting, not replacing, human pedagogy.

In STEM disciplines, the role of AI is nuanced. A review by [17] found that for well-structured procedural knowledge, AI tutors performed on par with human tutors. However, for developing conceptual understanding and knowledge transfer – tasks requiring deep reasoning – AI tutors significantly underperformed their human counterparts. The researchers attributed this gap to the sophisticated pedagogical reasoning of human tutors, who can diagnose misconceptions, adapt explanations, and use strategic questioning in ways current AI systems cannot replicate. AI tutors tend to operate via pattern matching, which is less effective when students’ difficulties deviate from anticipated patterns.

[16] provide a balanced overview of large language models in education, identifying key affordances (e.g., immediate, personalized support) alongside critical risks (e.g., factual inaccuracies or “hallucinations,” fostering over-reliance). This highlights the need for explicit instruction in AI literacy, where learners are taught to critically evaluate AI outputs.

A systematic review by [18] identified significant gaps in the literature that this study aims to address. They found a scarcity of research directly comparing AI-supported learning to face-to-face instruction, particularly in the context of teacher education. The existing research has disproportionately focused on AI for student learning rather than for teacher preparation – a critical omission, as teachers must be prepared to integrate these tools pedagogically. The present study directly addresses these gaps by comparing AI-assisted and face-to-face problem-solving modalities within physics teacher education, examining the development of both conceptual understanding and the pedagogical skill of analogical reasoning [39].

3. Research Objectives and Hypotheses

3.1. Research objectives

This study is guided by an overarching aim: to comparatively evaluate the efficacy of AI-assisted and traditional face-to-face problem-solving modalities in developing the analogical reasoning of pre-service physics teachers for the learning of thermodynamics. To achieve this aim, the following specific objectives are pursued, each corresponding directly to our research questions:

  1. To comparatively evaluate the effectiveness of AI-assisted versus traditional face-to-face (FF) problem-solving modalities in developing participants’ ability to generate analogical reasoning for thermodynamic concepts of varying conceptual difficulty (addresses RQ1).

  2. To identify and analyze the patterns of divergence and convergence in performance between the two modalities, particularly concerning the generation of analogies for high-complexity concepts such as entropy and irreversibility (addresses RQ2).

  3. To elucidate how the inherent limitations of each problem-solving modality, as evidenced by their performance patterns, explain the persistence of fundamental cognitive barriers to mastering thermodynamics (addresses RQ3).

3.2. Research hypotheses

Based on the theoretical framework, the following hypotheses are proposed:

H1: The traditional face-to-face (FF) problem-solving modality will yield superior performance in the generation of analogies for thermodynamic concepts of low-to-intermediate difficulty compared to the AI-assisted modality. This hypothesis is predicated on the established role of human-led interaction in providing adaptive pedagogical scaffolding, nuanced diagnostic feedback, and multimodal communication cues. These affordances, which are critical for establishing foundational knowledge, remain challenging for current AI systems to replicate effectively [17].

H2: Both the AI-assisted and FF modalities will result in comparably low performance on tasks of high conceptual complexity, demonstrating a strong concordance in failure. This outcome is anticipated due to the profound and well-documented cognitive barriers associated with advanced thermodynamic concepts like entropy [30, 40]. Overcoming these barriers likely requires targeted, specialized pedagogical interventions that fall outside the scope of either a generalized AI tool or a traditional, unstructured problem-solving dialogue.

H3: The AI-assisted modality will enable marginal, yet non-zero, success on analogical reasoning tasks of extreme complexity, whereas the FF modality will yield zero success on these same items. This hypothesized slight advantage is attributed to the unique affordances of AI, specifically its computational capacity to rapidly access, process, and synthesize vast interdisciplinary datasets. This capability may facilitate the construction of novel, non-obvious analogical mappings that are intractable through conventional human pedagogical reasoning alone [16].

4. Research Methodology

4.1. Research design

This study adopts an Exploratory Comparative Case Study design. This methodological approach was selected to conduct an in-depth investigation into the differential efficacy of two distinct problem-solving modalities – AI-assisted versus traditional face-to-face (FF) – within a specific, real-world educational context. Given the participant pool of 14 pre-service physics teachers (N = 14), this design functions as a robust pilot study, prioritizing the richness of qualitative cognitive data and the identification of reasoning patterns over broad statistical generalization. The study involves a within-subjects comparison, where the same cohort participated in both experimental conditions, allowing for a direct contrast of cognitive trajectories.

4.2. Participants and context

The research was conducted at the University of Antioquia (Colombia), within its established physics teacher education program. The participant sample consisted of 14 pre-service teachers (N=14) from the bachelor’s programs in Physics and Mathematics Education, selected via an open call. The sole inclusion criterion was the prior successful completion of at least one university-level thermodynamics course. Data were collected without any preceding training intervention, thereby capturing the participants’ extant competencies as they engaged with each modality.

4.3. Instruments and measures

Conceptual understanding questionnaire: A 21-item assessment, administered in Spanish, was used to measure participants’ post-intervention conceptual mastery of thermodynamics, adapted from the validated Standardized Test of Thermodynamics for First-year Students at University Level (STPFaSL).

Analogical reasoning task: An open-response instrument required participants to generate and justify pedagogical analogies for thermodynamic phenomena of varying complexity. This task served as the primary measure of performance within each modality and included two extreme-difficulty items (Maxwell’s Demon; urban heat island) designed to probe the limits of analogical reasoning (see Appendices A and B, respectively). Responses were scored using a detailed analytic rubric.

Think-Aloud protocols: Concurrent think-aloud protocols were employed during the analogical reasoning tasks. Participants were instructed to continuously verbalize their thoughts, providing direct access to reasoning strategies and cognitive obstacles [37, 38]. Sessions were audio-recorded for analysis.

4.4. Procedure

Data collection followed a fixed, sequential procedure for each participant spanning a period of two months to prevent cognitive fatigue and ensure high-quality engagement: (1) Informed consent was obtained. (2) In weekly sessions, participants completed the Analogical Reasoning Task under the AI-Assisted Modality, using Claude. To avoid “hallucinations” or irrelevant deviations, the AI was pre-configured with specific prompts and constraints designed to focus the dialogue on thermodynamic analysis. (3) Subsequently, participants completed the same task under the Face-to-Face Modality. The human facilitators were experienced Physics Education professors and public school teachers who underwent a rigorous training process over several weeks. This training included sensitization to the project goals and active participation in the design of the instruments (specifically for the Urban Heat Island and Maxwell’s Demon scenarios) to ensure standardized pedagogical scaffolding. (4) Finally, participants completed the Conceptual Understanding Questionnaire. A fixed order of conditions was used, and this potential limitation is addressed in the discussion of results.

4.5. Data analysis

The analysis plan integrated quantitative and qualitative methods to provide a comprehensive evaluation of the research hypotheses.

Data Preparation and Reliability: All open-response data were independently scored by two trained raters. Inter-rater reliability was calculated using Gwet’s AC1 coefficient, chosen for its stability and robustness with small samples and imbalanced data, which is common in tasks with very high or very low success rates [41].

Quantitative analysis: Statistical analyses were performed using JASP, Jamovi, and R-Studio.

Descriptive Statistics: Means, standard deviations, and frequency distributions were calculated to profile performance.

Item-Level comparative analysis: To compare the dichotomous success rates (Yes/No) on individual conceptual and analogical items between the two paired conditions (AI vs. FF), McNemar’s test was employed. This non-parametric test is specifically designed for analyzing paired categorical data to detect significant changes in proportions. Cramér’s V was calculated as a measure of the effect size for these comparisons.

Concordance analysis: To assess the degree of agreement in outcomes (both success and failure) between the two modalities on an item-by-item basis, Gwet’s AC1 coefficient was again utilized. This analysis moves beyond simple performance comparison to quantify the extent to which the two modalities produce the same result for the same participant, with specific attention paid to “concordance in failure” on high-difficulty items.

Qualitative analysis: The verbatim transcripts from the think-aloud protocols and the content of the generated analogies were analyzed to construct latent semantic networks for the two most complex tasks (Maxwell’s Demon and urban heat island). This qualitative technique synthesizes the collective understanding of the participants, mapping the convergences (shared core concepts) and divergences (unique interpretations and analogies) in their reasoning.

The construction of latent semantic networks was performed through computational analysis of unstructured textual data using a combination of Python (version 3.10) and JavaScript libraries. Specifically, we employed Python’s Natural Language Processing (NLP) libraries – including NLTK (Natural Language Toolkit) and spaCy – for initial text preprocessing, tokenization, and extraction of key conceptual elements from participants’ verbal protocols and written responses. The extracted concepts, mechanisms, and analogies were then structured into network representations using NetworkX (Python library for network analysis) to establish nodes (representing distinct conceptual elements) and edges (representing explicit relationships articulated by participants). For visualization, we utilized the ForceAtlas2 physics simulation algorithm, implemented via JavaScript D3.js library, which employs gravitational and repulsive forces to spatially position strongly connected nodes closer together, thereby naturally clustering convergent understanding while separating divergent interpretations. This approach produces an intuitive visual representation of the collective conceptual landscape that reveals both shared core concepts (convergences) and unique individual reasoning pathways (divergences). The resulting networks were exported as static two-dimensional representations suitable for publication, though the underlying computational models remain dynamic and interactive in the original code.

5. Results

This section presents the study’s results, beginning with an assessment of the participants’ baseline conceptual understanding to contextualize the main experimental findings. Following this, the section presents the findings organized to reveal how the pre-service teachers mobilized their thermodynamic knowledge and analogical reasoning when confronted with two complex, open-ended problems: the Maxwell’s Demon paradox and the urban heat island phenomenon (see Appendices for complete problem descriptions).

It is crucial to clarify that the conceptual (designated C1–C8) and analogical (A1–A8) categories presented below were not direct questions but rather emerged from an inductive analysis of the 14 participants’ responses. After collecting the solutions generated in both the AI and FF modalities, they were systematically coded to identify the key concepts and types of analogies that the pre-service teachers used to construct their arguments. Therefore, the results do not reflect success on a series of discrete items but rather the frequency and quality with which participants activated and articulated these conceptual and analogical constructs within their solutions to the two main problems. In the analyses, a response was coded as “Yes” (success) for a conceptual or analogical category if the participant demonstrated a correct and robust application of said concept or analogy in service of solving the target problem.

5.1. Baseline conceptual understanding of thermodynamics

Prior to the experiment, the conceptual understanding of thermodynamics of the 14 pre-service teachers was assessed using a 21-item questionnaire adapted from Brown (2015). The purpose of this assessment was to establish a baseline of the group’s conceptual domain mastery.

As shown in Figure 1, the group’s performance was markedly heterogeneous, revealing a functional yet fragile command of the discipline. The average success rate across the 21 items was moderate, and item-level analysis identified specific conceptual strengths and weaknesses. The pre-service teachers demonstrated high performance (64.3% success) on only one item concerning the principle that internal energy remains unchanged in a complete cycle (Q3). They showed medium performance (40–60%) on foundational concepts such as adiabatic expansion (Q1), isothermal compression (Q2), and direct applications of the Second Law (Q20, Q21). However, significant difficulties were evident in items requiring deeper reasoning, such as comparing work and heat across different thermodynamic paths (Q5, with only 7.1% success), analyzing entropy in free expansion (Q7), and justifying the spontaneity of processes via the Second Law (Q15, with 21.4% success). This profile suggests that while participants possessed declarative knowledge of basic principles, their ability to apply them in more complex or comparative scenarios was limited.

Figure 1
Understanding of thermodynamics concept among pre-service physics teachers.

5.2. Generation of analogical reasoning inproblem-solving

The following presents results from the main assessment, where pre-service teachers’ capacity to mobilize their knowledge and generate analogical reasoning was evaluated through two complex, open-ended problems: Maxwell’s demon paradox and the urban heat island phenomenon (See Appendices A and B, respectively, for complete problem descriptions.).

For the purpose of this analysis, a response was coded as “Yes” (Success) only if the participant demonstrated a correct and robust application of the concept or analogy to solve the target problem. Partial or superficial mentions were coded as “No”.

Table 1 reveals a performance pattern that is clearly dependent on both the modality and the complexity of the emergent construct.

Table 1
Distribution of key concepts and analogies in thermodynamics comprehension between AI and FF modalities.

Activation of Low-to-Intermediate Difficulty Concepts (C1–C3): The FF modality demonstrated a significant advantage in mobilizing foundational concepts. The difference was statistically significant for the correct application of isothermal compression (C2) (83.3% success in FF vs. 58.3% in AI; McNemar’s χ2=4.17, p=0.041, Cramér’s V=0.48) and cyclic processes (C3) (66.7% in FF vs. 33.3% in AI; McNemar’s χ2=5.14, p=0.023, Cramér’s V=0.52). These findings provide strong support for Hypothesis 1 (H1), suggesting that guided human interaction was more effective for enabling pre-service teachers to apply their basic conceptual knowledge correctly.

Generation of analogies for extreme-difficulty phenomena (A7–A8): The most telling finding emerged from the analysis of the most sophisticated analogies. In the FF modality, none of the 14 participants succeeded in constructing a successful analogy for either Maxwell’s Demon (A7) or the urban heat island (A8). In contrast, in the AI-assisted modality, a small but non-zero number of participants did succeed (16.7% for A7 and 8.3% for A8). This categorical difference – zero success versus non-zero success – provides compelling support for Hypothesis 3 (H3). It suggests that collaboration with AI enabled some participants to formulate analogies of a complexity that was otherwise beyond their unassisted pedagogical reasoning capabilities. These results can be seen in Figure 2.

Figure 2
Absolute distribution of responses for key concepts and analogies in thermodynamics comprehension between artificial intelligence (AI) and face to face (FF).

5.3. Concordance Analysis: Cognitive trajectories in problem-solving

A concordance analysis was conducted to determine whether the AI and FF modalities represent interchangeable or distinct cognitive pathways.

Table 2 shows that the two modalities are not functionally equivalent. For the activation of low-difficulty concepts (C1–C4), the agreement coefficients were very low (AC1 from 0.0 to 0.290), indicating they are divergent pedagogical trajectories. However, for high-complexity concepts (C7–C8), the concordance was almost perfect (AC1 > 0.89), but this consistency was driven by concurrent failure in both modalities, providing strong support for Hypothesis 2 (H2). For further details, see Appendix C.

Table 2
Concordance analysis of analogies and key concepts in the understanding of fundamental thermodynamics concepts between AI and FF.

5.4. Qualitative Synthesis: Mapping emergent reasoning structures

To visualize the knowledge structures that participants constructed, latent semantic networks were generated from the content of their solutions.

5.4.1. Latent semantic network: urban heat island phenomenon: Analysis of convergences and divergences from physics and thermodynamics perspectives

Figure 3 illustrates a The latent semantic network that synthesizes collective understanding of urban heat island formation through convergent thermodynamic mechanisms (material heat retention, vegetation deficit, and ventilation obstruction) counteracted by corresponding mitigation strategies, while divergent frameworks reveal advanced theoretical applications of entropy and Second Law inefficiencies, complemented by unique analogical reasoning demonstrating creative pedagogical approaches to complex thermal phenomena explanation.

Figure 3
Latent semantic network representing collective understanding of the urban heat island phenomenon derived from 12 participant responses. Node colors indicate conceptual categories: orange (central phenomenon), light blue (convergent understanding), green (divergent interpretations), red (causal factors), dark blue (mitigation solutions), dark orange (advanced thermodynamic concepts), and purple (unique analogies). Node size reflects hierarchical level (35-18 units). Edge styles denote relationship types: solid lines (direct relationships), dashed lines (analogical/theoretical connections); edge colors correspond to source node categories. Green edges specifically indicate mitigation relationships between solutions and causal factors. The network reveals a coherent causal structure (red nodes) addressed by complementary solutions (blue nodes), grounded in thermodynamic principles (orange nodes) and enriched by creative pedagogical analogies (purple nodes with participant codes P1-P12). ForceAtlas2 physics simulation ensures optimal spatial distribution for visual clarity.
5.4.2. Latent semantic network: Maxwell’s Demon: Analysis of convergences and divergences in reconciliation with the second law of thermodynamics

Figure 4 presents a Latent Semantic Network synthesizing the resolution of the Maxwell’s Demon paradox through three convergent thermodynamic (external work requirement, global entropy increase, and energy redistribution) illustrated via nineteen pedagogical analogies, while divergent frameworks establish the selective separation paradox and hypothetical violation scenarios, demonstrating how analogical reasoning bridges abstract thermodynamic formalism with accessible concrete metaphors in physics education.

Figure 4
Latent semantic network representing collective understanding of Maxwell’s Demon paradox derived from student and AI responses employing analogical reasoning (n=26 nodes, 41 edges). Node colors indicate conceptual categories: orange (central paradox), light blue (convergent reconciliation with Second Law), green (divergent paradox illustration), deep blue (external work requirement), green (global equilibrium and total entropy), turquoise (energy redistribution), red (selective particle separation), dark red (hypothetical differential ordering), purple (specific analogies), and dark orange (closing synthesis). Node size reflects hierarchical level (35-18 units). Edge styles distinguish structural relationships (solid gray), hierarchical inclusion (solid colored matching source category), analogical illustration (dashed colored), mechanistic progression (solid green), and paradox resolution (dashed orange). The network reveals three complementary resolution mechanisms – external work cost, total entropy increase, and energy redistribution – each supported by multiple pedagogical analogies, with a critical “resolved by” connection (dashed orange line) linking the selective separation paradox to the external work requirement. ForceAtlas2 physics simulation with increased gravitational constant ensures spatial clustering by conceptual function.

Although quantitative data provides a macro view of performance, the semantic networks presented here are essential to understand how students connected concepts. We utilized ForceAtlas2 physics simulation not merely for aesthetics, but to spatially cluster concepts that were cognitively linked by the students, revealing the density and centrality of specific ideas (e.g., ‘entropy’ or ‘disorder’) within their reasoning process.

6. Discussion

This study investigated the differential efficacy of two problem-solving modalities – one assisted by artificial intelligence (AI) and the other mediated through a face-to-face (FF) dialogue – on the development of analogical reasoning for learning thermodynamics among a group of pre-service physics teachers. By having the same 14 participants engage with complex problems under both conditions (a within-subjects design), our findings reveal a nuanced yet theoretically coherent pattern: (1) FF interaction was superior for activating and correctly applying foundational concepts; (2) both modalities proved equally ineffective for overcoming the cognitive barriers associated with high-complexity concepts, a pattern we term concordance in failure; and (3) the AI-assisted modality enabled a marginal yet categorically distinct success on analogical reasoning tasks of extreme complexity, where the FF modality resulted in complete failure. Herein, we interpret these findings, discuss their implications for physics teacher education, and propose a hybrid pedagogical model.

6.1. The superiority of human interaction for foundational concepts: adaptive scaffolding versus algorithmic assistance

The first key finding, which provides preliminary support for our initial hypothesis (H1), is the significant advantage of the FF modality in mobilizing low-to-intermediate difficulty concepts. The pre-service teachers demonstrated a notably greater ability to correctly apply concepts such as isothermal compression and cyclic processes when engaged in a Socratic dialogue with a human facilitator. This result aligns with research highlighting the superiority of human tutors over AI systems for the development of deep conceptual understanding [17].

The most plausible explanation for this advantage lies not in simple information transmission but in the nature of the adaptive pedagogical scaffolding that human interaction affords. A human facilitator can interpret subtle multimodal cues – a pause, an inflection of the voice, a facial expression – that signal doubt or an incipient misconception. This allows for real-time diagnostic intervention: the facilitator can rephrase a question, offer a counter-example, or guide the participant toward a more productive line of reasoning [31]. Current AI systems, though linguistically fluent, lack this contextual and pedagogical sensitivity. Their pattern-based assistance is less effective when a student’s cognitive obstacle does not conform to a predictable error. The FF interaction, therefore, appears superior for building and solidifying foundational knowledge because it is inherently diagnostic and adaptive.

6.2. Concordance in failure: the ceiling of pedagogical inefficacy in the face of complexity

Perhaps the most sobering finding is the concordance in failure of both modalities when confronted with high-complexity concepts such as entropy, irreversibility, and the Carnot cycle (H2). The fact that the vast majority of participants failed to successfully apply these concepts, regardless of the modality, reveals a problem deeper than the choice of a pedagogical tool. It demonstrates that neither a general Socratic dialogue nor assistance from a general-purpose AI is sufficient to overcome the well-documented and persistent cognitive barriers associated with the Second Law of Thermodynamics [12, 36].

This shared failure suggests that the problem lies not with the modality of interaction but with the absence of an explicit, research-based pedagogical content knowledge (PCK) intervention. Overcoming misconceptions about entropy as “disorder” or understanding the statistical nature of irreversibility requires a set of specific didactic strategies: the use of microscopic models, the explicit confrontation of intuitive ideas, and metacognitive reflection on probabilistic reasoning [37]. Neither the human facilitator (who, in our design, refrained from direct instruction) nor the AI (which lacks an embedded PCK model) provided this specialized scaffolding. This finding carries a critical implication for teacher education: mastery of disciplinary content does not guarantee the ability to teach it effectively. The preparation of future teachers must move beyond problem-solving and explicitly cultivate knowledge of research-based pedagogical strategies for the most challenging topics.

6.3. The incremental yet significant success of AI: Extending the boundaries of human reasoning

The most novel and provocative finding is the advantage, albeit marginal, of the AI modality on the analogical reasoning tasks of extreme complexity (H3). The fact that a few participants succeeded in generating effective analogies for Maxwell’s Demon and the urban heat island with AI assistance, while absolutely none did in the FF modality, represents a qualitative difference. This result suggests that AI can play a unique role not as a substitute for an instructor, but as a cognitive augmentation tool that allows learners to transcend their individual cognitive limitations.

The theoretical explanation for this phenomenon lies in AI’s computational capabilities. To resolve the Maxwell’s Demon paradox, for instance, one must bridge thermodynamics with information theory (Landauer’s principle) – a domain that may lie outside the immediate knowledge of a pre-service teacher. AI, with its instant access to vast interdisciplinary databases, can propose these non-obvious connections, acting as a catalyst for analogical creativity. It enables a process of iterative refinement and unconstrained exploration, where the participant can test and discard ideas at a pace and depth that would be difficult to sustain in a human dialogue limited by time and the facilitator’s own knowledge. Although the absolute success rate was low, this finding points to a strategic niche for AI in education: as a tool for complex problem-solving at the boundaries of knowledge, where its capacity to synthesize diverse information can unlock new pathways of thought [16].

6.4. Implications for teacher education: Toward a hybrid pedagogical model

Collectively, our findings argue against a dichotomous view that pits AI against traditional teaching. Instead, they advocate for a hybrid and differentiated pedagogical model that strategically leverages the strengths of each modality:

  1. For Foundational Concepts: Face-to-face, interactive, and feedback-rich instruction should remain the pillar for building a solid conceptual foundation. Its capacity for adaptive scaffolding is, for now, irreplaceable.

  2. For Complex Concepts: An explicitly hybrid approach is required. Students can use AI for initial exploration and problem-solving, but this must be supplemented with structured, instructorled FF sessions that implement research-based PCK strategies to directly address known cognitive barriers.

  3. For Frontier Reasoning: AI should be positioned as a cognitive augmentation tool for high-complexity tasks that demand the integration of interdisciplinary knowledge and the development of creative thinking, such as the generation of novel analogies.

Finally, an implication that transcends the choice of modality is the imperative to integrate explicit instruction in metacognitive strategies into teacher education programs. The participants’ difficulty in mobilizing their knowledge in both modalities suggests they lack not only content knowledge but also the metacognitive skills to plan, monitor, and evaluate their own problem-solving processes.

The education of future physics teachers must be built upon three integrated pillars: first, a deep command of disciplinary content that goes beyond procedural knowledge to encompass conceptual understanding; second, robust pedagogical content knowledge (PCK) that enables teachers to transform complex physics concepts into accessible and accurate representations for diverse learners; and third, a comprehensive repertoire of metacognitive strategies that empower both teachers and their future students to regulate their own learning processes effectively.

6.5. Limitations

The interpretation of these findings requires a prudent acknowledgement of the study’s exploratory nature and methodological boundaries. Specifically, the modest sample size (N = 14) and the fixed-order experimental design – while appropriate for a depth-oriented comparative case study – constrain the statistical power to detect subtle modality effects and introduce potential order-related confounds, such as learning carry-over or fatigue. Consequently, null findings should be viewed as indicative rather than definitive evidence of equivalence. Furthermore, the reliance on concurrent think-aloud protocols and the inherent variability of human facilitators suggests that the observed performance reflects reasoning under specific experimental demands rather than spontaneous pedagogical practice.

Beyond internal design, the study is bounded by its contextual and temporal specificity. Conducted within a distinct Colombian higher education setting and utilizing a specific iteration of AI technology (Claude, late 2024), the results may not be immediately generalizable to other cultural contexts or rapidly evolving AI architectures. Additionally, by operationalizing pedagogical competence primarily through analogical reasoning within a binary modality comparison, the research offers a focused but partial view of the multifaceted nature of physics teaching. Therefore, these constraints should be understood not as fundamental deficits, but as clear demarcations that define the scope of this initial inquiry, highlighting the necessity for future longitudinal, counterbalanced, and large-scale investigations to validate and extend these preliminary insights.

7. Conclusion

This exploratory comparative case study investigated the differential efficacy of AI-assisted versus face-to-face learning modalities in developing analogical reasoning for thermodynamics among pre-service physics teachers. Beyond a simple binary comparison, this research systematizes its findings into three distinct contributions – theoretical, practical, and methodological – that advance the scholarship on science teacher education.

The study’s primary theoretical contribution is the delineation of a framework of differential affordances. Our findings indicate that the choice between human and AI instruction is not a dichotomy but a matter of complementary strengths. Direct human interaction demonstrated a clear superiority in the adaptive scaffolding of foundational concepts, where the facilitator’s ability to interpret multimodal cues proved irreplaceable for solidifying basic understanding. Conversely, the AI modality offered tentative but distinct potential for supporting reasoning at the frontiers of complexity, enabling marginal success in interdisciplinary tasks where unassisted human reasoning stalled.

However, this differentiation has a critical limit, leading to our second theoretical construct: “concordance in failure.” We introduce this term to describe the phenomenon where both modalities reached a shared ceiling of inefficacy when confronted with high-complexity concepts like entropy and irreversibility. This construct serves as an analytical lens to distinguish between the inherent cognitive difficulty of a subject and the limitations of a delivery mechanism. It suggests that debates about “AI versus human instruction” are misframed when the fundamental challenge lies not in the modality, but in the absence of explicit, research-based Pedagogical Content Knowledge (PCK) to dismantle deeply rooted conceptual barriers.

In practice, these theoretical insights argue decisively against one-size-fits-all approaches. We propose an evidence-based hybrid pedagogical model for teacher education that leverages face-to-face interaction for foundational mentorship and metacognitive regulation, while strategically deploying AI tools for adaptive practice and complex synthesis. Furthermore, the observed gap between participants’ ability to generate analogies and their struggle to evaluate them underscores an urgent practical need: teacher preparation programs must move beyond content mastery to include explicit instruction in PCK and metacognitive monitoring, specifically for counter-intuitive topics like thermodynamics.

Methodologically, this work demonstrates the utility of Latent Semantic Network Analysis for visualizing complex qualitative data in small-sample contexts. This approach offered a replicable tool for identifying reasoning patterns – such as the semantic disconnectedness in AI-assisted explanations – that traditional coding might overlook, providing a robust avenue for future research in physics education.

Ultimately, this study posits that the future of physics education lies not in a forced choice between human and artificial intelligence, but in orchestrating a symbiotic relationship. As institutions navigate the permanent integration of AI, our findings offer empirical guidance: prioritize human mentorship for deep conceptual grounding and metacognitive development, while harnessing AI’s unique capacity to extend the boundaries of reasoning. The “concordance in failure” on entropy serves as a final, critical reminder that technology amplifies, but cannot replace, the need for sophisticated, specialized physics pedagogy.

Supplementary Material

The following online material is available for this article

Appendice A:

Appendice B:

Appendix C:

Acknowledgments

This research was supported by the Faculty of Education, University of Antioquia. We thank all pre-service physics teachers who participated in this study.

Data Availability

Anonymized data supporting the findings of this study are available from the corresponding author upon reasonable request and subject to institutional review board approval for secondary data use.

References

  • [1] NATIONAL RESEARCH COUNCIL, A framework for K-12 science education: Practices, crosscutting concepts, and core ideas (The National Academies Press, Washington, 2012).
  • [2] OECD, PISA 2018 assessment and analytical framework (OECD Publishing, Paris, 2019).
  • [3] W.K. Adams and C.E. Wieman, Am. J. Phys. 83, 459 (2015).
  • [4] J.L. Docktor and J.P. Mestre, Phys. Rev. ST Phys. Educ. Res. 10, 020119 (2014).
  • [5] M.V.J. Veenman and B.H.A.M. Van Hout-Wolters, Metacognition and Learning 1, 3(2006).
  • [6] M.C. Wang, G.D. Haertel and H.J. Walberg, Rev. Educ. Res. 63, 249 (1993).
  • [7] M.T.H. Chi, P.J. Feltovich and R. Glaser, Cogn. Sci. 5, 121 (1981).
  • [8] C. Dignath and G. Büttner, Metacogn. Learn. 3, 231 (2008).
  • [9] L. Darling-Hammond, Educ. Res. 45, 83 (2016).
  • [10] NATIONAL RESEARCH COUNCIL, Preparing teachers: Building evidence for sound policy (The National Academies Press, Washington, 2013).
  • [11] M.T.H. Chi, R.D. Roscoe, J.D. Slotta, M. Roy and C.C. Chase, Cogn. Sci. 36, 1 (2012).
  • [12] M.E. Loverude, C.H. Kautz and P.R.L. Heron, Am. J. Phys. 70, 137 (2002).
  • [13] B.J. Zimmerman and T.J. Cleary, in: Handbook of motivation at school, edited by K.R. Wentzel and A. Wigfield (Routledge, New York, 2009).
  • [14] E. Etkina, B. Gregorcic and S. Vokos, Phys. Rev. Phys. Educ. Res. 14, 010120 (2018).
  • [15] L. Chen, P. Chen and Z. Lin, IEEE Access 8, 75264 (2020).
  • [16] E. Kasneci, K. Sessler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. Günnemann, E. Hüllermeier et al., Learn. Individ. Differ. 103, 102274 (2023).
  • [17] W. Holmes, M. Bialik and C. Fadel, Artificial intelligence in education: Promises and implications for teaching and learning (Center for Curriculum Redesign, Boston, 2019).
  • [18] O. Zawacki-Richter, V.I. Marín, M. Bond and F. Gouverneur, Int. J. Educ. Technol. High Educ. 16, 39 (2019).
  • [19] J.H. Flavell, in: Metacognition, motivation, and understanding, edited by F.E. Weinert and R.H. Kluwe (Lawrence Erlbaum Associates, Hillsdale, 1987).
  • [20] G. Schraw and R.S. Dennison, Contemp. Educ. Psychol. 19, 460 (1994).
  • [21] M.V.J. Veenman and D. van Cleef, ZDM Math. Educ. 51, 691 (2019).
  • [22] H. de Boer, A.S. Donker, D.D. Kostons and G.P. van der Werf, Educ. Res. Rev. 24, 98 (2018).
  • [23] B.J. Zimmerman, in: Handbook of self-regulation, edited by M. Boekaerts, P.R. Pintrich and M. Zeidner (Academic Press, San Diego, 2000).
  • [24] E. Panadero, Front. Psychol. 8, 422 (2017).
  • [25] A. Bandura, Self-efficacy: The exercise of control (W.H. Freeman, New York, 1997).
  • [26] B.J. Zimmerman and A.R. Moylan, in: Handbook of metacognition in education, edited by D.J. Hacker, J. Dunlosky and A.C. Graesser (Routledge, New York, 2009).
  • [27] S. Yerdelen-Damar and H. Peşman, J. Educ. Res. 106, 280 (2013).
  • [28] M.L. Bernacki, J.P. Byrnes and J.G. Cromley, Contemp. Educ. Psychol. 39, 321 (2014).
  • [29] D. Gentner, Cogn. Sci. 7, 155 (1983).
  • [30] R. Duit, Sci. Educ. 75, 649 (1991).
  • [31] L.R. Novick, J. Exp. Psychol. Learn. Mem. Cogn. 14, 510 (1988).
  • [32] L.E. Richland, O. Zur and K.J. Holyoak, Science 316, 1128 (2007).
  • [33] R. Duit and D.F. Treagust, Int. J. Sci. Educ. 25, 671 (2003).
  • [34] D.F. Treagust, R. Duit, P. Joslin and I. Lindauer, Int. J. Sci. Educ. 14, 413 (1992).
  • [35] A.G. Harrison and D.F. Treagust, in: Metaphor and analogy in science education, edited by P.J. Aubusson, A.G. Harrison and S.M. Ritchie (Springer, Dordrecht, 2006).
  • [36] W.M. Christensen, D.E. Meltzer and C.A. Ogilvie, Am. J. Phys. 77, 907 (2009).
  • [37] D.F. Styer, Am. J. Phys. 68, 1090 (2000).
  • [38] H.S. Lee, O.L. Liu and M.C. Linn, J. Res. Sci. Teach. 59, 243 (2022).
  • [39] K.A. Ericsson and H.A. Simon, Protocol analysis: Verbal reports as data (MIT Press, Cambridge, 1993).
  • [40] M.T.H. Chi, J. Learn. Sci. 6, 271 (1997).
  • [41] K.L. Gwet, Handbook of inter-rater reliability: The definitive guide to measuring the extent of agreement among raters (Advanced Analytics, LLC, Gaithersburg, 2014), 4 ed.

Edited by

Publication Dates

  • Publication in this collection
    29 May 2026
  • Date of issue
    2026

History

  • Received
    15 Sept 2025
  • Reviewed
    11 Feb 2026
  • Accepted
    22 Mar 2026
location_on
Sociedade Brasileira de Física - SBF Rua do Matão, travessa R, 187 - Edifício Sede - Cidade Universitária, São Paulo, SP, Brasil, CEP 05508-090, Tel: +55 (11) 3034-0429 - São Paulo - SP - Brazil
E-mail: rbef@sbfisica.org.br, marcellof@unb.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro