Open-access MTNN based meta-modeling and PSO driven multi-response optimization of FSW parameters for enhancing mechanical attributes of AZ80A Mg alloy joints

ABSTRACT

In this work, an interpretable machine learning framework was developed for optimizing friction stir welding (FSW) parameters during joining of AZ80A Mg alloy plates for lightweight structural applications. The proposed strategy integrates a multi-task neural network (MTNN) with particle swarm optimization (PSO) to enhance tensile strength (TS) and hardness (H). Welding speed (WS), rotational speed (RS), axial force (AF), tool tilt angle (TA), and pin geometry (PG) were considered as input variables. A face-centered central composite design was adopted for generating training data. MTNN architecture effectively captured nonlinear interactions among process variables and mechanical responses through shared feature learning and task-specific prediction branches, achieving prediction accuracies (R²) of 95.13% and 93.35% during training, and 86.57% and 84.02% during testing for TS and H, respectively, with the lowest mean prediction deviations among all evaluated models. SHapley Additive exPlanations (SHAP) analysis revealed WS, AF, and tapered cylindrical (TC) pin geometry significantly influenced joint performance. Subsequently, PSO was coupled with the trained MTNN to identify optimal parameter combinations yielding superior TS and H, and experimentally validated. SEM fractography confirmed ductile fracture behavior and Compared with RR, SVR, ANN, and RFR models, MTNN–PSO framework demonstrated lower prediction error, improved robustness, and enhanced interpretability.

Keywords:
Multi task neural network; Particle swarm optimization; Friction stir welding; AZ80A Mg alloy; Tensile strength

1. INTRODUCTION

1.1. FSW

Invented by TWI (i.e., The Welding Institute), friction stir welding (i.e., FSW) is an entirely environmentally friendly, green joining process, in which joining of materials is possible, without melting the parent metals [1]. In this process, a non-consumable type tool possessing uniquely designed pin geometry was employed and heat gets generated when this tool rotates on the surface of the materials to be welded. This frictional heat leads to plastic deformation and these deformed materials gets stirred at the region of interface. As the rotating tool travels along the line of joint, it merges the softened materials, thereby leading to the generation of flaw free joints possessing superior mechanical attributes. Contradictory to the conventional joining methodologies, involving reasonable thermal cycles leading to cracks, coarsening of grains, porosity etc., which deteriorates the quality of the joint, during the process of FSW, joining of materials takes place below their melting temperature [2, 3]. Due to this, not only the metallurgy defects get minimized, occurrence of flaws like distortion, residual stresses also get reduced, making FSW process a preferable welding technique for joining hard to join and temperature sensitive materials including alloys of titanium, magnesium (Mg), stainless steel etc., [4].

In the FSW process, the parameters (like tool’s rotational speed, speed of traverse, its pin geometry, shoulder diameter, angle of tilt, offset distance, axial force etc.,) being employed play a vital role in determining the volume of heat, flow of the plasticized material, consolidation of the plasticized metal, conduction & distribution of frictional heat, flow behavior of plasticized metal etc., and thereby impact the micro-structural modifications, structural integrity, mechanical attributes of the fabricated joints [5]. As a result, these process parameters have to be employed in a suitable combination, such that the employed combination is ideal for attaining flaw free joints exhibiting superior mechanical properties. Identifying suitable, perfect combination of process parameters during FSW entails choosing and adjusting these parameters in a systematic manner, so as to attain the perfect joints and processing outcomes [6]. The major objective of optimizing the parameters of the FSW process is not only to ensure fabrication of superior quality, flaw free joints, but also to minimize relevant fabrication costs, speed up the production rate, reduce consumption of energy, increase reliability and consistency of the fabricated weldment. In the recent years, several researchers had focused their research work on optimizing the process parameters during FSW of various metals [7, 8].

Among the various commercially available magnesium alloys, AZ80A has been selected in the present investigation owing to its excellent combination of low density, high specific strength, good castability, and superior strength-to-weight ratio, making it an attractive material for lightweight structural components in the automotive, aerospace, and defense industries. The relatively higher aluminum content in AZ80A promotes precipitation strengthening through the β-Mg₁₇Al₁₂ phase, resulting in higher mechanical strength than many conventional magnesium alloys. However, its friction stir weldability is strongly influenced by complex thermo-mechanical interactions, dynamic recrystallization, and material flow behavior, making optimization of FSW process parameters essential for achieving defect-free joints with enhanced mechanical performance.

1.2. Need for optimization

Conventional optimization techniques like RSM (i.e., response surface methodology), Taguchi method, Analysis of variance, grey relational analysis, design of experiments and unconventional techniques like genetic algorithm (GA), ant colony optimization, simulated annealing (SA), multi criteria decision making approaches etc., had been frequently employed by several experimental investigators to optimize the parameters of the FSW process [9, 10]. In the recent years, employing traditional optimization techniques was proven to be not effective for FSW process due to the fact that majority of the traditional techniques were focused towards single objective optimization. On the contrary, FSW process comprises several contradicting objectives (like minimizing residual stresses and maximizing tensile strength) [11].

Most traditional techniques focus on single-objective optimization, whereas FSW often involves conflicting objectives (e.g., maximizing tensile strength while minimizing residual stresses). These techniques were reported to identify and analyze the intricate, nonlinear interaction amidst the parameters, demand for reasonable experimental investigation, making them practically not preferable for industrial application-oriented process like FSW, which demands for rapid optimization [12]. At the same time, optimization techniques like SA, GA were reported to function frequently as black box patterns, providing minimal insight into the elementary relationships amidst the outcomes and process parameters, thereby affecting the decision-making capability [13, 14]. Moreover, with the escalation in the parameter’s levels, sample numbers, the relationship amidst the process parameters and experimental outcomes becomes more intricate, frequently found to exhibit larger levels of non-linearity and uncertainty. RSM, Analysis of variance, multi criteria decision making approaches basically rely on uniquely formulated numerical models and as a result, when employed for optimizing scenarios inheriting larger levels of non-linearity and uncertainty, the models formulated by these techniques were reported to oversimplify the prevailing scenario and were found to be inefficient in providing adequate flexibility to adjust to the system’s complexity [15].

1.3. ML & DL

ML (i.e., machine learning) is a subset of AI (i.e., artificial intelligence) that permits systems to learn and improve automatically from past experience, eliminating the need for specific programming. It was made possible with the help of ML to forecast and complete tasks involving decision making, by acquiring data for hidden trends and patterns [16]. ML encompasses formulation of algorithms and numerical models that assess and demonstrate data for making suitable decisions. Another unique feature of ML is that they do not rely on predefined principles, instead ML possess the capability of automatically adapting and improving their performance, when exposed to numerous data. ML have also exhibited breakthroughs when employed for high dimensional and large scale data sets. Apart from its merits, ML techniques also possess some drawbacks which limit their applications [17]. For instance, ML based techniques were found to exhibit inferior performance in formulating models for scenarios involving large dimensional process spaces. It was also found that models framed by ML fail to share their parameter networks betwixt multiple forecasting tasks, thereby failing to manipulate feedback from distinctive tasks required for accelerating learning efficacy. It is also proven to be difficult to interpret the ML formulated models, thereby making it complicated to understand the fundamental impact of parameters on experimental outcomes [18].

Deep Learning (DL) is a specialized branch of Machine Learning (ML) that employs multi-layered artificial neural networks to automatically learn complex patterns and hierarchical feature representations directly from data [19]. Unlike conventional ML algorithms, which often require manual feature engineering, DL progressively extracts low-level and high-level features through multiple hidden layers, enabling accurate modeling of highly nonlinear relationships. Inspired by the structure and functioning of the human nervous system, deep neural networks consist of interconnected computational neurons that collaboratively learn intricate input–output mappings through iterative training. Owing to their exceptional capability for handling complex, high-dimensional datasets, DL models have been increasingly adopted in manufacturing, materials engineering, and welding applications for prediction, optimization, defect detection, and intelligent decision-making. It is inspired by the structure and functioning of the human brain, where multiple layers of artificial neurons are used to model and learn from data. The “deep” in deep learning refers to the use of multiple layers in the neural network, enabling it to learn hierarchical representations of the data [20, 21].

1.4. Neural networks, RR & RFR

Neural networks were found to have made reasonable breakthroughs in heath and prognostic management, computer vision, processing natural languages etc. They are perfectly suited for handling large scale information, analyzing complicated non-linear relationships in several manufacturing including forming, welding processes, inheriting intricate thermal behaviors [22]. Artificial Neural Networks (ANNs) are machine learning models inspired by the structure and functioning of the human brain, consisting of interconnected neurons organized into multiple computational layers. Owing to their ability to capture complex nonlinear relationships, ANNs have been extensively employed for modeling manufacturing processes. However, conventional ANN-based optimization studies generally adopt single-task architectures, requiring separate models for individual responses, which increases computational cost and overlooks shared information among correlated outputs [23, 24]. Moreover, their limited interpretability and reduced generalization capability often restrict predictive performance. These limitations motivate the adoption of advanced deep learning frameworks, such as Multi-Task Neural Networks (MTNN), for simultaneous multi-response prediction and optimization.

RR (i.e., ridge regression) is a category of linear regression that involves a regularization term for avoiding overfitting by chastising higher coefficients in the model. This is most suitable for scenarios in which the raw data inherits larger correlation values amidst forecasting variables (i.e., multi–collinearity) [25]. RFR (i.e., random forest regression) is a unique learning process that employs numerous decision trees for forecasting the responses. RFR also incorporates the prognosis of singular trees, for enhancing accuracy and minimizing overfitting, thereby making it an efficient tool for both classification and regression tasks [26]. It was proven by researchers that both RFR and RR are very much efficient in formulating regression models w.r.t. parameters of various manufacturing processes. By combining together, ANN, RFR and RR, their merits gets leveraged in tandem, thereby making it possible for formulating superior performance, intelligible, robust models for various forecasting tasks like optimization of parameters related to manufacturing sectors including FSW process. These techniques in an interconnected manner can address various diversified challenges from acquiring intricate non–linear relationships to dealing multi–collinearity. At the same time, there also prevails a gap in exploring the application of deep learning based neutral networks for establishing relationship amidst the performance indicators and parameters of the FSW process.

1.5. MTNN & SHAP

During the FSW process, there prevails a relative large non–linear category relationship amidst the parameters being employed and the experimental outcomes and this relationship cannot be expressed and described easily by means of numerical models. In this paper, with the objective of addressing all the above mentioned complex issues related with optimization of FSW process, an intelligible MTNN (i.e., multi–task neural network) was employed. An intelligible MTNN, in the background of optimizing the FSW parameters pertains to a neural network model formulated for simultaneously forecasting several responses or outcomes (i.e., multi-task learning), while providing perceptions on how the employed parameters impact these responses [27].

Generally, FSW comprises of numerous interconnected responses (like hardness, strength, toughness, flaw rate, corrosion resistance etc.,) that demands for optimizing the parameters simultaneously. The model formulated using MTNN can address the relationships between responses and outcomes efficiently. In addition to this, as MTNNs are intelligible, the formulated model will deliver valuable insights like which parameter (like tool offset) strongly influenced a particular outcome (like hardness) and how far the outcomes are sensitive to variation in every parameter. At the same time, intelligible MTNNs also inherit some limitations. For instance, even though MTNN was basically designed for optimizing the multiple responses in a simultaneous manner, there exists difficulty in specifically explaining the trade–offs amidst responses (like hardness vs. tensile strength) [28]. Moreover, MTNNs can only provide generalized perceptiveness on how the parameters impact the outcomes and fail to provide accurate explanations w.r.t explicit instances like how a specific parameter setting contributed for escalation in hardness.

SHAP (i.e., Shapley additive explanation) is an efficient interpretability ML technique employed for describing the outcomes of the formulated forecasting models [29]. SHAP when employed for optimizing the FSW parameters, identifies the impact of every employed parameter (like speed of traverse, rotational speed, axial force, tool pin geometry etc.,) on the experimental outcomes (like hardness, tensile strength etc.,). By visualizing and quantifying these contributions, SHAP furnishes perceptiveness on how the employed parameters had influenced the joint quality, thereby permitting informative and data driven optimization. Another attractive feature of SHAP is that it identifies shared significant parameters and parameter based specific trade–offs, thereby facilitating researchers to design parameter combinations that will balance the contradicting objectives [30]. SHAP provides localized explanations for specific trial joints, thereby revealing the accurate contribution of every parameter of FSW process w.r.t the forecasted response. By merging the global insights attained from MTNN with these SHAP localized explanations, investigators will gain a detailed understanding of both specific cases as well as generalized trends.

Compared with conventional ANN, RFR, and SVR models, the proposed MTNN framework offers several scientific advantages for FSW parameter optimization. Unlike ANN, which generally develops independent models for each response, the MTNN simultaneously predicts tensile strength and hardness through a shared feature-learning architecture with task-specific output branches. This enables the model to exploit the inherent correlation between multiple mechanical responses, resulting in improved learning efficiency, enhanced generalization capability, and reduced computational redundancy. In contrast, RFR and SVR primarily learn response-specific mappings without utilizing shared information among correlated outputs, thereby limiting their ability to capture complex multi-response interactions. Furthermore, the integration of SHAP with MTNN enhances model interpretability by quantifying the contribution of each process parameter, while coupling the trained MTNN with PSO facilitates efficient multi-response optimization. These features collectively represent a significant advancement over conventional ML-based FSW optimization approaches.

1.6. Proposed optimization methodology

The optimization methodology being adapted for this research work for FSW of AZ80A Mg alloys is described below:

  • Trial FSW experimental runs for joining AZ80A Mg alloy plates were conducted based on the face centered central composite design incorporating 3 levels and 5 factors (including 4 numerical factors and 1 categorical factor). In the data generation stage, CLS (i.e., curriculum based learning strategy), which is a ML concept, was employed for reorganizing and structuring the set of raw data. This plays a vital role in diminishing the complexity in learning, overfitting risk and enhances the learning efficacy.

  • An optimization strategy merging PSO (i.e., particle swarm optimization) and MTNN was employed. Numerical model was formulated based on MTNN which comprised of 2 branches for forecasting the non-linear relationship amidst process parameters and experimental outcomes namely, hardness and tensile strength. These 2 branches shared weights for speeding up the training process and for enhancing the learning about associated attributes. During the optimization phase, constraints and variables were taken into consideration for attaining the preferred objective function, furnishing guidance for design optimization.

  • SHAP was used for computing the significance of distinctive parameters of the FSW process on the experimental outputs, which was essential for gaining detailed understanding of the influence of every employed parameter on outcomes and guiding rational adjustment & optimization of the parameters. Moreover, numerical model formulated based on intelligible MTNN was merged with this SHAP, for attaining evidential foundations for exceptional decision making.

  • In this research work, the proposed strategy was used for comparing the results generated by distinctive approaches. In addition to this, the results attained from various employed optimization comparison techniques were also used for validating the forecasted results in a practical manner.

2. OPTIMIZATION BACKGROUND

2.1. PSO algorithm

PSO algorithm falls under the category of heuristic optimization technique, which relies on procedures derived from the combined behavior of bird flocks and fish schools. In a PSO, potential solutions termed as particles, make their movement in the space on their past individual experience (individual ideal position) and on the entire swarm’s experience (globalized ideal position). Every particle alters its position and velocity repeatedly, equalizing exploration of the potential solutions by exploring known best solutions. This collective behavior permits the swarm to assemble towards the optimized solution efficiently [31, 32]. By employing PSO, the perfect combination of the parameters during FSW of AZ80A Mg alloy plates can be determined with the objective of attaining superior quality flaw free joints exhibiting enhanced mechanical attributes. Equations 1 to 4 describe the complete working behavior of PSO in optimizing the parameters of the FSW process for joining AZ80A Mg alloys:

(1) X i = ( x 1 i , x 2 i , , x i D )
(2) P i = ( p 1 i , p 2 i , , p i D )
(3) V i = ( v 1 i , v 2 i , , v i D )
(4) v i d = w v i d + c 1 r 1 ( P i d X i d ) + c 2 r 2 ( P g d X i d )

where viD is the particle’s velocity, xiD is the particle’s current position, w is the inertia factor, relative impact of the intelligible component is determined by c1, relative impact of the social component is determined by c2, Pid is the pbest of the ith particle, Pad is the gbest of the group, w is the weight of inertia and r1, r2 are incidental numbers that are spread equally amidst 0 and 1.

Relationship amidst parameters of the FSW process and the performance outputs is intricate and superiorly non–linear, especially when employed for joining alloys of Mg, like AZ80A. PSO algorithm, being self–dependent of the scenario’s explicit nature, is especially well suitable for addressing multi–objective and non–linear optimization problems associated with the FSW process [33]. In this research work, the parameters taken into consideration includes tool’s traverse speed, rotational speed, tilt angle, pin geometry and axial force. Differences amidst anticipated and objective values of the joint quality including hardness, tensile strength and minimization of flaws, deduced from a stand–in model was employed as the observational function for evaluating the performance of each & every combination of parameters. Every particle in PSO describe a 5-dimensional vector conforming to the above mentioned 5 parameters [34]. First, a set of particles were generated randomly, each portraying a probable combination of FSW parameters, with initial random velocities and positions assigned to every particle. These maiden values and their modification directions describe the PSO earliest guesses for the optimized parameters.

During every iteration, the value of the intention function was computed based on the prevailing position of every particle, which quantifies the distinctness amidst the desired and forecasted joint quality outcomes. The position and velocity of every particle was updated depending on the perfect position attained by every particle (personally perfect) and swarm’s global ideal position. The particles assemble towards the optimized solution by means of repetitive updates. The process gets terminated automatically once the predestined repetition count is attained and the optimized combination of FSW parameters will be generated. PSO was chosen over other metaheuristic techniques such as GA, GWO and DE for this optimization problem due to its comparatively simpler parameter tuning (requiring only inertia weight and two acceleration coefficients), its natural suitability for continuous search spaces without requiring discrete crossover or mutation operators, and its faster convergence behavior for low-dimensional problems such as the present 5-variable FSW optimization task [33, 34].

2.2. SHAP optimization methodology

SHAP is an effective tool for interpreting forecasting made by ML based models and it figures out to what extent every parameter contributes for the model’s forecasting. SHAP relies on two major principles consistency and additivity [35]. Contributions made by every individual parameter are being added up to the final forecasting was assured by additivity and consistency assures that importance of a parameter escalates with increase in its contribution, thereby making SHAP one of the most unbiased and fair technique for analyzing the importance of a parameter. Moreover, SHAP simplifies the forecasting of an intricate ML model into intelligible parts, ascribing an accurate value to the impact made by every parameter. Computing the values of SHAP for a definite parameter of FSW process entails totaling the insignificant contributions of the parameter I, calculated across all the subsets S of N, excepting parameter I, was given by:

(5) φ i = S N { i } | S | ! ( n 1 | S | ) ! [ f ( S { i } ) f ( S ) ] n !

where f is the forecasting model, parameters subset is indicated by S, N portrays the set of all the FSW parameters, specific parameter is denoted by i and n is the total number of features.

Equation (6) explains the SHAP technique’s concept for giving significance to a specific parameter i based on the forecasting result of a single sample:

(6) f ( x ) g ( x ) = φ 0 + i = 1 M φ i x i

where x denotes the reference under explanation, x′ indicate the input in a simplified manner and the relationship amidst these two was governed by the mapping function x = hx(x′). In addition to this, ϕ0 denotes the foundation value when all the inputs were absent and the number of simplified input parameters was indicated by M. Offering global interpretability is an attractive feature of SHAP and this permits the researchers to comprehend about every parameter’s contribution, thereby finding the collective dependency and significance of distinctive parameters w.r.t experimental outcomes [36]. By employing SHAP to FSW of AZ80A Mg alloys, models can be formulated for forecasting key responses namely hardness, tensile strength and it will be possible to identify the extent to which the FSW parameters impact the anticipated outputs. This helps us to completely understanding the formulated model’s reliability on various parameters of FSW process. By comprehending the interdependencies of these FSW parameters, researchers can generate valuable decisions for attaining preferred joint quality and efficacy.

3. EXPERIMENTAL SETUP

3.1. Base metal and FSW machine

Rectangular plates of AZ80A Mg alloy having a dimension of 110 mm (length) X 50 mm (width) X 6mm (thickness) were taken as the base metal in this research work. Several chemical ingredients and mechanical attributes of this base metal are described in the Table 1. The FSW machine being used in this research work is illustrated in the Figure 1(a), which was a semi-automatic category machine, equipped with a 5-kW capacity motor spindle. It comprised of a 410 X 820 mm side work platform and the entire platform of this FSW machine had the capability to traverse in multiple directions longitudinally (510 mm), vertically (400 mm), horizontally (400 mm).

Table 1
Parent metal – chemical constituents and mechanical attributes.
Figure 1
(a) FSW machine (b) fixture used for holding AZ80A Mg alloy plates and (c) friction stir welded AZ80A Mg alloy plates.

Rectangular plates of AZ80A Mg alloy were held rigidly together with the help of a uniquely designed and fabricated work fixture, inheriting clamps on the four sides, as seen in the Figure 1(b). A tool fabricated out of M35 grade high speed steel possessing unique pin geometries was employed in this research work. The principle of operation of the FSW process, where the AZ80A Mg alloy plates were welded as butt joints, at 90° angle to their direction of rolling. During this process, the joining of the AZ80A Mg alloy plates had occurred even before the plates had reached their melting point and the heat required for joining the plates was generated owing to the friction amidst the surface of the plates and the shoulder of the rotating tool. Photograph of the friction stir welded AZ80A Mg alloy plates is portrayed in the Figure 1(c).

3.2. Investigational design

Based on the expertise achieved from the vast literature review’s experimental inferences [8, 13, 15] and results of the preliminary trial runs, significant FSW process parameters of this experimental investigation were identified carefully. Parameters being taken into consideration for this experimental investigation include tool’s speed of traverse, i.e., welding speed (WS), rotational speed (RS), tilt angle (TA), pin geometry (PG) and axial force (AF). Reasons for choosing these parameters includes their inevitable role in ascertaining the volume of friction heat being generated, integrity of the joints, flow of the plasticized metal. With the objective of ensuring a detailed exploration of the FSW parameters and their impacts during FSW of the AZ80A Mg alloy, suitable ranges for every parameter were fixed. These ranges were chosen based on the guidance of prevailing literature, reality in process constraints and initial runs carried out for ensuring relevance and feasibility [5, 7, 10]. Table 2 describes the selected ranges of the parameters and these chosen ranges permit for acquiring the intricate interdependencies amidst the parameters while guaranteeing soundness of the fabricated joints.

Table 2
Parameters of FSW process and their distinctive levels.

Based on the initial trial runs, the limits of every parameter were set as follows: 3kN < AF < 5kN, 1.5 < WS < 2.5, 900 < RS < 1100, 0 < TA < 1 and tool with three distinctive pin geometries (PG) namely straight cylindrical (SC), taper cylindrical (TC) and straight square (SS). For avoiding the impact of the physical investigational errors on the response function, it is significant to attain the attributes of the experimental data within the range of parameters with minimal realistic experiments, thereby generating a most efficient objective function. Face centered central composite investigational design was employed for generating investigational points, which will give design space uniformity and fills the space of the sampling points. As described in Table 2, proposed investigational design with 5 factors was implemented based on the attributes of the FSW process parameters.

4. PROPOSED STRATEGY

In this research work, an optimization strategy combining PSO and MTNN was proposed. MTNN was employed for establishing a meta model incorporating chosen parameters for the outcomes namely tensile strength (TS) and hardness (H). This meta model was constructed with 5 parameters of the FSW process as inputs and tensile strength, hardness as outcomes. MTNN comprises of two branches and a backbone. The backbone architecture was employed for sharing the chosen parameter’s network and the two branches were utilized for forecasting the hardness and tensile strength. The objective function was designed by means of meta model, based on which the globalized optimization search was carried out using PSO. The employed optimization strategy is portrayed in the Figure 2, and the step-by-step procedure is mentioned below:

Figure 2
Optimization strategy employed in this research work.
  • 1st step: Dataset for the FSW process comprising of process parameters & outputs (TS & H) was defined. Five process parameters include AF, RS, WS, PG and TA, followed by the outputs TS and H.

  • 2nd step: FSW experiments were carried out to attain the datasets of the observation points attained using the RSM (i.e., response surface methodology) face–centered central composite design, in which w.r.t employed 3 pin geometry, 30 trials per individual pin geometry were carried out.

  • 3rd step: Hierarchical strategy of learning was employed for reorganizing the acquired dataset.

  • 4th step: MTNN was trained by means of the well-structured datasets and its accuracy was tested by means of randomly chosen additional experiments.

  • 5th step: Perfectly trained MTNN model was employed for constructing the objective function of the optimization model.

  • 6th step: PSO algorithm was employed for identifying the globalized optimization parameters and the process was designed to stop, once the highest number of iterations was reached.

  • 7th step: Effectiveness of the generated optimization parameter values were validated by conducting realistic experiments.

  • 8th step: Contribution of every parameter w.r.t the formulated model was explained by employing SHAP methodology.

4.1. Preparation of data

FSW experiment was employed for collecting the sets of data and was described as:

(7) D = { x A F i , x W S i , x R S i , x P G i , x T A i , y T S i , y H i } i = 1 n

where n is the total number of experimental runs, xAFi is the axial force, xwSi is the welding speed, xRSi is the rotational speed, xPGi is the pin geometry, xTAi is the tilt angle, yTSi is the tensile strength of the fabricated joint and yHi is the hardness of the joint. Statistically designed experiments were carried out to acquire 90 experimental data sets and the outcomes of these trial runs for tensile strength and hardness are illustrated in the Figure 3(a) & (b) respectively.

Figure 3
Circular dendrogram graphs of the 90 trial runs illustrating the impact of parameters on (a) TS and (b) H.

In Figure 3(a), the outer nodes are provided with distinct colors for describing the various values of TS ranging from 154 to 242 MPa. From this circular dendrogram, it can be visualized that the tapered cylindrical (TC) pin geometry had contributed for larger values of TS consistently and the lowest value of tensile strength was generated during the employment of tool possessing straight cylindrical pin geometry. Similarly, from Figure 3(b), it can be inferred that both SC and TC had contributed for higher values of hardness. These graphs reveal the significant role of tool geometry and interacting dynamics of the process parameters in determining the quality of the friction stir welded joints and provides significant interpretations for optimizing the parameters during FSW of AZ80A Mg alloys [11].

Figure 4(a) & (b) graphically illustrates the experimental data distributed before progressive training framework. From these figures, it can be inferred that the source data was intricate and shuffled, making it challenging for neural networks for learning efficient patterns from it. In case of minor samples, input inconsistencies will lead to fluctuating model performance and lack of robustness.

Figure 4
Data repository for TS and H (a) & (b) before and (c) & (d) after the progressive training framework.

As no sequential correlation amidst samples before and after the statistically designed dataset, the experimental order does not impact the consequent orthogonal sets results and the data set can be identified by means of a legitimate approach. Hierarchical training arranges the trained data to attain quicker convergence and improved performance [18, 23].

By restructuring the actual dataset employing the hierarchical strategy of optimization, this strategy does not impact the overall attributes of the source data and at the same time, reduces the complexity of the data set. Major objective of the hierarchical strategy of training is to position complicated samples at the last and elementary samples at the starting, thereby permitting the model to adapt slowly to the data complexity. This, in turn improves the performance and efficiency of the training [16]. As a result, a structured data processing method relying on hierarchical learning methodology was proposed. Initially, the aggregated values of ysum of TS and H were computed based on the Equation No. 8. Then, the data is arranged in a high to low sequence based on ysum.

(8) y s u m = λ T S y T S + λ H y H

where yrs and yH are the tensile strength and hardness values respectively, λH and λH are the weight of TS and H respectively, normally set at a value of 0.5, ysum is the cumulative score.

Figure 4(c) & (d) portrays sequenced data set which follows a distinct pattern. Data preprocessing is essential for improving network stability and ensuring convergence during model training. In the present study, the input variables were normalized using the Min–Max normalization technique, which linearly scales each feature to the range of [0, 1], thereby eliminating the influence of differing feature magnitudes and ensuring that all input variables contribute uniformly during model learning. The normalization procedure adopted is expressed by Equation No. 9

(9) X n o r m a l i z e d = X min ( X ) max ( X ) min ( X )

As the volume of the data was minimal, verification subsets were not partitioned additionally. A 72:18 (80:20) split was adopted for the learning and evaluation subsets, consistent with common practice for small-sample regression tasks, ensuring a sufficiently large training set for the MTNN to learn shared representations while retaining an adequately sized held-out set for unbiased performance evaluation. A batch size of 4 was used during training, consistent with the limited training set size (n = 72); this allowed a sufficient number of gradient updates per epoch (18 updates/epoch) while introducing mild stochasticity beneficial for regularization on small datasets. Given the limited dataset size (n = 90), a separate validation set was not partitioned; instead, dropout regularization (10%) and the shared-backbone architecture of MTNN were relied upon to mitigate overfitting during training.

4.2. MTNN training procedure

In contrast to the conventional neural networks that are capable of performing an individual definite task, MTNNs possess the capability of carrying out various tasks simultaneously. This was attained by integrating various outcome layers in the network, each addressing a specific task and distributing the undisclosed layers around all the tasks [17, 21]. By simultaneously optimizing all these specific tasks, MTNN enhances the effectiveness and performance around the tasks. MTNN plays a significant role in modeling the FSW parameters and simultaneously provides the intricate mapping amidst those parameters, TN and H of the joints.

In this paper, common internal layers of MTNN were employed for extracting unified feature patterns of the FSW parameters and these patterns possess the ability to grab the correlations prevailing amidst those distinctive FSW parameters, thereby enhances the model’s learning versatility. Then for forecasting the outcomes TS and H separately, independent forecasting nodes were employed and, in this work, parameters of common layers as well as customized task response layers were optimized for precisely forecasting TS and H. The objective function comprised of error metrics of all relevant tasks, harmonizing the significance of distinctive tasks by means of priority-based aggregation. Formula described in the Equation No. 10 was employed as a part of MTNN and the outcome of the common internal layer was:

(10) H = R e L U ( W s h . X + b s h )

where the outcome of the common layer was denoted by H, biases & weights of the common internal layers are denoted by Wsh and bsh respectively and function of activation represented by ReLU.

Likewise, the outcome of the customized task response layer was given by the Equation Nos. 11 and 12 and are as follows:

(11) y T S = S i g m o i d ( W T S . H + b T S )
(12) y H = S i g m o i d ( W H . H + b H )

where ŷTS and ŷH represent the outcomes TS and H respectively, bTS, bH and WTS, WH represent the biases and weights of those outcome layers respectively. The Objective function is mentioned below:

(13) L = λ T S . L T S ( y T S , y T S ) + λ H . L H ( y H y H )

where weights for H and TS are denoted by λH and λTS respectively, objective functions for H and TS are denoted by LH and LTS respectively and the computed values of H and TS are denoted by yH and yTS respectively. Since tensile strength and hardness possess different numerical ranges and units, both response variables were subjected to Min–Max normalization prior to MTNN training. Consequently, the task-specific loss functions in Equation (13) were computed using normalized values, ensuring balanced learning and preventing either response from dominating the optimization process due to scale differences.

To minimize overfitting, both the MTNN and ANN models were trained using normalized input data with an appropriate network architecture. The dataset was divided into training and testing subsets to independently evaluate model performance. The final model was selected based on its consistent prediction accuracy on both datasets, thereby ensuring improved generalization capability and model robustness.

Although the experimental dataset comprises only 90 samples, the proposed MTNN architecture is well suited for this application because it simultaneously learns two correlated responses (tensile strength and hardness) through shared hidden layers, thereby improving feature utilization and learning efficiency compared with training separate single-task models. Unlike conventional deep learning models requiring large datasets, the adopted MTNN employs a compact architecture with a limited number of trainable parameters, making it appropriate for engineering datasets of modest size. Furthermore, input variables were normalized, the network architecture was carefully selected to avoid unnecessary complexity, and independent training and testing subsets were employed to assess the model’s generalization capability. These measures effectively reduce the likelihood of overfitting while ensuring robust and reliable prediction performance.

4.3. Optimization model formulation

A multi-objective optimization model typically consists of three aspects: parameter variables, constraints, and objective functions [14, 22]. The variables involved in the objective optimization are the five process parameters, which can be expressed as follows:

(14) V = ( x A F , x R S , x T S , x P G , x T A ) T

Based on the FSW parameters requirements, the spectrum of the chosen parameters was fixed and is mentioned below:

(15) s . t . { 3 k N < A F < 5 k N 900 r p m < R S < 1100 r p m 1.5 m m / s e c < T S < 2.5 m m / s e c 0 ° < T A < 1 ° P G : S S , S C , T C

Since H and TS are the performance measures, the optimization strategy being employed for this research paper is a reduction-based optimization task. The optimal span for hardness (H) was fixed in the range of 47 HV to 58 HV and for tensile strength (TS), the range was 223 to 236 MPa. The nugget zone exhibited finely refined uniformly scattered grain structure whenever the TS reached 225 MPa and for a hardness of 55 HV. As a result, the objective function for the optimization was framed as described below:

(16) O b j e c t i v e f u n c t i o n = { min ( f T S ( A F , R S , T S , T A , P G ) 225 ) 2 min ( f H ( A F , R S , T S , T A , P G ) 55 ) 2

Accordingly, the objective function was formulated as the minimization of the squared deviation between the MTNN-predicted responses and their respective target values. This formulation guarantees non-negative objective values and guides the PSO algorithm toward parameter combinations that simultaneously achieve the desired tensile strength and hardness. For the above-described optimization model, it must be remembered that fTS and fH do not describe explicit analytical expressions; instead, they are approximated using the proposed MTNN architecture, comprising two distinctive outcome branches dedicated for forecasting TS and H, respectively. The optimized model was addressed by PSO algorithm employing a swarm population of 100 particles and restricting the iteration count to 500. The inertia weight parameter w = 0.7 together with cognitive and social learning factors, C1 and C2 both being equal to 0.9 were chosen based on the realistic FSW scenario and fabrication of defect free joints. Proposed optimization approach was implemented by employing a Python–based computational framework.

4.4. Execution of SHAP

In a friction stir welded joint, the influence of every parameter on the mechanical attributes of the joints was important for modifying the employed parameters for attaining flaw free joints. Moreover, comprehending the relationship amidst the FSW parameters and mechanical attributes permits the optimization process to attain an ideal state, thereby enhancing the joint efficiency and quality [10, 15]. So, there prevails an inevitable need for computing the impact of every FSW parameter on the mechanical attributes by means of employing the SHAP optimization technique. The procedure being employed for implementing the SHAP technique in this research work is described below:

  • 1st step: A non-linear dependency architecture amidst the chosen FSW parameters and experimental outcomes (namely hardness and tensile strength) was formulated by employing ML techniques.

  • 2nd step: Computing the SHAP values

    • a)

      Reference point determination: Mean of hardness (H) and tensile strength (TS) data set was selected as the reference point for identifying the model outcomes modifications in an easier manner

    • b)

      Generate parameter configurations: Distinctive combinations of FSW parameters were generated for computing the values of SHAP.

    • c)

      Anticipatory model: For every parameter configuration, the formulated model’s outcome was computed and this was accomplished by forecasting the model several times, every time employing a unique combination of FSW parameters

    • d)

      Compute changes in prediction: Change in the anticipated value (for every sample) associated with the reference point was computed. This quantifies the deviation between the predicted outcome of the model and the established reference output for the respective parameter subset.

    • e)

      Weightage computation: A statistical model was employed for computing the contribution of every parameter w.r.t the modifications in the anticipated outcomes.

    • f)

      Attain SHAP values: For every chosen FSW parameter, a SHAP values was attained, describing the extent to which that particular parameter impacts the forecasted model’s outcome.

  • 3rd step: Interpretation of SHAP scores: After computing the SHAP scores, the outcome behavior of the formulated learning model was analyzed. This was accomplished by analyzing the graphical representation of the contribution scores. By examining the Shapley scores, the most dominant parameter which had highly impacted the hardness and tensile strength, can be identified, including their interdependencies [29]. This enables an in-depth comprehension of parameter combination mechanisms during FSW process, permitting fine control of input parameters settings and providing more robust insights for smart FSW process control. In the present work, the mean values of tensile strength and hardness across the experimental dataset were adopted as the SHAP baseline (reference) values. Consequently, the SHAP value of each input parameter quantifies its contribution toward increasing or decreasing the model prediction relative to this average prediction, thereby providing a consistent and interpretable measure of individual parameter influence.

5. INFERENCES AND DISCUSSIONS

5.1. Baseline techniques and model assessment indicators

From the detailed literature review, it was inferred that ANNs and some unique ML techniques including RR, RFR and SVR (i.e., support vector regression) were effective for addressing manufacturing problems involving minimal sample numbers and as a result, in this research paper, ANN, RFR, RR and SVR were chosen as the optimization comparison techniques. Table 3 describes in detail the comparison methodology being used in this paper. ANN and MTNN were iterated individually for 500 training cycles for formulating an optimized prophetic network. Moreover, all the comparison methodologies were executed in a Python-based computational environment. In the training phase, MSE (i.e., mean squared error) metric was employed for quantifying the deviation amidst the measured and predicted values. This evaluation metric regulates the training modules by means of reverse error propagation and enables incremental optimization of the formulated model’s biases and weights.

Table 3
Comprehensive configuration details of the evaluated methodologies.

At the same time, model’s over complexity is a frequent issue while training minimal samples [24, 35]. To solve this issue, regularization layers were included following the 1st two common layers and these regularization layers haphazardly deactivate a particular percentage of neurons, especially 10% in our scenario, for minimizing variance of the formulated model and for enhancing model’s robustness. The forecasting accuracy is very important w.r.t the PSO based optimization outcomes. So, it is essential to estimate the meta model’s prediction accuracy in a quantitative manner by choosing suitable performance metrics. Performance indicators like mean prediction deviation (MPD) and R2 (i.e., coefficient of determination), etc., were employed frequently to evaluate the forecasting accuracy of data driven model. Formulas employed for these scenarios are described below:

(17) M P D = 1 n i = 1 n | y i y i |
(18) R 2 = ( 1 i = 1 n ( y i y t ) 2 i = 1 n ( y i y t ) 2 ) X 100 %

where the measured values are denoted by yi, anticipated values by ŷi, average values by y and the number of samples in the collected data are denoted by n.

5.2. Analysis of data

The collected data D was subjected to statistical analysis and Table 4 describes the associated analytical metrics. Table 4 presents the descriptive statistics of the input parameters and output responses used for model development. Welding speed (WS), axial force (AF), rotational speed (RS), and tool tilt angle (TA) are continuous variables; therefore, their statistical measures are reported. In contrast, pin geometry (PG) is a categorical parameter representing different tool pin profiles, and hence conventional statistical descriptors are not applicable (denoted as NA). The measured tensile strength (TS) and hardness (H) exhibited adequate variability across the experimental trials, providing a reliable dataset for training the proposed MTNN model and evaluating its predictive performance.

Table 4
Statistical analysis of collected data.

Even though serious efforts were put forward to eliminate the occurrence of the flaws, during realistic experimental runs, the intricate relationships amidst the FSW parameters makes it impossible to attain entirely flaw free AZ80A Mg alloy joints, possessing superior mechanical attributes (H and TS). This demands for generating an interdependency evaluation matrix for assessing the strength of the relationships amidst the chosen FSW parameters, hardness, tensile strength and Figure 5 illustrates the interdependency evaluation matrix for FSW of AZ80A Mg alloy.

Figure 5
Interdependency evaluation matrix for the employed FSW parameters and outcomes.

By observing the interdependency evaluation matrix illustrated in the Figure 5, it can be inferred that there prevails a robust correlation amidst the outcome TS and the parameters AF, RS, WS, TA, PG with the dependence indicators as 0.40, 0.41, 0.35, 0.24 respectively. In addition to these parameters, TS was also impacted significantly by the tool possessing a tapered cylindrical (TC) pin geometry with a 0.52 dependence indicator as seen in the Figure 5. At the same time, a low-level negative dependency prevails with the pin geometries, namely SS (–0.27) and SC (–0.17) respectively for TS. Likewise, H (i.e., hardness) of the friction stir welded AZ80A Mg alloy joints can be observed to have high degree of association with AF, WS, RS, TA and tool Pin geometry TC with their relevant dependent indicators being 0.76, 0.59, 0.23, 0.18 and 0.38 respectively and the linkage of H with other pin geometries namely SS and SC are marginally weak with –0.22 and –0.06 dependent indicators respectively. It can also be inferred that all the chosen 5 FSW parameters had demonstrated a dynamic, intricate interdependencies with H and TS.

5.3. Predictive effectiveness of the evaluated techniques

The predictive effectiveness of the evaluated techniques on the model fitting and assessment sets are graphically illustrated in the Figure 6(a) & (b), (c) & (d) respectively. From these graphs, it can be inferred that the largest value of R2 and minimal MPD values have been attained in the model fitting (i.e., learning) phase. Moreover, the formulated MTNN model had demonstrated improved robustness in R2 for both the H and TS data segments. At the same time, w.r.t MPD, the statistical dispersion of the TS data segments in the assessment phase are higher than that of the H, leading to a higher MPD on the TS assessment set, despite identical R2 performance. From these graphs, it can also be inferred that among the five evaluated techniques, formulated MTNN model had demonstrated most accurate predictions, by exhibiting the largest R2 (i.e., goodness of fit) and minimal forecasting errors (i.e., MPD) [6, 11].

Figure 6
Predictive effectiveness of the evaluated techniques on the model fitting and assessment sets (a) & (c) MPD and (b) & (d) R2.

Statistical performance indicators of all the 5 evaluation techniques during the model fitting and assessment sets are described in Table 5 and 6 respectively. From the Table 5, it can be seen that in the network adaptation (training) phase, the model formulated using RR had demonstrated the largest MPD values on both TS and H on the data segments, spanning around 13.23 and 2.61 respectively, MPD of the MTNN’s model is in the range of 1.45 for H and 5.83 for TS, and the MPDs of the remaining 3 formulated models were spread in the vicinity of 1.78 to 2.21 for H and 6.94 to 10.67 for TS. This helps to understand that during the network adaptation phase, the model formulated by MTNN had demonstrated the minimal forecast discrepancies, whereas RR based model had exhibited the highest approximation inaccuracies. MPD values of the MTNN’s model lies in the interval of 2 to 6.5, whereas the MPD of SVR’s model lies in the band of 2.7 to 10 and the MPDs of the remaining 3 models are in the vicinity of 2 and 18, revealing us that the model formulated using MTNN had demonstrated minimized MPDs during both model fitting and assessment datasets.

Table 5
MPDs of the distinctive evaluation techniques.

From Table 6, it can be observed that during the model fitting phase, for both H and TS data samples, model formulated by MTNN had demonstrated the largest R2 value around 93–95% and the lowest value of R2 was exhibited by RR in the span of 83 to 84% and the remaining 3 models had exhibited R2 values in the spectrum of 84.5 to 92.5%. From this, it can be apprehended that during the model fitting phase, MTNN exhibits the superior modeling accuracy. During the assessment phase, for both H and TS data samples, the R2 value of the MTNN’s model was the largest nearer to 84–87%, whereas the R2 value of the SVR’s model was the least in the range of 75–79% and the R2 values of the remaining 3 models were roughly amidst 80–87%.

Table 6
R2 values of the distinctive evaluation techniques.

So, in the assessment phase also, model formulated by MTNN will be the perfect choice. Taking into consideration, both the R2 and MPD criterions, it can be understood that the MTNN model had demonstrated robust performance across both the evaluation and model fitting phases, with comparatively minimized error levels [4, 16].

To statistically confirm the superiority of MTNN over the baseline models, paired t-tests were performed on the absolute prediction errors of each model against MTNN, using the 18-sample assessment set for both TS and H. Wilcoxon signed-rank tests were also conducted as a non-parametric cross-check given the limited sample size.

As evident from Table 7, MTNN demonstrated statistically significant lower prediction error compared to RR, ANN, and SVR for both TS and H (p < 0.05, with 95% confidence intervals excluding zero). Against RFR, MTNN showed significant improvement for TS (p = 0.006), while the improvement for H showed a favorable trend but did not reach conventional significance (p = 0.051). This is attributed to RFR’s comparatively strong ensemble-based performance on the hardness dataset. It should be noted that the assessment set comprises only 18 samples, and larger validation datasets would allow for narrower confidence intervals and more conclusive significance testing in future work.

Table 7
MTNN vs each model—paired t-test on absolute errors (n = 18).

5.4. Interpretation of process parameter contributions based on SHAP values

The SHAP based comparative significance of the FSW parameters in impacting the outputs namely TS and H is illustrated in the Figure 7(a)(e). In the illustrated bar plots, the green-coloured bars describe the impact of parameters on TS, whereas the orange-coloured bars describe the FSW parameters on H. From these plots, it can be visualized that the impact of every parameter not only varies amidst the two outcomes, but also across distinctive forecasting models.

Figure 7
SHAP based comparative significance of the FSW parameters in impacting the outputs (a) RR (b) MTNN (c) RFR (d) ANN and (e) SVR.

Among all the 5 evaluation techniques, MTNN exhibited the largest forecasting accuracy based on the prior formulated model assessment outcomes. With MTNN taken as the guiding model, the significant ranking of the FSW parameters for forecasting TS is WS > AF > PG (TC) > RS > TA > PG (SS) > PG (SC), whereas for the outcome H, the ranking changes and is WS > PG (TC) > AF > RS > TA > PG (SS) > PG (SC). These rankings were found to coincide highly with the experimental inferences and statistical results of previously conducted FSW investigational experiments.

For example, larger welding speed (WS) was found to ensure balanced heat input and facilitates effective stirring of the plasticized metal, thereby improving the tensile strength of the fabricated joints. At the same time, tapered cylindrical pin geometry (TC) was reported to play a significant part in regulating the flow of the plasticized metal and refinement of grains in the nugget zone, thereby highly impacting the joint’s hardness. Moreover, axial force (AF) and rotational speed (RS) were also proven to impact both H and TS in a reasonable manner.

The differing SHAP rankings observed for the RR and SVR models are primarily attributed to their comparatively lower prediction accuracy and limited ability to capture the complex nonlinear interactions among the FSW parameters. Consequently, their feature importance rankings deviate from the experimentally observed metallurgical trends. In contrast, the proposed MTNN more effectively learned these nonlinear relationships, resulting in SHAP interpretations that are consistent with both the experimental observations and the underlying physical metallurgy of the FSW process.

In addition to this, models possessing relatively inferior forecasting accuracy, like SVR and RR, the SHAP based ranking of the parameter significance differs from already proven trends. This divergence highlights the restricted analytical insight of low precise models and again substantiates the MTNN model’s superiority in grabbing both the intricate interdependencies and physical relevance of the FSW parameters [29, 30]. From these outcomes, it can be understood that the model possessing larger forecasting accuracy not only generate enhanced predictions about the outcomes, but also improve the transparency by consistently reflecting the actual impact of the FSW parameters on performance indicators. This emphasizes the merit of combining ML approaches with process specific understanding in optimizing the parameters during FSW of AZ80A Mg alloys [19, 37].

5.5. Comparative analysis of optimization outcomes

Above described 5 evaluation techniques namely MTNN, RR, RFR, ANN and SVR were employed for formulating 5 distinctive meta models respectively and PSO was used for attaining optimized parameter combinations exhibiting maximized performance. The convergence behavior of the PSO algorithm was monitored by tracking the best fitness value obtained at each iteration. A rapid improvement in the objective function was observed during the initial iterations owing to the extensive exploration capability of the swarm, followed by gradual refinement as the particles converged toward the global optimum. The fitness value became nearly constant after successive iterations, indicating stable convergence without noticeable oscillations or premature stagnation. This convergence trend confirms the effectiveness and stability of the PSO algorithm in identifying the optimal FSW parameter combination.

Below mentioned equation was employed for computing the percentage of deviation from the forecasted performance metric and was used for comparing the optimization outcomes of those distinctive evaluation techniques.

(19) P e r c e n t a g e o f d e v i a t i o n = | y 0 y i | y 0 X 100 %

where the target metrics are denoted by y0 and forecasted optimal values attained from the multi criteria decision making model are denoted by yi.

Table 8 describes the forecasted optimal values and the target metrics. From this table, it can be understood that all the employed 5 evaluation techniques have the same target metrics, namely 225 MPa for TS and 55 HV for H. From the attained optimization outcomes for TS, it can be inferred that the result generated by the MTNN driven PSO is 222.25 MPa, which is very nearer to the target metrics and the remaining 4 techniques have generated comparable outputs. In the H scenario, ANN driven PSO and SVR driven PSO had yielded values of 54 HV and 57 HV respectively, which were nearer to the target metrics of H. On the contrary, the MTNN driven PSO model had yielded a relatively larger value of 58 HV.

Table 8
Forecasted optimal values and target metrics.

Deviation metrics on TS and H are illustrated in Figure 8(a) and (b) respectively, where the lesser divergence values imply enhanced optimization efficacy across the employed evaluation techniques. For TS, the discrepancy of MTNN driven PSO stands nearly at 1.22%, noticeably outperforming the remaining 4 evaluation techniques. On the contrary, for H, the discrepancy in the MTNN driven PSO exceeds the deviation metrics of the remaining models.

Figure 8
Optimization outcomes, target metrics and deviation percentages of different evaluation techniques on (a) TS and (b) H.

At the same time, this discrepancy lies within the acceptable engineering spectrum, provided that the permissible window for H lies amidst 47 HV to 58 HV. Taking into consideration, the aggregate discrepancies for both TS and H, it can be visualized that the MTNN driven PSO had recorded the minimal combined deviation of about 3.34%, succeeded by ANN driven PSO (4.16%), SVR driven PSO (4.82%), RFR driven PSO (5.30%), and RR driven PSO (6.63%). As a result, the formulated MTNN driven PSO had exhibited the most robust and accurate forecasting capability among all other employed evaluation techniques.

5.6. Experimental confirmation of predictions

Confirmatory experiments were performed by employing the optimized parameter combinations attained from the above mentioned 5 evaluation techniques and a total of 5 tensile specimens was extracted from the friction stir welded AZ80A Mg alloy joints. With the objective of validating the optimization outcomes and for understanding the mode of failure of the joints, extracted tensile specimen were subjected to tensile test. SEM (i.e., scanning electron microscopy) images of the fractured surfaces of these tensile specimens extracted from the joints fabricated under parameter combinations proposed by the above mentioned 5 evaluation techniques are illustrated in the Figure 9(a)(e).

Figure 9
SEM images of the fractured surfaces of these tensile specimen extracted from the joints fabricated under parameter combinations proposed by (a) MTNN (b) RR (c) RFR (d) ANN and (e) SVR.

From these SEM images, it can be inferred that the majority of the tensile specimens had encountered a ductile mode of fracture in varying degrees including void coalescence, tear ridges, dimples etc. At the same time, it can be inferred that there prevails variation in the size and morphological distribution of these structures, based on their welding conditions, indicating that the quality of the fabricated joint and input of frictional heat had reasonably impacted their failure mode [7, 38]. Amidst all, the fractured surface of the specimen fabricated by employing the parameter combinations optimized by the MTNN driven PSO, illustrated in the Figure 9(a) inherits homogeneous scattered, deep & hemisphere shaped dimple structures, which is an indication of occurrence of plastic deformation in a reasonable manner, before fracture and this is normally associated with larger tensile strength and superior joint integrity.

From this SEM image, it can be understood that the employed optimized parameter combinations had led to the generation of flaw free, perfectly consolidated joint. At the same time, the SEM image of the joint fabricated using the optimized scenarios attained from RR driven PSO, illustrated in the Figure 9(e) exhibits superficial dimples distributed between non-textured, smooth surfaces, revealing a combined fracture mode comprising both brittle as well as ductile attributes. Existence of small pore accumulations and grain boundary cracks reveals that inadequate flow of the plasticized metal and inappropriate plasticization caused due to generation of frictional heat in reduced volumes, which in turn had led to the decline in the mechanical properties of the joints [2, 5].

Optimized parameters, optimized outputs and validated outcomes for the above mentioned 5 evaluation techniques are described in the Table 9. Minimal deviations amidst the validated and optimized outputs are computed using the below mentioned equation:

(20) M i n i m a l d e v i a t i o n ( % ) = | y v y o | y 0 X 100 %
Table 9
Validated experimental outputs, forecasted outcomes and percentage deviations.

where the validated experimental outputs and optimized outputs attained from the above mentioned 5 evaluation techniques are denoted by yv and y0 respectively.

Provided that all the 5 evaluation techniques were combined with PSO for optimizing FSW parameters, the dissimilarities in the mechanical attributes are mainly due to the forecasting accuracy of the meta models being used. As described in the previous sections that the models formulated by MTNN, ANN, RFR had exhibited perfect forecasting capability and contributed for more stable optimization outputs. For instance, WS suggested by these models were 2.67 mm/sec, 2.36 mm/sec, and 2.29 mm/sec had contributed for larger tensile strengths of 280.04 MPa, 270.63 MPa, and 260.42 MPa respectively. From this it can be understood that the welding speed plays an important role in ascertaining the tensile strength of the friction stir welded joints.

In addition to this, the impact of the tool’s tilt angle (TA) and its pin geometry (PG) are also evident. Larger values of tensile strengths and hardness were exhibited by AZ80A Mg alloy joints fabricated by employing a tapered cylindrical pin and a 1° tilt angle, when compared with that of the joints fabricated using other PG and at other TAs. From this, it can be understood that the PG and TA also plays an inevitable part in impacting frictional heat generation, flow of plasticized metal etc., and in turn influence the mechanical attributes of the joints.

It can be summarized that the confirmatory experiments substantiate the optimization outcomes, highlighting the key aspects of the accurate forecasting meta models in predicting the ideal combination of optimized parameters of FSW process. These research findings also reaffirm the significant impact of WS, AF, RS, PG and TA on the TS and H of the friction stir welded AZ80A Mg alloy joints. Validated experimental outputs, target metrics and percentage deviations for TS and H, by employing the above mentioned 5 evaluation techniques are illustrated in the Figure 10(a) & (b) respectively.

Figure 10
Validated experimental outcomes, target metrics and deviation percentages of different evaluation techniques on (a) TS and (b) H.

From these graphs, it can be visualized that the MTNN driven PSO model had exhibited the best forecasting performance with minimized deviations of 2.42% for TS and 3.64% for H, reflecting perfect alignment with the experimental outcomes. Comprehensively, MTNN driven PSO recorded a mean deviation of 3.03% surpassing ANN driven PSO (5.47%), RFR driven PSO (6.82%), RR driven PSO (6.85%) and SVR driven PSO (7.32%). These outcomes affirm enhanced capacity of MTNN driven PSO in modeling intricate, multi parameter relations amidst key FSW parameters namely AF, WS, RS, PG, TA and mechanical attributes (TS and H). Integrated learning architecture of this MTNN driven PSO had efficiently reflected the interdependency amidst TS and H, thereby enabling enhanced adaptability. Comprehensively, these outcomes announce MTNN driven PSO as a robust tool for optimizing the FSW parameters for attaining superior quality, flaw free AZ80A Mg alloy joints

6. CONCLUSIONS

Optimizing the parameters during FSW of AZ80A Mg alloys is an intricate task owing to the substantially non-monotonic and interdependent linkages amidst various input parameters and the associated mechanical outcomes. While traditional ML algorithms can interpret fundamental behaviors, their efficiency declines when employed for parameter-intensive, multiple criteria architectures, where overlapping impacts on target metrics are prevalent. With the objective of addressing these limitations and challenges, in this research work, a contemporary framework integrating MTNN with PSO was employed for predicting and optimizing the FSW parameters. Significant findings of this research work are summarized below:

  • The MTNN model outperformed conventional models by effectively learning shared and task-specific representations for TS and H, which is critical for multi-output FSW modeling. The SHAP method further enhanced the model transparency by quantifying parameter importance without requiring additional physical experiments.

  • In traditional FSW optimization, process refinement relied heavily on trial-and-error and single-objective focus. However, MTNN-PSO efficiently balanced both mechanical outcomes. It was found that the parameter welding speed (WS) and axial force (AF) had significantly influenced both TS and H. Meanwhile, pin geometry (PG), especially tapered cylindrical (TC) pin had a dominant effect on tensile strength owing to its contribution in flow of plasticized metal & stirring action, while rotational speed (RS) and tool’s tilt angle (TA) showed varying impacts depending on the tilt configuration.

  • Validation experimental results confirmed that joints produced using MTNN-PSO optimized parameters had demonstrated the largest tensile strength (~230.45 MPa) and hardness (~57 HV), supported by SEM observations showing homogeneously distributed, deep, hemisphere-shaped dimples, which is a signature of ductile fracture and good metallurgical bonding.

  • Proposed MTNN-PSO-based framework not only improved the modeling accuracy and optimization reliability in FSW of AZ80A Mg alloy but also provided interpretable insights into parameter–response interactions, thereby offering a scalable and intelligent pathway for future solid-state welding process design.

7. DATA AVAILABILITY

All data that support the findings of this study are included within the article.

8. BIBLIOGRAPHY

  • [1] LASKA, A., SADEGHI, B., SADEGHIAN, B., et al, “Temperature evolution, material flow, and resulting mechanical properties as a function of tool geometry during friction Stir Welding of AA6082”, Journal of Materials Engineering and Performance, v. 32, n. 15, pp. 10655–10668, 2023. doi: https://doi.org/10.1007/s11665-023-08671-1.
    » https://doi.org/10.1007/s11665-023-08671-1
  • [2] NALLASAMY, V., BABUCHELLAM, A.K., VADIVELU, V., et al, “Optimizing underwater friction-stir welding parameters for AA356/SiC composites using CoCoSo and MEREC: enhancing joint performance and quality”, Matéria (Rio de Janeiro), v. 30, pp. e20240948, 2025. doi: https://doi.org/10.1590/1517-7076-rmat-2024-0948.
    » https://doi.org/10.1590/1517-7076-rmat-2024-0948
  • [3] YILDIZ, M., OZTURK, F., SHEIKH-AHMAD, J., “A comprehensive review on friction stir welding of high-density polyethylene”, Arabian Journal for Science and Engineering, v. 48, n. 8, pp. 11167–11210, 2023. doi: https://doi.org/10.1007/s13369-023-08048-5.
    » https://doi.org/10.1007/s13369-023-08048-5
  • [4] SINGH, V.P., KUMAR, A., KUMAR, R., et al, “Effect of rotational speed on mechanical, microstructure, and residual stress behaviour of AA6061-T6 alloy joints through friction stir welding”, Journal of Materials Engineering and Performance, v. 33, n. 6, pp. 3706–3721, 2024. doi: https://doi.org/10.1007/s11665-023-08527-8.
    » https://doi.org/10.1007/s11665-023-08527-8
  • [5] MANICKAM, S.K., PALANIVEL, I., “Practical implications of FSW parameter optimization for AA5754-AA6061 alloys”, Matéria (Rio de Janeiro), v. 29, n. 4, pp. e20240482, 2024. doi: https://doi.org/10.1590/1517-7076-rmat-2024-0482.
    » https://doi.org/10.1590/1517-7076-rmat-2024-0482
  • [6] NGUYEN, T.T., NGUYEN, C.T., VAN, A.L., “Sustainability-based optimization of dissimilar friction stir welding parameters in terms of energy saving, product quality, and cost-effectiveness”, Neural Computing & Applications, v. 35, n. 7, pp. 5221–5249, 2023. doi: https://doi.org/10.1007/s00521-022-07898-8.
    » https://doi.org/10.1007/s00521-022-07898-8
  • [7] NISHANT, JHA, S.K., PRAKASH, P., “Numerical analyses of underwater friction stir welding using computational fluid dynamics for dissimilar aluminum alloys”, Journal of Materials Engineering and Performance, v. 33, n. 17, pp. 12620–12637, 2024.
  • [8] PALANIVEL, V., JOHNSON, P., MUNIMATHAN, A., et al, “Finite element analysis of friction stir welding process to predict temperature distribution”, Matéria (Rio de Janeiro), v. 29, n. 4, pp. e20240465, 2024. doi: https://doi.org/10.1590/1517-7076-rmat-2024-0465.
    » https://doi.org/10.1590/1517-7076-rmat-2024-0465
  • [9] ZHANG, C., WANG, Y., ZHAO, Z., et al, “Performance-driven closed-loop optimization and control for smart manufacturing processes in the cloud-edge-device collaborative architecture: a review and new perspectives”, Computers in Industry, v. 162, pp. 104131, 2024. doi: https://doi.org/10.1016/j.compind.2024.104131.
    » https://doi.org/10.1016/j.compind.2024.104131
  • [10] DARWINS, A.K., LEWISE, K.A.S., FAHAD, M., et al, “Parametric optimization of friction stir welding of ZE42 using NSGA-II”, International Journal on Interactive Design and Manufacturing, v. 18, n. 8, pp. 3193–3205, 2024. doi: https://doi.org/10.1007/s12008-023-01480-9.
    » https://doi.org/10.1007/s12008-023-01480-9
  • [11] DESIKAN, S., VITTEL, R., SIVASUNDAR, V., et al, “Microstructure, mechanical properties of dissimilar friction stir welded AA6063/AA5052 alloys, and optimization of process parameters using Box Behnken-TOPSIS approach”, Kovove Materialy, v. 62, pp. 363–375, 2024.
  • [12] LIU, B., YANG, J., ZHANG, X., et al, “Development and application of magnesium alloy parts for automotive OEMs: a review”, Journal of Magnesium and Alloys, v. 11, n. 1, pp. 15–47, 2023. doi: https://doi.org/10.1016/j.jma.2022.12.015.
    » https://doi.org/10.1016/j.jma.2022.12.015
  • [13] KASMAN, S., “The effects of pin offset for FSW of dissimilar materials: a study for AA 7075–AA 6013”, Matéria (Rio de Janeiro), v. 25, n. 2, pp. e-12612, 2020. doi: https://doi.org/10.1590/s1517-707620200002.1012.
    » https://doi.org/10.1590/s1517-707620200002.1012
  • [14] KARUMURI, S., HALDAR, B., PRADEEP, A., et al, “Multi-objective optimization using Taguchi based grey relational analysis in friction stir welding for dissimilar aluminium alloy”, International Journal on Interactive Design and Manufacturing, v. 18, n. 3, pp. 1627–1644, 2024. doi: https://doi.org/10.1007/s12008-023-01529-9.
    » https://doi.org/10.1007/s12008-023-01529-9
  • [15] VIJAYAKUMAR, S., MANICKAM, S., SEETHARAMAN, S., et al, “Examination of friction stir-welded aa 6262/5456 joints through the optimization technique”, Advances in Materials Science and Engineering, v. 11, pp. 4527595, 2022. doi: https://doi.org/10.1155/2022/4527595.
    » https://doi.org/10.1155/2022/4527595
  • [16] WANG, P.E., GHASSEMI-ARMAKI, H., POUR, M., et al, “Applicable and generalizable machine learning for intelligent welding in automotive manufacturing”, Welding in the World, v. 69, n. 5, pp. 1349–1384, 2025. doi: https://doi.org/10.1007/s40194-025-01951-5.
    » https://doi.org/10.1007/s40194-025-01951-5
  • [17] SAFDAR, M., PAUL, P.P., LAMOUCHE, G., et al, “Fundamental requirements of a machine learning operations platform for industrial metal additive manufacturing”, Computers in Industry, v. 154, pp. 104037, 2024. doi: https://doi.org/10.1016/j.compind.2023.104037.
    » https://doi.org/10.1016/j.compind.2023.104037
  • [18] ZENG, D., WU, D., LUO, Z., et al, “A performance comparison of deep learning and shallow machine learning in acoustic emission monitoring of aluminium alloy pulsed laser welding”, Soft Computing, v. 28, n. 14, pp. 10263–10279, 2024. doi: https://doi.org/10.1007/s00500-024-09778-w.
    » https://doi.org/10.1007/s00500-024-09778-w
  • [19] ZHOU, K., FENG, P., FENG, F., et al, “A deep transfer learning model for online monitoring of surface roughness in milling with variable parameters”, Computers in Industry, v. 164, pp. 104199, 2025. doi: https://doi.org/10.1016/j.compind.2024.104199.
    » https://doi.org/10.1016/j.compind.2024.104199
  • [20] GUO, B., LI, X., “Arc bubble edge detection method based on deep transfer learning in underwater wet welding”, Scientific Reports, v. 14, n. 1, pp. 22628, 2024. doi: https://doi.org/10.1038/s41598-024-73516-3. PubMed PMID: 39349710.
    » https://doi.org/10.1038/s41598-024-73516-3
  • [21] TIAN, W., HU, P., ZHANG, C., “Optimization framework of laser oscillation welding based on a deep predictive reward reinforcement learning net”, Journal of Intelligent Manufacturing, 2024.
  • [22] SELVAM, T., SOLAIYAPPAN, A., “Optimization of friction stir welding parameters for joining dissimilar aluminum alloys AA 6061-T651 and AA 7075-T651 with Mg and Cr reinforcement”, Matéria (Rio de Janeiro), v. 30, pp. e20250152, 2025. doi: https://doi.org/10.1590/1517-7076-rmat-2025-0152.
    » https://doi.org/10.1590/1517-7076-rmat-2025-0152
  • [23] JEYAKRISHNAN, S., VIJAYAKUMAR, S., NAGA SWAPNA SRI, M., et al, “An integration of RSM and ANN modelling approach for prediction of FSW joint properties in AA7178/AA5456 alloys”, Canadian Metallurgical Quarterly, v. 64, n. 1, pp. 43–60, 2025. doi: https://doi.org/10.1080/00084433.2024.2310344.
    » https://doi.org/10.1080/00084433.2024.2310344
  • [24] FARAJ, A.K., SALIH, A.K., AHMED, M.A., et al, “Fracture pressure prediction in carbonate reservoir using artificial neural networks”, Petroleum Chemistry, v. 64, n. 5, pp. 796–803, 2024. doi: https://doi.org/10.1134/S0965544124050050.
    » https://doi.org/10.1134/S0965544124050050
  • [25] AKHTAR, N., ALHARTHI, M.F., “Enhancing accuracy in modelling highly multicollinear data using alternative shrinkage parameters for ridge regression methods”, Scientific Reports, v. 15, n. 1, pp. 10774, 2025. doi: https://doi.org/10.1038/s41598-025-94857-7. PubMed PMID: 40155439.
    » https://doi.org/10.1038/s41598-025-94857-7
  • [26] DEJENE, N.D., LEMU, H.G., GUTEMA, E.M., “Effects of process parameters on the surface characteristics of laser powder bed fusion printed parts: machine learning predictions with random forest and support vector regression”, International Journal of Advanced Manufacturing Technology, v. 133, n. 11–12, pp. 5611–5625, 2024. doi: https://doi.org/10.1007/s00170-024-14087-5.
    » https://doi.org/10.1007/s00170-024-14087-5
  • [27] QIU, Y., PING, J., SHU, L., et al, “Defect monitoring of high-power laser-arc hybrid welding process based on an improved channel attention convolutional neural network”, Journal of Intelligent Manufacturing, v. 36, n. 7, pp. 2657–2676, 2025. doi: https://doi.org/10.1007/s10845-024-02354-x.
    » https://doi.org/10.1007/s10845-024-02354-x
  • [28] LI, S., CORNEY, J., “Multi-view expressive graph neural networks for 3D CAD model classification”, Computers in Industry, v. 151, pp. 103993, 2023. doi: https://doi.org/10.1016/j.compind.2023.103993.
    » https://doi.org/10.1016/j.compind.2023.103993
  • [29] KHAN, A., KHAN, M., KHAN, W.A., et al, “Predicting pile bearing capacity using gene expression programming with SHapley Additive exPlanation interpretation”, Discover Civil Engineering, v. 2, n. 1, pp. 58, 2025. doi: https://doi.org/10.1007/s44290-025-00215-x.
    » https://doi.org/10.1007/s44290-025-00215-x
  • [30] KILIC, K., IKEDA, H., NARIHIRO, O., et al, “A soft ground micro TBM’s specific energy prediction using an eXplainable neural network through Shapley additive explanation and Optuna”, Bulletin of Engineering Geology and the Environment, v. 83, n. 5, pp. 175, 2024. doi: https://doi.org/10.1007/s10064-024-03670-5.
    » https://doi.org/10.1007/s10064-024-03670-5
  • [31] TANG, D., DAI, M., SALIDO, M.A., et al, “Energy-efficient dynamic scheduling for a flexible flow shop using an improved particle swarm optimization”, Computers in Industry, v. 81, pp. 82–95, 2016. doi: https://doi.org/10.1016/j.compind.2015.10.001.
    » https://doi.org/10.1016/j.compind.2015.10.001
  • [32] WANG, Y., ZHANG, Y., SHUANG, Z., et al, “A novel hybrid differential particle swarm optimization based on particle influence”, Cluster Computing, v. 28, n. 1, pp. 65, 2025. doi: https://doi.org/10.1007/s10586-024-04783-y.
    » https://doi.org/10.1007/s10586-024-04783-y
  • [33] DAS, P.P., CHAKRABORTY, S., “In search of the best multi-criteria decision making-particle swarm optimization-based hybrid approach for parametric optimization of friction stir welding processes”, Opsearch, v. 61, n. 5, pp. 1764–1794, 2024. doi: https://doi.org/10.1007/s12597-024-00757-1.
    » https://doi.org/10.1007/s12597-024-00757-1
  • [34] KAHHAL, P., GHASEMI, M., KASHFI, M., et al, “A multi-objective optimization using response surface model coupled with particle swarm algorithm on FSW process parameters”, Scientific Reports, v. 12, n. 1, pp. 2837, 2022. doi: https://doi.org/10.1038/s41598-022-06652-3. PubMed PMID: 35181705.
    » https://doi.org/10.1038/s41598-022-06652-3
  • [35] DORBANE, A., HARROU, F., SUN, Y., et al, “Machine learning for modeling and defect detection of friction stir welds: a review”, Journal of Failure Analysis and Prevention, v. 25, n. 1, pp. 110–139, 2025. doi: https://doi.org/10.1007/s11668-025-02118-6.
    » https://doi.org/10.1007/s11668-025-02118-6
  • [36] VIJAYAKUMAR, S., “Optimization of friction stir welding parameters for dissimilar aluminium alloys using RSM-GRA and RSM-TOPSIS: towards sustainable manufacturing in industry 4.0”, Results in Engineering, v. 27, pp. 107054, 2025. doi: https://doi.org/10.1016/j.rineng.2025.107054.
    » https://doi.org/10.1016/j.rineng.2025.107054
  • [37] SUN, S., YU, P., XING, X., et al, “Multi-objective collaborative optimization of active distribution network operation based on improved particle swarm optimization algorithm”, Scientific Reports, v. 15, n. 1, pp. 8999, 2025. doi: https://doi.org/10.1038/s41598-025-90907-2. PubMed PMID: 40089551.
    » https://doi.org/10.1038/s41598-025-90907-2
  • [38] ABDULLAH, I., MEJBEL, M., AL-BHADLE, B., “Double stage friction stir spot extrusion welding: a novel manufacturing technique for joining sheets”, Experimental Techniques, v. 48, n. 2, pp. 323–342, 2024. doi: https://doi.org/10.1007/s40799-023-00660-2.
    » https://doi.org/10.1007/s40799-023-00660-2

Publication Dates

  • Publication in this collection
    28 Aug 2026
  • Date of issue
    2026

History

  • Received
    21 May 2026
  • Accepted
    23 July 2026
location_on
Laboratório de Hidrogênio, Coppe - Universidade Federal do Rio de Janeiro, em cooperação com a Associação Brasileira do Hidrogênio, ABH2 Av. Moniz Aragão, 207, 21941-594, Rio de Janeiro, RJ, Brasil, Tel: +55 (21) 3938-8791 - Rio de Janeiro - RJ - Brazil
E-mail: revmateria@gmail.com
rss_feed Stay informed of issues for this journal through your RSS reader
Go to top Report error