Abstract
HIV-1 exhibits tropism for the CCR5 and/or CXCR4 receptors, which are essential for viral entry into CD4 cells. With the advent of coreceptor antagonists, understanding HIV-1 tropism has become crucial for patient management. The V3 loop, a highly variable region of approximately 35 amino acids, plays a key role in determining tropism. Due to the genetic diversity of HIV-1 subtypes, V3 loop sequences may vary in amino acid composition depending on tropism and subtype. This study aimed to identify critical positions within the V3 loop for defining CCR5, CXCR4, and R5X4 tropisms based on HIV-1 subtype. The random forest algorithm was employed to assess variable importance. CCR5 tropism was predominant, accounting for 80.06% of the sample, while 15% exhibited dual-tropism (R5X4) and 6.37% had CXCR4 tropism. The most prevalent subtype was B, representing 54.45% of the sample. Random forest analysis identified positions 2, 5, 8, 11, 12, 13, 18, 20, 21, 22, 24, 25, 32, and 34 as crucial for HIV-1 tropism. Most of these positions aligned with previous findings. Interestingly, the widely cited 11/25 rule was not universally applicable across subtypes, as recombinant forms exhibited different key positions. Position 25 was important across all subtypes, whereas position 11 ranked among the top five only in subtype B. The random forest algorithm effectively identified key positions for tropism characterization.
Keywords:
HIV-1 tropism; V3 loop; Random Forest algorithm; CCR5 and CXCR4 receptors; HIV-1 subtypes.
HIGHLIGHTS
Random forest efficiently identified key tropism positions in the HIV-1 V3 region.
Positions 11, 24, 25, and 32 emerged as vital for tropism determination.
Position 5 is important for tropism determination, particularly in non-B subtypes.
INTRODUCTION
HIV is a significant global threat, with more than 38 million individuals affected worldwide and 650,000 deaths in 2021 [1]. The virus requires a co-receptor (either CCR5 or CXCR4) along with the CD4 receptor to infiltrate target cells. The viral envelope protein gp120 attaches to the CCR5 or CXCR4 co-receptor after binding to CD4, allowing the virus to enter the cytoplasm with the help of gp41 and triggering fusion of the viral capsule with the cell membrane [2-5].
R5-tropic viruses infect cells that express the CCR5 co-receptor and are typically the first type transmitted during HIV infection. X4-tropic viruses infect cells that express the CXCR4 co-receptor and are usually detected in later stages of the disease. Dual/mixed-tropic viruses can infect cells expressing either CCR5 or CXCR4 co-receptors. These viruses are a mixture of R5-tropic and X4-tropic variants and are associated with an intermediate rate of disease progression [6].
To block viral entry into cells, scientists have been developing molecules that can antagonize CCR5 or CXCR4 [7]. The first approved CCR5 antagonist was Maraviroc, which is effective only against HIV-1 variants with exclusive tropism for this co-receptor. Therefore, determining HIV-1 tropism in infected individuals is crucial for clinical decision-making regarding drug administration.
The molecular structure of the V3 loop of gp120 primarily determines HIV-1 tropism through co-receptors, where a single amino acid alteration in the V3 loop can change HIV-1 tropism [8]. Two methods have been developed to assess viral tropism: (i) in vitro phenotypic tests that rely on cells and (ii) in silico methods based on viral genotypic information. In silico prediction methods provide a low-cost and faster alternative due to computational advancements [9].
Most genotypic prediction models rely on the viral V3 loop sequence. The earliest approach to tropism prediction was the 11/25 rule, which states that a virus has X4 tropism if the amino acids at positions 11 or 25 are positively charged (arginine or lysine) [10]. More sophisticated models have since been developed using machine learning techniques, such as gradient boosting [11], support vector machines [12], random forests [13], and artificial neural networks [14].
The accuracy of computational models in predicting HIV-1 tropism depends heavily on the database used to construct the model. Despite subtype C's global prevalence, most models and studies have relied on V3 sequences obtained from subtype B strains, which dominate North America and Europe. As a result, the applicability of these models to non-B strains remains uncertain [15].
Given the vast diversity of HIV-1, understanding variations in the V3 region across different subtypes and identifying key characteristics that distinguish tropism among subtypes is crucial. Therefore, it is imperative to consider the variability in the V3 region across diverse subtypes to improve the reliability of computational models. This study aimed to identify the key positions in the V3 region of gp120 in defining the three types of tropism based on the HIV-1 subtype.
MATERIAL AND METHODS
Dataset
The data used in this study were obtained from the preprocessing developed by Soares and Raposo [16]. They consist of HIV-1 sequences related to the V3 region of the viral gp120, obtained from the Los Alamos database [17]. These sequences contain 35 amino acids, initiated and terminated by cysteine (C), as this characterizes the V3 region, according to Chiou and co-authors [18]. Codons that did not correspond to any amino acid, generated after the sequence alignment step, were represented by the 'X' character. An example of a V3 sequence is CTRPNNNTRKSIHIGPGRAFYTTGEIIGDIRQAHC. In addition to the 35 V3 positions, each observation in the database included tropism classification according to the virus subtype and phenotype. A total of 6,630 unique V3 sequences were analyzed in the study.
Exploratory data analysis
Descriptive statistics of the V3 sequence positions, tropisms, and subtypes were analyzed using frequency (%) analysis and sequence logos, a graphical representation of nucleotide or amino acid conservation, commonly used in bioinformatics research. In this study, each position displayed the distribution of identified amino acids in the alignment, with letter height indicating frequency - the taller the letter, the higher its frequency.
Identification of the most important positions - Random Forest
To identify the most important positions in defining tropism for each subtype of HIV-1, the random forest algorithm was used. This algorithm combines the bagging of decision trees with the randomization of predictor variables. The bagging method aims to improve model performance by combining multiple models built from the same machine learning algorithm but using training sets obtained through bootstrap sampling. Bootstrap sampling involves resampling a dataset with replacement to generate new datasets of the same size as the original.
Decision trees are simple machine learning algorithms based on "if-then" rules, one of which is the CART (Classification and Regression Trees) algorithm, proposed by Breiman and colleagues [19] and used in the random forest. The CART algorithm uses the Gini index as a measure of purity. In classification problems, such as the present study, the Gini index measures the frequency with which a randomly selected observation from the set would be misclassified if it were randomly labeled according to the class label distribution of the dataset.
One of the major advantages of random forest is its ability to provide measures for evaluating the importance of individual variables for the problem. By using variable importance measures, it is possible to identify the most important predictors and reduce the large set of variables to those containing the most relevant information for the problem. This approach allows random forest to be used not only as a predictive model but also as a strategy for identifying relevant features or performing variable selection.
The simplest technique for calculating variable importance is counting how many times the variable appears in the set of decision trees. Variables with the highest number of appearances are considered the most important for the problem. Two other popular strategies are the mean decrease in impurity (measured by the Gini index in classification problems) and the mean decrease in accuracy. In the first measure, the average amount that a variable's partition reduces node impurities across all trees is calculated. In the second measure, the out-of-bag (OOB) error is used, which is calculated from the set of observations not used in constructing each tree. These samples are particularly useful for estimating generalization errors, acting as an internal validation dataset in parameter tuning and variable importance evaluation. To measure the importance of a variable after training, the values of the jth variable are permuted (maintaining the original distribution) in the OOB sets, and the OOB error is computed again on this perturbed dataset. The jth variable's importance score is calculated as the average difference between the OOB error before and after permutation across all trees. The score is normalized by the standard deviation of these differences. Thus, for unimportant variables, permutation should have little or no effect on the model's accuracy, while permutation of important variables should significantly decrease it.
In this study, random forests were constructed with 1,000 trees, and five variables were used (the floor of the square root of 35 positions), as suggested by Breiman [20]. Variable importance was measured using mean decrease in accuracy (MDA), mean decrease in Gini index (MDG), and the number of times the variable was used in the random forest (number of nodes, NN). The data were not split into training and testing sets because the study's objective was not to develop a predictive model but to identify and detail the five most important variables in defining the tropism of each HIV-1 subtype.
All data analysis was performed in R software version 4.0.3 (R Development Core Team, 2021).
RESULTS
CCR5 tropism was the most prevalent, representing 80.06% of the sample. Dual-tropic sequences (R5X4) accounted for approximately 14% of the sequences, and 6.37% exhibited CXCR4 tropism. Subtype B was the most frequent, accounting for 54.45% of the sample. Subtypes C and CRF01_AE were the second (18.28%) and third (8.69%) most frequent, respectively. The remaining subtypes each accounted for less than 5% of the total.
Table 1 presents the tropism distribution according to subtype. It can be observed that, except for CRF02_AG, CCR5 tropism was the most frequent in the sample obtained from the Los Alamos database. Due to the limited number of F and G subtype sequences, these were excluded from the analysis.
Circulating recombinant form CRF01_AE
Positions 5, 8, 12, 25, and 32 were identified as the most important for predicting tropism in CRF01_AE. Position 12 had the highest MDA and appeared most frequently in the trees. Position 8 showed the highest mean decrease in the Gini index (MDG). The estimated OOB error was 8.18%. In the five selected positions, Serine (S) was predominant at position 5 for CCR5 (91.3%) and R5X4 (73.3%) tropisms, while Phenylalanine (F) was the most frequent for CXCR4 tropism (37.5%). Threonine (T) was found in CCR5 tropism (97.2%), while a mixture of Threonine (T) and Isoleucine (I) was observed in CXCR4 and R5X4 tropisms at position 8. Position 12 displayed different mixtures, with Isoleucine (I) predominating in CCR5 tropism (66.8%), and Valine (V) and Phenylalanine (F) being present in CXCR4 and R5X4 tropisms. Position 25 predominantly had Aspartic acid (D) in all three tropisms, while position 32 had Lysine (K) and a mixture of Glutamine (Q) and Arginine (R) in CCR5 and CXCR4 tropisms, respectively (Figure 1).
The polarities of amino acids at each of the 35 positions in the V3 sequences of the circulating recombinant form 01_AE.
Circulating recombinant form 02_AG
In the CRF02_AG, positions 2, 5, 25, 32, and 34 were identified as the most relevant for tropism. Position 5 showed the highest DMA and DMG. The estimated OOB error of the model was 16.19%.
The amino acid Threonine (T) was the most frequent at position 2 for CCR5 (58.7%) and R5X4 (76.6%) tropisms. At position 5, Asparagine (N) and Glycine (G) were most prevalent for CCR5, while Glycine (G) was dominant for R5X4 (62.4%). Position 25 showed a prevalence of Aspartic acid (D) and Glutamic acid (E) for CCR5 (50.5%) and R5X4 (44.8%), respectively. Position 32 exhibited a mixture of amino acids, with Glutamine (Q) predominating for both tropisms. Position 34 was predominantly Histidine (H) for both tropisms (Figure 2).
The polarities of amino acids at each of the 35 positions in the V3 sequences of the circulating recombinant form 02_AG.
Subtype A
Subtype A tropism was analyzed, identifying the five most significant positions for tropism as 5, 12, 22, 24, and 25. Position 25 had the highest frequency, while position 5 had the highest DMA and position 24 had the highest DMG. The OOB error was estimated to be 6.06%. CXCR4 tropism was not evaluated due to the low number of sequences.
The prevalence of amino acids at the five selected positions varied between CCR5 and R5X4 tropisms. Position 5 had the highest frequency of Asparagine (N) in CCR5 (76.0%) and R5X4 (55.9%) tropisms, followed by Glycine (G) (10.5% and 35.3%, respectively). Position 12 had the hydrophobic amino acid Valine (V) as the most frequent in CCR5 tropism (66.3%) and Isoleucine (I) in R5X4 tropism (47.1%). At position 22, CCR5 tropism predominantly had Alanine (A) (86.0%), while R5X4 presented A in 44.1% and Threonine (T) in 41.2% of the sequences. Position 24 had Glycine (G) in 83.3% of CCR5 sequences, while R5X4 presented a mixture with more than four amino acids. Position 25 had Aspartic acid (D) in 54.3% of CCR5 tropism sequences and in 44.1% of R5X4 sequences (Figure 3).
The polarities of amino acids at each of the 35 positions in the V3 sequences of the subtype A.
Subtype B
The top five positions associated with subtype B tropism, as identified by random forest analysis, were 11, 13, 20, 24, and 25. Position 25 was the most frequently identified. Position 11 had the highest DMA and DMG values, and the estimated OOB error was 6.12%. Among the five positions selected, position 11 was predominantly associated with the amino acid Serine (S) for CCR5 tropism (75.71%), with mixtures of Arginine (R) and Glycine (G) for CXCR4 and R5X4 tropisms. Position 13 exhibited Histidine (H) as the prominent residue for all three tropisms (60.31% for CCR5; 26.5% for CXCR4, and 47.7% for R5X4). At position 20, the hydrophobic amino acid Phenylalanine (F) was the most common residue for all three tropisms (77.7% for CCR5, 56.5% for CXCR4, and 51.4% for R5X4), with a mixture of the hydrophobic amino acid Valine (V) in CXCR4 tropism and the hydrophobic amino acid Tryptophan (W) in R5X4 tropism. Position 24 was dominated by Glycine (G) for CCR5 tropism (88.36%), while different amino acid mixtures were observed for CXCR4 and R5X4 tropisms. Position 25 showed several amino acid mixtures for all three tropisms, with Glutamic Acid (E) and Glutamine (Q) being the most frequently observed amino acids for CCR5/R5X4 and CXCR4 tropisms, respectively (Figure 4).
The polarities of amino acids at each of the 35 positions in the V3 sequences of the subtype B.
Subtype C
The random forest algorithm identified positions 5, 11, 18, 21, and 25 as relevant for subtype C. Among these positions, position 25 was the most frequent. Position 11 showed the highest values for DMA and DMG. The estimated OOB error was 3.98%. For position 5, mixtures of amino acids were found for all three tropisms. Position 11 showed Serine (S) as prevalent for CCR5 tropism (98.7%) and Arginine (R) for CXCR4 tropism (59.3%). Position 18 presented Glutamine (Q) for all three tropisms (98.14% for CCR5; 66.7% for CXCR4, and 58.5% for R5X4), with a relevant presence of Arginine (R) only for CXCR4 and R5X4 tropisms. For position 21, the polar amino acid Tyrosine (Y) was frequent for all three tropisms (81.43% for CCR5, 44.4% for CXCR4, and 86.8% for R5X4), with the occurrence of Aspartic acid (D) only for CXCR4 tropism. Position 25 showed various amino acid mixtures for all three tropisms, with the R5X4 tropism presenting several different amino acids (Figure 5).
The polarities of amino acids at each of the 35 positions in the V3 sequences of the subtype C.
Subtype D
In subtype D, positions 12, 18, 20, 24, and 25 were identified as the most important for tropism based on random forest analysis. Position 24 was the most frequently selected, while position 20 had the highest MDA, and position 12 had the highest MDG. The estimated OOB error was 4.71%.
Specifically, position 12 was dominated by the hydrophobic amino acid Isoleucine (I) for CCR5 tropism (53.8%) and the polar amino acid Threonine (T) for CXCR4 tropism (99.1%), while position 18 had the neutral Glutamine (Q) for both CCR5 and R5X4 tropisms (58.1% and 83.3%, respectively) and Arginine (R) for CXCR4 tropism (78.4%). Position 20 showed mixtures of polar and hydrophobic amino acids, with the hydrophobic amino acid Leucine (L) being more frequent in CCR5 and R5X4 tropisms (41.6% and 64.2%, respectively) and the polar amino acids Tyrosine (Y) and Serine (S) being more prevalent in CXCR4 tropism (41.4% and 36.0%, respectively). Position 24 had diverse mixtures for CCR5 and R5X4 tropisms, with Leucine (L) and Arginine (R) being prominent in CXCR4 tropism. Finally, position 25 showed different amino acids for all three tropisms, with Arginine (R) being prominent in CXCR4 and R5X4 tropisms (Figure 6).
The polarities of amino acids at each of the 35 positions in the V3 sequences of the subtype D
DISCUSSION
This investigation aimed to identify the critical positions in the V3 region of HIV-1 responsible for determining the virus's tropism across different subtypes. The random forest algorithm was utilized to achieve this goal, and the findings were compared with existing biological and computational studies.
The algorithm identified positions 2, 5, 8, 11, 12, 13, 18, 20, 21, 22, 24, 25, 32, and 34 as crucial for HIV-1 tropism. Among these, positions 11, 24, 25, and 32 were particularly significant in determining tropism, especially about the substitution of neutral or acidic amino acids with basic ones. The results revealed that position 11 ranked among the top five most important for subtypes B and C, position 24 was key for subtypes A, B, and D, position 25 was critical across all subtypes, and position 32 was vital for the two evaluated CRFs. When the analysis was extended to the top ten positions, position 11 also appeared for CRF01_AE, position 24 for CRF02_AG and subtype C, and position 32 for subtype B. Notably, all the reported positions appeared within the top ten for subtype B, which is the most extensively studied in developed countries.
An investigation into the V3 loop sequences of HIV-1 in North America and Europe revealed frequent substitutions of basic amino acids at specific positions within the loop [21]. The authors highlighted substitutions at positions 11, 24, 25, and 32, which were responsible for approximately 70% of non-conservative basic substitutions. The charge at positions 11 and 25 was found to correlate with the transition from the R5 to X4 phenotype. These positions together form the "charge rule," which posits that the presence of a positively charged amino acid at positions 11 and 25 determines the virus's tropism. Specifically, if arginine or lysine is present at position 11 or 25, the virus is CXCR4-tropic; otherwise, it is R5-tropic. However, this rule may not apply uniformly to all HIV-1 subtypes, as the significance of positions 11 and 25 varies among subtypes. Further investigations are necessary to establish subtype-specific rules. While this rule represents the simplest approach for tropism classification, our study found that only position 25 was consistently selected among the top five positions for all subtypes and CRFs analyzed. Additionally, when examining the top ten positions, position 11 did not appear for CRF02_AG and subtype A. This outcome suggests that the application of this simplified rule may not be suitable for all subtypes, underscoring the need for more subtype-specific guidelines.
The inclusion of position 24, as demonstrated by Cardozo and coauthors [22], enhances the accuracy of predicting HIV-1 tropism (X4 or R5) when combined with positions 11 and 25, leading to the establishment of the 11/24/25 rule. Their research demonstrated that a positively charged amino acid at positions 11, 24, or 25 indicates X4 tropism, while the absence of these amino acids suggests R5 tropism. In our study, when position 24 was analyzed, the nonpolar amino acid glycine (G) predominantly appeared in R5 tropism, while a mixture of amino acids was observed in X4 tropism. Position 24 ranked among the top five for subtypes B and D, and among the top ten for all subtypes/CRFs analyzed.
Additionally, our study highlighted the significance of position 5 in determining tropism, particularly in non-subtype B viruses. Chen and coauthors [11] also emphasized the importance of position 5, noting its comparable relevance to positions 11 and 25. They found that the presence of the neutral polar amino acid tyrosine (Y) at position 5 is associated with X4 tropism. In our study, this association was observed in CRF01_AE and subtypes A and D. However, position 5 was not selected as a significant predictor in subtype D due to the presence of other tropisms exhibiting the amino acid Y. Chen and coauthors proposed the 5/11/25 rule for tropism prediction and suggested considering the importance of positions 13, 18, 22, and 24.
Structural information and the random forest algorithm were employed by Sander and coauthors [23] to identify key positions in the V3 loop of HIV-1 that determine tropism. Their analysis highlighted positions 3, 7, 11, 13, 18, 20, 22, 24, 25, and 32 as important, while our study confirmed the importance of positions 13, 18, 20, and 22. Subtype-specific analysis revealed that positions 13, 18, 20, and 22 were significant for certain HIV-1 subtypes.
Bioinformatics techniques were used by Zhou and coauthors [24] to examine position 22 of the V3 loop in subtype B, comparing R5 and X4 virus amino acid sequences and secondary structures. They found that position 22, along with positions 11 and 25, played a crucial role in determining tropism for CCR5 and CXCR4. In our study, position 22 ranked among the top five for subtype A and top ten for subtype B. Consensus amino acid frequencies differed slightly between the study by Zhou and colleagues and ours, particularly for position 25 in CCR5 tropism.
Random forests were applied to predict the biological phenotype of HIV-1 by considering 37 random features of the V3 loop (35 positions, net charge, and polarity), with positions 11, 13, 18, and 22 identified as the most significant for predicting co-receptor usage or phenotype. This approach, as described by Xu and coauthors [13], involved the analysis of various subtypes, primarily B and C.
The study by Díez-Fuertes and coauthors [25] introduced a Bayesian network classifier that employed specific nucleotide positions from the env gene to predict HIV-1 co-receptor tropism for subtypes B, C, and non-B/non-C. For subtype B, they selected eight positions from the V3 region, four of which were in agreement with the top ten positions identified by our random forest analysis (positions 11, 19, 20, and 22). For subtype C, they selected five positions, three of which matched our findings (positions 11, 18, and 25). Similar to our study, Díez-Fuertes and coauthors observed a higher proportion of S in R5 viruses and a higher proportion of R in X4 viruses at position 11 for subtypes B and C.
The 11/25 rule is inadequate for determining the tropism of CRF01_AE and CRF02_AG due to the prevalence of uncharged or negatively charged amino acid residues at these positions [26-28]. Our study revealed that position 11 predominantly featured a neutral polar amino acid, while position 25 consistently had an acidic polar amino acid for both CRFs. Although position 11 did not rank among the top five crucial positions for either CRF, it did appear in the top ten for CRF01_AE. Previous research employed a logistic model tree approach to identify significant positions, highlighting positions 5, 7, 11, 13, 14, 18, 19, 27, and 32 [29]. Our findings corroborated positions 5 and 32 as part of the top five positions and included positions 7, 11, 13, 14, and 27 among the top ten positions. Overall, seven out of the nine positions identified aligned with our study.
In 2018, Löechel and coauthors [30] devised a random forest-based tropism prediction model for subtype A viruses. Their study identified two important groups of positions: 10-14 and 22-25. Comparing these findings to our study, we observed agreement regarding positions 10, 12, and 13 in the first group and positions 22-25 in the second group.
In their investigation of HIV-1 subtype D sequences, Raymond and coauthors [31] determined that lysine (K) at position 25 is a polymorphic amino acid in this subtype and does not serve as a determinant for CXCR4 usage, unlike subtype B viruses. In our study, position 25 consistently exhibited K across all three tropisms in subtype D, suggesting a mixture of amino acids at this position.
Our study has several limitations. Firstly, it was conducted on a relatively small sample size, resulting in an insufficient number of sequences available for certain subtypes (F and G) and CXCR4 tropism. As a result, caution should be exercised when interpreting the results. This limitation restricted our ability to perform certain analyses and may have influenced the performance of the random forest algorithm. Additionally, unlike some previous studies that grouped X4 tropism and R5X4 tropism together as non-R5 or solely examined X4 tropism, we separately assessed all three tropisms. Although this approach may have affected the identification of specific positions, most of our findings remained consistent.
Moreover, our study primarily focused on the V3 loop of HIV-1, acknowledging that other regions of the virus may also contribute to tropism. We did not investigate how mutations in the V3 loop could impact the binding ability of HIV-1 to its receptors, thus leaving the effects of these mutations on tropism unclear.
The evaluation of three tropisms may have also affected the performance of the random forest algorithm due to significant class imbalance. Out-of-bag error rates were below 10% for all tropisms except for CRF02_AG, which exhibited an OOB error rate of 16.19%. It is important to emphasize that our study's objective was not to develop a predictive model but rather to utilize the random forest algorithm as a technique for variable selection. We acknowledge that the development of effective and generalizable classifiers requires careful consideration of data processing and hyperparameter tuning.
Furthermore, the successful application of the random forest algorithm in other studies focusing on tropism classification highlights its significance in the field of machine learning.
-
Funding:
This research was funded by Fundação de Amparo à Pesquisa do Estado do Rio de Janeiro - FAPERJ, grant number E-26/010.002582/2019.
Acknowledgments:
The authors acknowledge the support of the Federal University of the State of Rio de Janeiro in conducting this research.
REFERENCES
-
1 World Health Organization. HIV data and statistics [Internet]. World Health Organization; 2022 [cited 2023 May 19]. Available from: https://www.who.int/teams/global-hiv-hepatitis-and-stis-programmes/hiv/strategic-information/hiv-data-and-statistics
» https://www.who.int/teams/global-hiv-hepatitis-and-stis-programmes/hiv/strategic-information/hiv-data-and-statistics -
2 Deng H, Liu R, Ellmeier W, Choe S, Unutmaz D, Burkhart M, et al. Identification of a major co-receptor for primary isolates of HIV-1. Nature. 1996;381(6584):661-6. doi:10.1038/381661a0.
» https://doi.org/10.1038/381661a0. -
3 Dragic T, Litwin V, Allaway GP, Martin SR, Huang Y, Nagashima KA, et al. HIV-1 entry into CD4+ cells is mediated by the chemokine receptor CC-CKR-5. Nature. 1996;381(6584):667-73. doi:10.1038/381667a0.
» https://doi.org/10.1038/381667a0. -
4 Sierra S, Kupfer B, Kaiser R. Basics of the virology of HIV-1 and its replication. J Clin Virol. 2005;34(4):233-44. doi:10.1016/j.jcv.2005.09.004.
» https://doi.org/10.1016/j.jcv.2005.09.004. -
5 Lodowski DT, Palczewski K. Chemokine receptors and other G protein-coupled receptors. Curr Opin HIV AIDS. 2009;4(2):88-95. doi:10.1097/coh.0b013e3283223d8d.
» https://doi.org/10.1097/coh.0b013e3283223d8d. -
6 Nankya IL, Tebit DM, Abraha A, Kyeyune F, Gibson R, Jegede O, et al. Defining the fitness of HIV-1 isolates with dual/mixed co-receptor usage. AIDS Res Ther. 2015;12(1). doi:10.1186/s12981-015-0066-7.
» https://doi.org/10.1186/s12981-015-0066-7. -
7 Seibert C, Sakmar T. Small-molecule antagonists of CCR5 and CXCR4: A promising new class of Anti-HIV-1 Drugs. Curr Pharm Des. 2004;10(17):2041-62. doi:10.2174/1381612043384312.
» https://doi.org/10.2174/1381612043384312. -
8 De Jong JJ, De Ronde A, Keulen W, Tersmette M, Goudsmit J. Minimal requirements for the human immunodeficiency virus type 1 V3 domain to support the syncytium-inducing phenotype: Analysis by single amino acid substitution. J Virol. 1992;66(11):6777-80. doi:10.1128/jvi.66.11.6777-6780.1992.
» https://doi.org/10.1128/jvi.66.11.6777-6780.1992. -
9 Raymond S, Delobel P, Izopet J. Phenotyping methods for determining HIV tropism and applications in clinical settings. Curr Opin HIV AIDS. 2012;7(5):463-9. doi:10.1097/coh.0b013e328356f6d7.
» https://doi.org/10.1097/coh.0b013e328356f6d7. -
10 Shioda T, Levy JA, Cheng-Mayer C. Small amino acid changes in the V3 hypervariable region of gp120 can affect the T-cell-line and macrophage tropism of human immunodeficiency virus type 1. Proc Natl Acad Sci U S A. 1992;89(20):9434-8. doi:10.1073/pnas.89.20.9434.
» https://doi.org/10.1073/pnas.89.20.9434. -
11 Chen X, Wang ZX, Pan XM. HIV-1 tropism prediction by the xgboost and HMM methods. Sci Rep. 2019;9(1). doi:10.1038/s41598-019-46420-4.
» https://doi.org/10.1038/s41598-019-46420-4. -
12 Pillai S, Good B, Richman D, Corbeil J. A new perspective on V3 phenotype prediction. AIDS Res Hum Retroviruses. 2003;19(2):145-9. doi:10.1089/088922203762688658.
» https://doi.org/10.1089/088922203762688658. - 13 Xu S, Huang X, Xu H, Zhang C. Improved prediction of coreceptor usage and phenotype of HIV-1 based on combined features of V3 loop sequence using random forest. J Microbiol. 2007;45(5):441-446.
-
14 Resch W, Hoffman N, Swanstrom R. Improved success of phenotype prediction of the human immunodeficiency virus type 1 from Envelope Variable Loop 3 sequence using neural networks. Virology. 2001;288(1):51-62. doi:10.1006/viro.2001.1087.
» https://doi.org/10.1006/viro.2001.1087. - 15 Riemenschneider M, Cashin KY, Budeus B, Sierra S, Elham Shirvani-Dastgerdi, Saeed Bayanolhagh, et al. Genotypic Prediction of Co-receptor Tropism of HIV-1 Subtypes A and C. Sci Rep. 2016;6:1-9.
- 16 Soares RC, Raposo LM. Desempenho de Ferramentas genotípicas e stacking na predição de tropismo do subtipo C do HIV-1 [Performance of Genotypic Tools and Stacking in Predicting Tropism of HIV-1 Subtype C]. In: Concurso de Trabalhos de Iniciação Científica - Simpósio Brasileiro de Computação Aplicada à Saúde (SBCAS), 20; 2020; Sep 15-18; Online. Porto Alegre: Sociedade Brasileira de Computação; c2020. p. 99-104.
-
17 Los Alamos National Laboratory. HIV databases [Internet]. Los Alamos National Laboratory; 2020 [cited 2020 May 19]. Available from: https://www.hiv.lanl.gov/content/index
» https://www.hiv.lanl.gov/content/index -
18 Chiou SH, Freed EO, Panganiban AT, Kenealy WR. Studies on the role of the V3 loop in human immunodeficiency virus type 1 envelope glycoprotein function. AIDS Res Hum Retroviruses. 1992;8(9):1611-8. doi:10.1089/aid.1992.8.1611.
» https://doi.org/10.1089/aid.1992.8.1611. - 19 Breiman L. Classification and regression trees. 1st ed. Boca Raton: Chapman and Hall/CRC; 1984. 368 p.
-
20 Breiman L. Bagging predictors. Mach Learn. 1996;24(2):123-40. doi:10.1007/bf00058655.
» https://doi.org/10.1007/bf00058655. -
21 Milich L, Margolin B, Swanstrom R. V3 loop of the human immunodeficiency virus type 1 env protein: Interpreting sequence variability. J Virol. 1993;67(9):5623-34. doi:10.1128/jvi.67.9.5623-5634.1993.
» https://doi.org/10.1128/jvi.67.9.5623-5634.1993. -
22 Cardozo T, Kimura T, Philpott S, Weiser B, Burger H, Zolla-Pazner S. Structural basis for coreceptor selectivity by the HIV type 1 V3 loop. AIDS Res Hum Retroviruses. 2007;23(3):415-26. doi:10.1089/aid.2006.0130.
» https://doi.org/10.1089/aid.2006.0130. -
23 Sander O, Sing T, Sommer I, Low A, Peter, P. Richard Harrigan, et al. Structural descriptors of gp120 V3 loop for the prediction of HIV-1 coreceptor usage. PLoS Comput Biol. 2007;3(3). doi:10.1371/journal.pcbi.0030058.
» https://doi.org/10.1371/journal.pcbi.0030058. -
24 Zhou HZ, Xu HF, Xin XM, Guan XR, Zhou J. Position 22 of the V3 loop is associated with co-receptor usage and disease progression in HIV-1 subtype B isolates. Curr HIV Res. 2011;9(8):636-41. doi:10.2174/157016211798998727.
» https://doi.org/10.2174/157016211798998727. -
25 Díez-Fuertes F, Delgado E, Vega Y, Fernández-García A, Cuevas MT, Pinilla M, et al. Improvement of HIV-1 coreceptor tropism prediction by employing selected nucleotide positions of the env gene in a Bayesian network classifier. J Antimicrob Chemother. 2013;68(7):1471-85. doi:10.1093/jac/dkt077.
» https://doi.org/10.1093/jac/dkt077. -
26 Esbjörnsson J, Månsson F, Martínez-Arias W, Vincic E, Biague AJ, da Silva ZJ, et al. Frequent CXCR4 tropism of HIV-1 subtype A and CRF02_AG during late-stage disease - indication of an evolving epidemic in West Africa. Retrovirology. 2010;7(1). doi:10.1186/1742-4690-7-23.
» https://doi.org/10.1186/1742-4690-7-23. -
27 Li QH, Shao B, Li J, Wang JY, Song B, Lin YL, et al. Critical amino acid residues and potential N-linked glycosylation sites contribute to circulating recombinant form 01_AE pathogenesis in Northeast China. AIDS. 2019;33(9):1431-9. doi:10.1097/qad.0000000000002197.
» https://doi.org/10.1097/qad.0000000000002197. -
28 Soulie C, Morand-Joubert L, Cottalorda J, Charpentier C, Bellecave P, Le Guen L, et al. Performance of genotypic algorithms for predicting tropism for HIV-1 CRF01_AE recombinant. J Clin Virol. 2018;99-100:57-60. doi:10.1016/j.jcv.2017.12.014.
» https://doi.org/10.1016/j.jcv.2017.12.014. -
29 Shoombuatong W, Hongjaisee S, Barin F, Chaijaruwanich J, Samleerat T. HIV-1 CRF01_AE coreceptor usage prediction using kernel methods based logistic model trees. Comput Biol Med. 2012;42(9):885-9. doi:10.1016/j.compbiomed.2012.06.011.
» https://doi.org/10.1016/j.compbiomed.2012.06.011 -
30 Löchel HF, Riemenschneider M, Frishman D, Heider D. Scotch: Subtype a coreceptor tropism classification in HIV-1. Bioinformatics. 2018;34(15):2575-80. doi:10.1093/bioinformatics/bty170.
» https://doi.org/10.1093/bioinformatics/bty170. -
31 Raymond S, Delobel P, Chaix ML, Cazabat M, Encinas S, Bruel P, et al. Genotypic prediction of HIV-1 subtype D tropism. Retrovirology. 2011;8(1). doi:10.1186/1742-4690-8-56.
» https://doi.org/10.1186/1742-4690-8-56.
-
Editor-in-Chief: Paulo Vitor Farago
-
Associate Editor: Jaiesa Zych Nadolny












