To investigate the dispersion, distance, and evolutionary relationships among the TSPY gene sequences under study, their bi-plots were analyzed based on principal component analysis (PCA), and the first 3 components were depicted by software DARwin6 0.12 (
61). The share of the first 3 components was 82.95%, 3.89%, and 2.03%, respectively, and in total, 88.87% of the total data. These values indicate that the first 3 components have been able to accurately calculate a high percentage of variation. According to the results, the TSPY genes sequence was divided into 4 groups, and also it was largely indicative of cluster analysis.
After the alignment of TSPY gene sequences, major ORF and, then, the secondary structure of this ORF were predicted (
Figure 2A and
Figure 3B). The Ramachandran diagram (
62) for the second structure of the TSPY gene shows that this structure is good stroke chemistry and 91.3% of the remaining groups are in the red zone, which has the highest acceptability. It has several amino acids (324), with formula; C
1619H
2550N
472O
493S
14, a molecular weight of 36963.76 D, theoretical pI of 5.56. The amino acid composition was Ala (A), 10.2%, Arg (R) (9.0%), Asn (N) (3.7%), Asp (D) (3.1%), Cys (C) (1.2%), Gln (Q) (5.2%), Glu (E) (12.3%), Gly (G) (5.2%), His (H) (2.5%), Ile (I) (3.1%), Leu (L) (8.6%), Lys (K) (3.7%), Met (M) (3.1%), Phe (F) (3.4%), Pro (P) (5.2%), Ser (S) (6.2%), Thr (T) (2.8%), Trp (W) (1.2%), Tyr (Y) (3.1%), Val (V) (7.1%), Pyl (O) (0.0%), Sec (U) (0.0%), (B) (0.0%), (Z) (0.0%), X) (0.0%), total number of negatively-charged residues (Asp + Glu) (50), total number of positively-charged residues (Arg + Lys) (41), Atomic composition; Carbon (1619), Hydrogen (2550), Nitrogen (472), Oxygen (493), Sulfur (14), Aliphatic index (76.51), and Grand average of hydropathicity (GRAVY) (-0.571). Integral prediction of protein location is shown in
Table 5.