摘要
Objective: The purpose of this study was to screen for common diagnostic genes and identify possible molecular mechanisms underlying CKD combined with AS using bioinformatic methods. Methods: We retrieved the gene expression profiles of CKD and AS from the Gene Expression Omnibus (GEO). Using differential expression combined with WGCNA, we identified hub genes shared by CKD and AS. GO and KEGG enrichment analyses were conducted. To screen the diagnostic biomarker, we constructed PPI networks and applied three ML methods: least absolute shrinkage and selection operator (ASSO), random forest (RF), and Support Vector Machine-Recursive Feature Elimination (SVM-RFE). The diagnostic performance for the candidate gene was assessed by a receiver operating characteristic curve (ROC) on the training set and the external validation dataset. CIBERSORT is used to evaluate immune cell infiltration the DGIdb database predicts possible drugs for treatment. results: Differentially expressed genes screening The volcano plot revealed that by applying |log(FC)|>0.5 and a p-value 0.75 and p-value < 0.05 were utilized for screening, the GSE66494 dataset resulted in the identification of 1743 DEGs, including 481 genes that were up-regulated and 1262 genes that were down-regulated in terms of [removed]Fig. 2B). The heatmaps illustrated the top 30 genes that were up-regulated and down-regulated in each dataset, respectively, as shown in (Fig. 2C, D). Ultimately, the differential genes from GSE100927 and GSE66494 that met the specified screening criteria were intersected, resulting in Venn diagrams illustrating a total of 57 up-regulated and 73 down-regulated DEGs (Fig. 3A, B). Additionally, we established a dataset named "DEGs," which includes the aforementioned 130 genes that are either up-regulated or down-regulated. Fig.2 Visualization results of differential analyses of GSE100927(AS) and GSE66494(CKD). A The volcano plot of differentially expressed genes for AS. B The volcano plot of differentially expressed genes for CKD. In the volcano plots, up-regulated genes are shown in red, down-regulated genes are shown in blue, and genes with no significant difference are shown in grey. C The heatmap of differentially expressed genes for AS. D The heatmap of differentially expressed genes for CKD. Fig.3 Venn diagrams. A 57 up-regulated genes between CKD and AS. B 73 down-regulated genes between CKD and AS. C 119 overlapped shared genes between the “Module genes” and “DEGs” datasets. Dataset construction of Module genes Using WGCNA analysis, we identified modules with p-values below 0.05 in the GSE100927 (AS) and GSE66494 (CKD) datasets, resulting in two sets named “CKD_module” and “AS_module.” As illustrated in Figures 4A and 4B, the optimal soft power β for AS was determined to be 8, while for CKD, it was set to 18. We utilized the blockwise modules function to cluster the samples, leading to the creation of a cluster dendrogram (Fig. 4C, D) and subsequently developed a heatmap representing the module-trait relationship (Fig. 4E, F). The AS_module comprises three modules of statistical significance: “turquoise,” “brown,” and “blue” (Fig. 4E). The CKD_module includes four significant modules: “midnightblue,” “yellow,” “cyan,” and “brown” (Fig. 4F). This analysis allowed us to pinpoint gene co-expression patterns and uncover potential genes related to AS and CKD. After integrating the modules and eliminating duplicates, we constructed a “Module genes” dataset containing a total of 3765 genes. Ultimately, by intersecting the DEGs and module genes datasets, we identified 119 shared genes (Fig. 3C). Fig.4 Weighted gene co-expression network analysis. A The selection of soft threshold in AS (GSE100927). B The selection of soft threshold in CKD (GSE66494). C The cluster dendrogram of co-expression genes for AS (GSE100927). D The cluster dendrogram of co-expression genes for CKD (GSE66494). E Module–trait relationships for AS (GSE100927). F Module–trait relationships for CKD (GSE66494). Functional enrichment analysis results The findings from the GO and KEGG enrichment analyses indicated that the shared genes participated in biological processes (BPs) such as signal transduction, the negative regulation of angiogenesis, and apoptosis (Figure 5A). Regarding cellular components (CCs), these genes were primarily linked to the cytosol, actin cytoskeleton, and the extracellular region (Fig. 5A). The molecular functions (MFs) predominantly included identical protein binding, activity related to protein homodimerization, and actin filament binding (Fig. 5A). Additionally, the KEGG pathway analysis of the shared genes revealed significant enrichment in aspects such as the cytoskeleton in muscle cells, vascular smooth muscle contraction, the cGMP-PKG signaling pathway, the NF-kappa B signaling pathway, among others (Fig. 5B). Fig.5 GO and KEGG enrichment analysis. A The top 8 GO functions. B The top 12 KEGG pathways. Construction of PPI network and identification of core genes To enhance our understanding of the interactions among the previously identified DEGs, we utilized STRING to construct a PPI network based on these shared DEGs (Fig. 6A). The output was then imported into Cytoscape software, and we employed the CytoHubba plugin to pinpoint the core genes. We applied five distinct algorithms (Degree, DMNC, MCC, EPC, and MNC) to derive the top 10 genes (Fig. 6B-F), with the intersections being considered as core genes, which were illustrated using a Venn Diagram (Fig. 6G). The core genes identified comprised FLNA, MYH11, CNN1, TAGLN, TPM2, and ACTG2. Fig.6 Core genes identification. A PPI network (confidence > 0.4). B The top 10 genes of Degree algorithm. C The top 10 genes of MCC algorithm. D The top 10 genes of EPC algorithm. E The top 10 genes of DMNC algorithm. F The top 10 genes of MNC algorithm. G The overlaps of five algorithms. Machine learning screening results We utilized LASSO to analyze the previously mentioned six core genes within the AS and CKD datasets individually, identifying four crucial genes in the AS dataset: FLNA, TAGLN, TPM2, and ACTG2 (Fig. 7A, B). Meanwhile, in the CKD dataset, we found three significant genes, specifically FLNA, TAGLN, and ACTG2 (Fig. 8A, B). Furthermore, we conducted RF analysis, which led to the identification of six essential genes in AS (Fig. 7C, D) and three principal genes in CKD: FLNA, TAGLN, and ACTG2 (Fig. 8C, D). By consolidating these findings, we ultimately recognized three potential hub genes: FLNA, TAGLN, and ACTG2. Fig.7 Machine learning screening in AS (GSE100927). A, B 4 genes were screened for AS diagnosis using LASSO algorithm. C, D 6 genes were screened for AS diagnosis using RF algorithm. Fig.8 Machine learning screening in CKD (GSE66494). A, B 3 genes were screened for CKD diagnosis using LASSO algorithm. C, D 3 genes were screened for CKD diagnosis using RF algorithm. Diagnosis values evaluation As demonstrated in Figure 9, we generated ROC curves and determined the AUC values for these three central genes to assess their diagnostic efficacy. In the training datasets (GSE100927 and GSE66494), the AUC values for FLNA, TAGLN, and ACTG2 in the AS dataset were 0.771, 0.854, and 0.762, whereas in the CKD dataset, the values were 0.906, 0.950, and 0.868, respectively, reflecting strong diagnostic capabilities (Fig 9A, B). In the validation datasets (GSE43292 and GSE142153), the AUC values for FLNA, TAGLN, and ACTG2 in the AS dataset were recorded as 0.750, 0.821, and 0.766, respectively; for the CKD dataset, the values were 0.814, 0.779, and 0.886 (Fig. 9C, D), suggesting that these three genes possess significant diagnostic potential in external datasets. In conclusion, the results suggest that FLNA, TAGLN, and ACTG2 may serve as valuable diagnostic biomarkers for CKD when accompanied by AS. Fig. 9 ROC graphs. A ROC graphs of FLNA, TAGLN and ACTG2 in GSE100927 dataset. B ROC graphs of FLNA, TAGLN and ACTG2 in GSE66494 dataset. C ROC graphs of FLNA, TAGLN and ACTG2 in GSE43292 dataset. D ROC graphs of FLNA,TAGLN and ACTG2 in GSE142153 dataset. Immune cell infiltration analysis CIBERSORT algorithm was used to analyze the proportion of immune infiltrating cells. Figure 10A-D illustrates the distribution of 22 types of immune cells within AS and CKD, represented through barplots and heatmaps. Additionally, when compared to control samples, the AS samples demonstrated elevated levels of memory B cells, regulatory T cells (Tregs), gamma delta T cells, M0 macrophages, and activated mast cells, whereas they had lower levels of naive B cells, plasma cells, resting memory CD4+ T cells, activated memory CD4+ T cells, monocytes, M1 macrophages, activated dendritic cells, and resting mast cells (Fig. 10E). In the context of CKD, there was an observed increase in CD8+ T cells, monocytes, and M0 macrophages, along with a decrease in resting memory CD4+ T cells, follicular helper T cells, regulatory T cells (Tregs), and activated NK cells (Fig. 10F). Notably, we observed that M0 macrophages and resting memory CD4+ T cells exhibited the same expression pattern across both AS and CKD samples; specifically, M0 macrophages increased while resting memory CD4+ T cells decreased. The correlation heatmap for individual immune cells revealed that in AS samples, there were negative correlations between M0 macrophages and monocytes (r = -0.72), resting memory CD4+ T cells (r = -0.55), activated NK cells (r = -0.64), and plasma cells (r = -0.51). Conversely, naive B cells (r = 0.55) and resting mast cells (r = 0