作者
Alexis Piedrafita,Marta Sablik,Evgenia Preka,Valentin Goutaudier,Zeynep Demir,Fariza Mezine,Gillian Divard,Jessy Dagobert,Blaise Robin,Agathe Truchot,Thibaut Thalamas,Aurélie Sannier,Marion Rabant,Olivier Aubert,Mehdi Maanaoui,Moglie Le Quintrec,Lionel Couzi,Bertrand Chauveau,Oriol Bestard,Michelle Elias
摘要
KEY POINTS: Normalization methods affected the performance of gene-based diagnostic classifiers for kidney transplant rejection across a large multicenter cohort. Complex normalization reduced classifier performance by overcorrecting signal, while simpler methods preserved rejection signature and key robustness. By identifying optimal normalization methods, this work advances a standardized preprocessing framework for molecular diagnostics in transplantation. BACKGROUND: The Banff 2022 classification endorses intragraft gene expression profiling using the Banff Human Organ Transplant (B-HOT) consensus gene panel for rejection diagnosis. However, lack of standardized analytical pipelines, including data normalization, limits clinical implementation, with its effect on diagnostic performance yet to be determined. METHODS: We evaluated ten normalization methods in 868 kidney allograft biopsies from nine European and North American centers, all Banff-graded and B-HOT-profiled on nCounter, comprising derivation ( n =441), internal ( n =186), and external ( n =241) validation cohorts. Each method was assessed through its downstream impact on ( 1 ) gene count stability, ( 2 ) differential expression and cross-platform concordance with RNA sequencing (RNA-seq) data, and ( 3 ) discrimination and calibration of predictive models for antibody-mediated rejection (AMR) and T -cell-mediated rejection (TCMR). RESULTS: Most methods improved count stability and showed high concordance with RNA-seq for overall gene expression. They also produced robust differential expression signatures consistent with those detected by RNA-seq, except for RUVSeq and RCRNorm , which identified fewer differentially expressed genes and showed lower concordance. In the overall validation cohort ( n =427), diagnostic performance was consistently high across nSolver -based approaches, nanostringr , NanoStringDiff , MetaNorm , and RCRNormFast (AMR, area under the ROC curve [AUROC], 0.88-0.91; area under precision-recall curves [AUPRC], 0.86-0.89; TCMR AUROC, 0.90-0.92; AUPRC, 0.78-0.83). Performance declined with RCRNorm (AMR AUROC/AUPRC, 0.55/0.41; TCMR, 0.53/0.18) and, for TCMR, with RUVSeq (AUROC, 0.84-0.85; AUPRC, 0.64-0.65). Calibration was satisfactory for most methods, except for RCRNorm and for TCMR models after RUVSeq . CONCLUSIONS: Normalization choice significantly impacted gene expression profiles and diagnostic classifier performance. Most methods, including nSolver-based pipelines, achieved robust discrimination for both AMR and TCMR. Complex methods, including RCRNorm, and RUVSeq for TCMR, reduced performance, with simpler approaches consistently outperforming them for B-HOT-based molecular diagnostics.