Welcome to the IKCEST
Journal
IEEE/ACM Transactions on Computational Biology and Bioinformatics

IEEE/ACM Transactions on Computational Biology and Bioinformatics

Archives Papers: 644
IEEE Xplore
Please choose volume & issue:
Guest Editorial Guest Editorial for the 20th Asia Pacific Bioinformatics Conference
Su Datt LamWai Keat YamYi-Ping Phoebe Chen
Keywords:Special issues and sectionsMeetingsBioinformaticsGuest EditorialNeuroimagingAlzheimer’s DiseaseQuantitative Trait LociGene Regulatory NetworksHeat ResistanceChemical InformationComputational BiologyCell BiolGlobal StabilityNon-negative Matrix FactorizationEpigenetic LandscapeMulti-task LearningImprovements In SpeedComputational Systems BiologyCombination Of Machine LearningSpecific Chemical PropertiesImaging Genetics
Abstracts:The four papers in this special section were presented at the 20th Asia Pacific Bioinformatics Conference (APBC), which was held in Malaysia 26-28 April 2022.
Parallel Convolutional Contrastive Learning Method for Enzyme Function Prediction
Xindi YuShusen ZhouMujun ZangQingjun WangChanjuan LiuTong Liu
Keywords:EnzymesContrastive learningProteinsTrainingPredictive modelsComputational modelingFeature extractionBioinformaticscontrastive learningenzyme function predictionparallel convolution
Abstracts:The function labeling of enzymes has a wide range of application value in the medical field, industrial biology and other fields. Scientists define enzyme categories by enzyme commission (EC) numbers. At present, although there are some tools for enzyme function prediction, their effects have not reached the application level. To improve the precision of enzyme function prediction, we propose a parallel convolutional contrastive learning (PCCL) method to predict enzyme functions. First, we use the advanced protein language model ESM-2 to preprocess the protein sequences. Second, PCCL combines convolutional neural networks (CNNs) and contrastive learning to improve the prediction precision of multifunctional enzymes. Contrastive learning can make the model better deal with the problem of class imbalance. Finally, the deep learning framework is mainly composed of three parallel CNNs for fully extracting sample features. we compare PCCL with state-of-art enzyme function prediction methods based on three evaluation metrics. The performance of our model improves on both two test sets. Especially on the smaller test set, PCCL improves the AUC by 2.57%.
MLW-BFECF: A Multi-Weighted Dynamic Cascade Forest Based on Bilinear Feature Extraction for Predicting the Stage of Kidney Renal Clear Cell Carcinoma on Multi-Modal Gene Data
Liye JiaLiancheng JiangJunhong YueFang HaoYongfei WuXilin Liu
Keywords:Feature extractionForestryData modelsCancerPredictive modelsRandom forestsLogic gatesComputational modelingNoiseDeep learningRenal CellRenal CarcinomaDynamic CascadeBilinear FeatureCascade ForestSeries Of ExperimentsCharacterization Of GenesSelection AlgorithmRedundant InformationCopy Number AberrationsBilinear ModelNeural NetworkTraining SetDeep LearningConvolutional Neural NetworkMachine Learning MethodsPerformance MetricsHighest AccuracyGene DatasetDeep FeaturesDeep ForestCascade ModelReLU LayerRed Dashed BoxArea Under CurveMultimodal ModelBox In FigRed Box In FigSubmodulePathological ImagesCascade modelbilinear modelensemble forestsmultimodal gene dataCarcinoma, Renal CellKidney NeoplasmsHumansComputational BiologyAlgorithmsNeoplasm StagingGene Expression ProfilingDatabases, Genetic
Abstracts:The stage prediction of kidney renal clear cell carcinoma (KIRC) is important for the diagnosis, personalized treatment, and prognosis of patients. Many prediction methods have been proposed, but most of them are based on unimodal gene data, and their accuracy is difficult to further improve. Therefore, we propose a novel multi-weighted dynamic cascade forest based on the bilinear feature extraction (MLW-BFECF) model for stage prediction of KIRC using multimodal gene data (RNA-seq, CNA, and methylation). The proposed model utilizes a dynamic cascade framework with shuffle layers to prevent early degradation of the model. In each cascade layer, a voting technique based on three gene selection algorithms is first employed to effectively retain gene features more relevant to KIRC and eliminate redundant information in gene features. Then, two new bilinear models based on the gated attention mechanism are proposed to better extract new intra-modal and inter-modal gene features; Finally, based on the idea of the bagging, a multi-weighted ensemble forest classifiers module is proposed to extract and fuse probabilistic features of the three-modal gene data. A series of experiments demonstrate that the MLW-BFECF model based on the three-modal KIRC dataset achieves the highest prediction performance with an accuracy of 88.9 %.
circ2DGNN: circRNA-Disease Association Prediction via Transformer-Based Graph Neural Network
Keliang CenZheming XingXuan WangYadong WangJunyi Li
Keywords:DiseasesDatabasesSemanticsHeterogeneous networksRNAVectorsProteinsMolecular biophysicsCompoundsPredictive modelsNeural NetworkGraph Neural NetworkscircRNA-disease AssociationsHyperparametersTypes Of RelationshipsFive-fold Cross-validationVector RepresentationHeterogeneous NetworkParameter MatrixAttention ScoresNode RepresentationsLink PredictionHeterogeneous GraphModel PerformanceTraining SetPositive SamplesSequence InformationNegative SamplesAttention MechanismSemantic SimilarityMessage PassingEntity TypesTarget NodeEmbedding VectorsNode EmbeddingsTypes Of EdgesTypes Of NodesNodes In The GraphcircRNA SequencesMedical TermsCircRNA-disease associationheterogeneous graph neural networkmessage aggregationlink predictionRNA, CircularHumansNeural Networks, ComputerComputational BiologyAlgorithmsGraph Neural Networks
Abstracts:Investigating the associations between circRNA and diseases is vital for comprehending the underlying mechanisms of diseases and formulating effective therapies. Computational prediction methods often rely solely on known circRNA-disease data, indirectly incorporating other biomolecules' effects by computing circRNA and disease similarities based on these molecules. However, this approach is limited, as other biomolecules also play significant roles in circRNA-disease interactions. To address this, we construct a comprehensive heterogeneous network incorporating data on human circRNAs, diseases, and other biomolecule interactions to develop a novel computational model, circ2DGNN, which is built upon a heterogeneous graph neural network. circ2DGNN directly takes heterogeneous networks as inputs and obtains the embedded representation of each node for downstream link prediction through graph representation learning. circ2DGNN employs a Transformer-like architecture, which can compute heterogeneous attention score for each edge, and perform message propagation and aggregation, using a residual connection to enhance the representation vector. It uniquely applies the same parameter matrix only to identical meta-relationships, reflecting diverse parameter spaces for different relationship types. After fine-tuning hyperparameters via five-fold cross-validation, evaluation conducted on a test dataset shows circ2DGNN outperforms existing state-of-the-art(SOTA) methods.
Discriminative Domain Adaption Network for Simultaneously Removing Batch Effects and Annotating Cell Types in Single-Cell RNA-Seq
Qi ZhuAizhen LiZheng ZhangChuhang ZhengJunyong ZhaoJin-Xing LiuDaoqiang ZhangWei Shao
Keywords:Feature extractionAnnotationsTrainingAdaptation modelsSemanticsRobustnessData miningSequential analysisCorrelationComputer scienceCell TypesBatch EffectsDomain AdaptationCell Type AnnotationMachine LearningFeature RepresentationFluidicLocal AlignmentTarget DomainSource DomainscRNA-seq AnalysisContrastive LossBatch CorrectionIdentify Cell TypesBatch Effect CorrectionDiscriminative Feature RepresentationSelf-paced LearningSource Domain SamplesCD4 T CellsCell ClustersAlignment ModuleTarget Domain SamplesGlobal AlignmentPseudo LabelsCell Type ClassificationDomain DatasetDistribution Of Cell TypesTypes Of DatasetsTarget DatasetMaximum Mean DiscrepancyBatch effectscell type classificationdomain adaptationself-paced learningsemantic alignmentSingle-Cell AnalysisRNA-SeqAlgorithmsHumansComputational BiologyMachine LearningAnimalsMiceSequence Analysis, RNASingle-Cell Gene Expression Analysis
Abstracts:Machine learning techniques have become increasingly important in analyzing single-cell RNA and identifying cell types, providing valuable insights into cellular development and disease mechanisms. However, the presence of batch effects poses major challenges in scRNA-seq analysis due to data distribution variation across batches. Although several batch effect mitigation algorithms have been proposed, most of them focus only on the correlation of local structure embeddings, ignoring global distribution matching and discriminative feature representation in batch correction. In this paper, we proposed the discriminative domain adaption network (D2AN) for joint batch effects correction and type annotation with single-cell RNA-seq. Specifically, we first captured the global low-dimensional embeddings of samples from the source and target domains by adversarial domain adaption strategy. Second, a contrastive loss is developed to preliminarily align the source domain samples. Moreover, the semantic alignment of class centroids in the source and target domains is achieved for further local alignment. Finally, a self-paced learning mechanism based on inter-domain loss is adopted to gradually select samples with high similarity to the target domain for training, which is used to improve the robustness of the model. Experimental results demonstrated that the proposed method on multiple real datasets outperforms several state-of-the-art methods.
Hierarchical Hypergraph Learning in Association- Weighted Heterogeneous Network for miRNA- Disease Association Identification
Qiao NingYaomiao ZhaoJun GaoChen ChenMinghao Yin
Keywords:DiseasesFeature extractionSemanticsVectorsHeterogeneous networksData miningCorrelationBiomedical measurementAggregatesVelocity measurementAssociative LearningHeterogeneous NetworkHypergraph LearningChanges In Expression LevelsSemantic InformationmiRNA LevelsPair Of NodesmiRNA Expression LevelsSimilar InformationAbundant InformationHierarchical ApproachHeterogeneous GraphmiRNA-disease AssociationsSimilarity ScoreSimilar AssociationsFeed-forward NetworkNodes In The GraphBaseline MethodsNeighboring NodesGraph Convolutional NetworkDisease SimilarityNode EmbeddingsNode RepresentationsNeighbor DistanceGraph Neural NetworksEdge AttributesTypes Of EdgesGraph Attention NetworkKey VectorInformation AggregationmiRNA-disease associationhierarchical hypergraph learningassociation-weighted heterogeneous networkMicroRNAsComputational BiologyHumansAlgorithmsMachine LearningGenetic Predisposition to Disease
Abstracts:MicroRNAs (miRNAs) play a significant role in cell differentiation, biological development as well as the occurrence and growth of diseases. Although many computational methods contribute to predicting the association between miRNAs and diseases, they do not fully explore the attribute information contained in associated edges between miRNAs and diseases. In this study, we propose a new method, Hierarchical Hypergraph learning in Association-Weighted heterogeneous network for MiRNA-Disease association identification (HHAWMD). HHAWMD first adaptively fuses multi-view similarities based on channel attention and distinguishes the relevance of different associated relationships according to changes in expression levels of disease-related miRNAs, miRNA similarity information, and disease similarity information. Then, HHAWMD assigns edge weights and attribute features according to the association level to construct an association-weighted heterogeneous graph. Next, HHAWMD extracts the subgraph of the miRNA-disease node pair from the heterogeneous graph and builds the hyperedge (a kind of virtual edge) between the node pair to generate the hypergraph. Finally, HHAWMD proposes a hierarchical hypergraph learning approach, including node-aware attention and hyperedge-aware attention, which aggregates the abundant semantic information contained in deep and shallow neighborhoods to the hyperedge in the hypergraph. Our experiment results suggest that HHAWMD has better performance and can be used as a powerful tool for miRNA-disease association identification.
An End-to-End Knowledge Graph Fused Graph Neural Network for Accurate Protein-Protein Interactions Prediction
Jie YangYapeng LiGuoyin WangZhong ChenDi Wu
Keywords:ProteinsFeature extractionGraph neural networksBiological system modelingDrugsPredictive modelsDiseasesData modelsBiologyAccuracyNeural NetworkProtein InteractionsPrediction AccuracyGraph Neural NetworksProtein Interaction PredictionCell SignalingCellular ProcessesDrug DevelopmentProtein-protein Interaction NetworkMultilayer PerceptronSemantic FeaturesTopological FeaturesFeature FusionComplex StructureModel PerformanceConvolutional Neural NetworkPositive SamplesFalse Positive RateTypes Of RelationshipsEmbedding RepresentationGraph Neural Network ModelBiological EntitiesArea Under CurveMatthews Correlation CoefficientVariational AutoencoderDeep Convolutional Neural NetworkSpecific FormulaNode RepresentationsNeighboring NodesDeep neural networkgraph neural networkknowledge graphPPI networkprotein-protein interactions predictionNeural Networks, ComputerComputational BiologyProtein Interaction MappingDatabases, ProteinHumansAlgorithmsProteinsProtein Interaction MapsGraph Neural Networks
Abstracts:Protein-protein interactions (PPIs) are essential to understanding cellular mechanisms, signaling networks, disease processes, and drug development, as they represent the physical contacts and functional associations between proteins. Recent advances have witnessed the achievements of artificial intelligence (AI) methods aimed at predicting PPIs. However, these approaches often handle the intricate web of relationships and mechanisms among proteins, drugs, diseases, ribonucleic acid (RNA), and protein structures in a fragmented or superficial manner. This is typically due to the limitations of non-end-to-end learning frameworks, which can lead to sub-optimal feature extraction and fusion, thereby compromising the prediction accuracy. To address these deficiencies, this paper introduces a novel end-to-end learning model, the Knowledge Graph Fused Graph Neural Network (KGF-GNN). This model comprises three integral components: (1) Protein Associated Network (PAN) Construction: We begin by constructing a PAN that extensively captures the diverse relationships and mechanisms linking proteins with drugs, diseases, RNA, and protein structures. (2) Graph Neural Network for Feature Extraction: A Graph Neural Network (GNN) is then employed to distill both topological and semantic features from the PAN, alongside another GNN designed to extract topological features directly from observed PPI networks. (3) Multi-layer Perceptron for Feature Fusion: Finally, a multi-layer perceptron integrates these varied features through end-to-end learning, ensuring that the feature extraction and fusion processes are both comprehensive and optimized for PPI prediction. Extensive experiments conducted on real-world PPI datasets validate the effectiveness of our proposed KGF-GNN approach, which not only achieves high accuracy in predicting PPIs but also significantly surpasses existing state-of-the-art models. This work not only enhances our ability to predict PPIs with a higher precision but also contributes to the broader application of AI in Bioinformatics, offering profound implications for biological research and therapeutic development.
RFLP-Inator: Interactive Web Platform for In Silico Simulation and Complementary Tools of the PCR-RFLP Technique
Kiefer Andre Bedoya BenitesWilser Andrés García-Quispes
Keywords:EnzymesDNADatabasesSoftwareIn vitroSoftware algorithmsEncodingElectrophoresisSymbolsBrowsersComputer simulationopen educational resourcespolymerase chain reaction (PCR)restriction fragment length polymorphism (RFLP)software toolsweb design
Abstracts:Polymerase chain reaction - Restriction Fragment Length Polymorphism (PCR-RFLP) is an established molecular biology technique leveraging DNA sequence variability for organism identification, genetic disease detection, biodiversity analysis, etc. Traditional PCR-RFLP requires wet-laboratory procedures that can result in technical errors, procedural challenges, and financial costs. With the aim of providing an accessible and efficient PCR-RFLP technique complement, we introduce RFLP-inator. This is a comprehensive web-based platform developed in R using the package Shiny, which simulates the PCR-RFLP technique, integrates analysis capabilities, and offers complementary tools for both pre- and post-evaluation of in vitro results. We developed the RFLP-inator's algorithm independently and our platform offers seven dynamic tools: RFLP simulator, Pattern identifier, Enzyme selector, RFLP analyzer, Multiplex PCR, Restriction map maker, and Gel plotter. Moreover, the software includes a restriction pattern database of more than 250,000 sequences of the bacterial 16S rRNA gene. We successfully validated the core tools against published research findings. This new platform is open access and user-friendly, offering a valuable resource for researchers, educators, and students specializing in molecular genetics. RFLP-inator not only streamlines RFLP technique application but also supports pedagogical efforts in genetics, illustrating its utility and reliability.
Orientation Determination of Cryo-EM Projection Images Using Reliable Common Lines and Spherical Embeddings
Xiangwen WangQiaoying JinLi ZouXianghong LinYonggang Lu
Keywords:Three-dimensional displaysReliabilityMolecular biophysicsImage reconstructionBiologyPeriodic structuresImage resolutionElectronsVectorsReconstruction algorithmsAngular reconstructioncryo-electron microscopy (cryo-EM)reliable common linessingle-particle reconstructionspherical embedding
Abstracts:Three-dimensional (3D) reconstruction in single-particle cryo-electron microscopy (cryo-EM) is a critical technique for recovering and studying the fine 3D structure of proteins and other biological macromolecules, where the primary issue is to determine the orientations of projection images with high levels of noise. This paper proposes a method to determine the orientations of cryo-EM projection images using reliable common lines and spherical embeddings. First, the reliability of common lines between projection images is evaluated using a weighted voting algorithm based on an iterative improvement technique and binarized weighting. Then, the reliable common lines are used to calculate the normal vectors and local $X$X-axis vectors of projection images after two spherical embeddings. Finally, the orientations of projection images are determined by aligning the results of the two spherical embeddings using an orthogonal constraint. Experimental results on both synthetic and real cryo-EM projection image datasets demonstrate that the proposed method can achieve higher accuracy in estimating the orientations of projection images and higher resolution in reconstructing preliminary 3D structures than some common line-based methods, indicating that the proposed method is effective in single-particle cryo-EM 3D reconstruction.
A Knowledge Graph-Based Method for Drug-Drug Interaction Prediction With Contrastive Learning
Jian ZhongHaochen ZhaoQichang ZhaoJianxin Wang
Keywords:DrugsFeature extractionContrastive learningKnowledge graphsTrainingSemanticsDatabasesSafetyPredictive modelsKnowledge engineeringContrastive learningdrug–drug interaction predictionknowledge graph
Abstracts:Precisely predicting Drug-Drug Interactions (DDIs) carries the potential to elevate the quality and safety of drug therapies, protecting the well-being of patients, and providing essential guidance and decision support at every stage of the drug development process. In recent years, leveraging large-scale biomedical knowledge graphs has improved DDI prediction performance. However, the feature extraction procedures in these methods are still rough. More refined features may further improve the quality of predictions. To overcome these limitations, we develop a knowledge graph-based method for multi-typed DDI prediction with contrastive learning (KG-CLDDI). In KG-CLDDI, we combine drug knowledge aggregation features from the knowledge graph with drug topological aggregation features from the DDI graph. Additionally, we build a contrastive learning module that uses horizontal reversal and dropout operations to produce high-quality embeddings for drug-drug pairs. The comparison results indicate that KG-CLDDI is superior to state-of-the-art models in both the transductive and inductive settings. Notably, for the inductive setting, KG-CLDDI outperforms the previous best method by 17.49% and 24.97% in terms of AUC and AUPR, respectively. Furthermore, we conduct the ablation analysis and case study to show the effectiveness of KG-CLDDI. These findings illustrate the potential significance of KG-CLDDI in advancing DDI research and its clinical applications.
Hot Journals