Building NGS, AI and network methods, and the public web servers & databases that make them usable.
I develop pipelines for sequencing and multi-omics analysis, sequence-based machine learning models, and systems-level host-pathogen networks, and release them as web servers and databases. Applications span viral pathogens, crops, livestock, and human disease.

Dr. Naveen Duhan
Computational Biology & Genomics
naveen.duhan@outlook.com
From Sequencing Data to Predictive Models & Public Resources
NGS & Multi-Omics Analytics
Pipelines that take raw sequencing reads to biological insight: RNA-Seq, small RNA, metagenomics, amplicon sequencing and variant calling. Applied to viral surveillance (iSNVs, quasispecies), crop stress transcriptomics, livestock markers, and integrated multi-omics.
Artificial Intelligence & Machine Learning
Alignment-free deep learning and protein language models that infer function, localization and interaction properties directly from sequence: enzymes, resistance genes, transporters, subcellular localization, and binding affinity (ΔΔG), benchmarked with homology-aware splits.
Systems Biology & Host–Pathogen Networks
Protein-protein interactomes within and between species, integrated with expression, pathway and regulatory data, to explain susceptibility, virulence and resistance in human, animal and plant systems, including host co-factors and innate immune signaling.
Web Servers, Databases & Reproducible Software
Cost-effective prediction servers, curated databases (microsatellite markers, host-pathogen interactomes, surveillance trackers) and containerized Nextflow DSL2 workflows, used by more than 11,000 researchers worldwide.
From sequence-level classification to predictive interactomes and public resources.
I have built an independent methodological trajectory spanning from sequence-level classification to proteome-scale host–pathogen interactomes. To resolve functional properties directly from sequence, I developed SNVguru for variant analysis (adopted by 50+ research groups worldwide). As first author of deepNEC, I conceived and developed an alignment-free architecture achieving >95% accuracy in classifying metabolic enzymes, establishing the deep learning foundation to predict viral variant effects in host cellular contexts.
I next expanded this foundation to transcriptomics and interactomics. To integrate expression dynamics with interactome networks, I developed pySeqRNA, and as first author of HuCoPIA, conceived a coronavirus interactome atlas spanning viral families. For deepHPI, I designed feature extraction and modeling pipelines achieving AUROC > 0.90, part of 13 deployed software packages and web resources accessed by researchers worldwide.
In recent computational genomics research, I translated these capabilities into high-throughput pathogen surveillance and outbreak analytics. To track rapidly evolving viruses during outbreaks, I led computational genomics for targeted amplicon sequencing studies of emerging avian metapneumovirus (AMPV) genomes as first and co-corresponding author, and co-authored studies on viral genomic diversity.
Education, Academic Trajectory & Honors
Education & Academic Timeline
Ph.D. in Plant Sciences (Bioinformatics and Computational Biology)
M.Sc. in Bioinformatics
Add-on Certificate Diploma in Bioinformatics
B.Sc. in Biology
Honors & Awards
Doctoral Student Researcher of the Year
College of Agriculture and Applied Sciences (CAAS), Utah State University
Awarded by the College of Agriculture and Applied Sciences for pioneering computational research in high-throughput genomics, deep learning enzyme prediction, and viral interactomics.
Young Scientist Award
National Conference on Technological Challenges (TECHSEAR-2017)
Conferred at ICAR-Indian Institute of Rice Research (IIRR), Hyderabad, India, in recognition of significant contributions to crop bioinformatics and molecular marker discovery.
Media & Institutional Press
Methodological Deep Dives & Computational Insights
Demystifying Intra-Host Single Nucleotide Variants (iSNVs) in Emerging Viral Surveillance
Why consensus genomes miss early transmission dynamics, and how to calibrate technical error thresholds to accurately detect low-frequency viral quasispecies.
Architecting Production Metagenomics Pipelines with Nextflow DSL2 and Singularity
Building scalable, deterministic bioinformatics workflows that survive high-throughput diagnostic loads on institutional HPC clusters.
Protein Language Models vs. Alignment-Free Classifiers in Enzyme Commission (EC) Prediction
Comparing deep contextual sequence embeddings (ESM-2, ProtTrans) against k-mer representations for enzymatic reaction classification below the twilight zone.
Teaching, Coursework & Computational Mentoring
Bioinformatics and Big Data Mining
Advanced computational methods for biological big data: high-throughput sequencing analysis, sequence alignment algorithms, structural modeling, machine learning in genomics, and high-performance computing cluster utilization.
Bioinformatics Tools and Their Application in Agriculture
Graduate curriculum covering biological databases, molecular phylogenetics, pairwise and multiple sequence alignment, protein secondary structure prediction, and applied crop genomics.
Advances in Bioinformatics
Doctoral curriculum on computational frontiers: comparative genomics, transcriptomic differential expression modeling, protein-protein interaction networks, and structural docking algorithms.
Terminal-First Learning & Rigorous Trainee Ownership
Bridging abstract biological concepts with hands-on computational execution. Equipping life scientists with command-line proficiency, reproducible workflows, and rigorous algorithmic reasoning as core competencies for modern genomics.