Predictive genomics of viral emergence & host response using machine learning & scalable web systems.
I develop computational methods for genomic and transcriptomic analysis, sequence-based machine learning, and molecular interactomes. My work delivers advanced bioinformatics infrastructure, high-throughput surveillance pipelines, and predictive algorithms for emerging viral pathogens and disease outbreaks.

Dr. Naveen Duhan
Computational Biology & Genomics
naveen.duhan@outlook.com
Host–Pathogen Genomics & Predictive Interactomes
Viral Surveillance & Multi-Omics Discovery
Investigates the relationship between viral genomic variation and host transcriptional responses during infection, reservoir maintenance, and cross-species spillover. Integrates targeted amplicon and metagenomic sequencing with host transcriptomics to resolve low-frequency intra-host single nucleotide variants (iSNVs) and quasispecies diversity.
Machine Learning & Context-Aware AI for Interactions
Develops multimodal machine learning architectures and protein language models that combine sequence representations, structural interfaces, and host-receptor orthology. Predicts continuous biophysical binding affinities (ΔΔG) and receptor specificity (including α-2,3 and α-2,6 sialic acids) to prioritize cross-species spillover risk.
Comparative Systems Biology & Host Immune Networks
Investigates how sequence-divergent pathogens converge on shared host regulatory networks. Analyzes species-specific host co-factors (such as ANP32A/B) and viral disruption of conserved innate immune signaling pathways (RIG-I, MDA5, interferon cascades) across reservoir species and susceptible hosts.
Reproducible Software Architecture & Public Infrastructure
Translates algorithmic discoveries into production-grade pipelines and publicly accessible web infrastructure. Engineers containerized Nextflow DSL2 workflows (MetaNextViro) for high-performance computing clusters and maintains 19 public web servers accessed by over 24,000 researchers across 140 countries.
From molecular classification to predictive host–pathogen interactomes.
I have built an independent methodological trajectory spanning from sequence-level classification to proteome-scale host–pathogen interactomes. To resolve functional properties directly from sequence, I developed SNVguru for variant analysis (adopted by 50+ research groups worldwide). As first author of deepNEC, I conceived and developed an alignment-free architecture achieving >95% accuracy in classifying metabolic enzymes, establishing the deep learning foundation to predict viral variant effects in host cellular contexts.
I next expanded this foundation to transcriptomics and interactomics. To integrate expression dynamics with interactome networks, I developed pySeqRNA, and as first author of HuCoPIA, conceived a coronavirus interactome atlas spanning viral families. For deepHPI, I designed feature extraction and modeling pipelines achieving AUROC > 0.90, part of 13 deployed software packages and web resources accessed by researchers worldwide.
In recent computational genomics research, I translated these capabilities into high-throughput pathogen surveillance and outbreak analytics. To track rapidly evolving viruses during outbreaks, I led computational genomics for targeted amplicon sequencing studies of emerging avian metapneumovirus (AMPV) genomes as first and co-corresponding author, and co-authored studies on viral genomic diversity.
Education, Academic Trajectory & Honors
Education & Academic Timeline
Ph.D. in Plant Sciences (Bioinformatics and Computational Biology)
M.Sc. in Bioinformatics
Add-on Certificate Diploma in Bioinformatics
B.Sc. in Biology
Honors & Awards
Doctoral Student Researcher of the Year
College of Agriculture and Applied Sciences (CAAS), Utah State University
Awarded by the College of Agriculture and Applied Sciences for pioneering computational research in high-throughput genomics, deep learning enzyme prediction, and viral interactomics.
Young Scientist Award
National Conference on Technological Challenges (TECHSEAR-2017)
Conferred at ICAR-Indian Institute of Rice Research (IIRR), Hyderabad, India, in recognition of significant contributions to crop bioinformatics and molecular marker discovery.
Media & Institutional Press
Methodological Deep Dives & Computational Insights
Demystifying Intra-Host Single Nucleotide Variants (iSNVs) in Emerging Viral Surveillance
Why consensus genomes miss early transmission dynamics, and how to calibrate technical error thresholds to accurately detect low-frequency viral quasispecies.
Architecting Production Metagenomics Pipelines with Nextflow DSL2 and Singularity
Building scalable, deterministic bioinformatics workflows that survive high-throughput diagnostic loads on institutional HPC clusters.
Protein Language Models vs. Alignment-Free Classifiers in Enzyme Commission (EC) Prediction
Comparing deep contextual sequence embeddings (ESM-2, ProtTrans) against k-mer representations for enzymatic reaction classification below the twilight zone.
Teaching, Coursework & Computational Mentoring
Bioinformatics and Big Data Mining
Advanced computational methods for biological big data: high-throughput sequencing analysis, sequence alignment algorithms, structural modeling, machine learning in genomics, and high-performance computing cluster utilization.
Bioinformatics Tools and Their Application in Agriculture
Graduate curriculum covering biological databases, molecular phylogenetics, pairwise and multiple sequence alignment, protein secondary structure prediction, and applied crop genomics.
Advances in Bioinformatics
Doctoral curriculum on computational frontiers: comparative genomics, transcriptomic differential expression modeling, protein-protein interaction networks, and structural docking algorithms.
Terminal-First Learning & Rigorous Trainee Ownership
Bridging abstract biological concepts with hands-on computational execution. Equipping life scientists with command-line proficiency, reproducible workflows, and rigorous algorithmic reasoning as core competencies for modern genomics.