Acta Computationis Biologicae

Supplementary Information

S5. Experiments

Numbered as in a results section. Status included.


Each experiment is a question, a method, and a present tense. Code repositories will be linked when they are public. Until then the GitHub profile is the working shelf: github.com/shweta-rai11.

Experiment 1. Sex-Specific Endocannabinoid System Biomarkers in Rheumatoid Arthritis

2025 - Present · PhD Research, University of Lethbridge · Genomics & Transcriptomics

Question. Do endocannabinoid-system genes show sex-stratified differential expression in rheumatoid arthritis, and is that association consistent with a causal role under Mendelian randomization?

Approach. Testing whether endocannabinoid-system genes show sex-stratified differential expression in rheumatoid arthritis, and whether that pattern holds up under Mendelian randomization. Expression is normalized and modeled separately by sex with limma, with FDR-corrected hits cross-referenced against WGCNA co-expression modules; the Mendelian randomization step is still running.

Pipeline. RNA → QC → Sex-stratified DE → WGCNA modules → MR causal inference → ECS biomarkers.

Tools. R, limma, WGCNA, TwoSampleMR, GEOquery.

Experiment 2. Multi-Omics Biomarker Discovery in Rheumatoid Arthritis

2024 - Present · PhD Research, University of Lethbridge · Multi-Omics & Cross-Omics Integration

Question. Can integrating transcriptomic and methylomic signal from public repositories reveal robust, ML-tractable biomarkers for rheumatoid arthritis (RA)?

Approach. Built a pipeline that takes raw GEO transcriptomics and methylomics data, integrates the two layers, and narrows them down to a ranked set of candidate rheumatoid arthritis biomarkers. DESeq2/limma differential expression feeds a Boruta feature-selection step, then cross-validated classifiers (logistic regression, random forest, XGBoost) rank the surviving candidates. This is the computational backbone of my ongoing PhD research.

Pipeline. Transcriptome → Methylome → Integration → Feature selection → Classification → Biomarker panel.

Tools. R, DESeq2, limma, Python, scikit-learn, XGBoost, Boruta.

Experiment 3. Interactive Web Application for Multi-Omics RA Data Visualization

2024 - Present · PhD Research, University of Lethbridge · Web Development

Question. How can multi-omics analysis for rheumatoid arthritis be made accessible to collaborators without requiring them to run R scripts directly?

Approach. Wrapped the biomarker pipeline in an R Shiny app so lab members can explore differential expression and methylation results interactively instead of running R scripts themselves. Still a work in progress alongside the PhD research.

Pipeline. Pipeline outputs → Shiny UI → Interactive plots → Lab-wide deployment.

Tools. R, Shiny, ggplot2, DESeq2.

Experiment 4. In-Silico Screening of Anti-Arthritic Phytochemicals

2021 - 2022 · MSc Thesis, Bangalore University · Proteomics & Molecular Docking

Question. Do phytochemicals from Cinnamomum zeylanicum and related sources show binding affinity for rheumatoid arthritis drug targets competitive with standard RA pharmaceuticals?

Approach. For my MSc thesis, I screened around 3,000 phytochemicals against rheumatoid arthritis drug targets using molecular docking, to see if any natural compounds could rival standard RA drugs. A handful of candidates came out ahead in the in-silico binding affinity comparison.

Pipeline. Compound library → Target prep → Docking → Affinity ranking → Drug benchmarking.

Tools. AutoDock, PyMOL, R, Python.

Experiment 5. Automated Mining of GEO, PubMed & ClinicalTrials Metadata

2024 - Present · PhD Research, University of Lethbridge · Data Engineering

Question. How can candidate transcriptomics/methylomics datasets and their associated publications be discovered and linked systematically, rather than manually?

Approach. Wrote a scraper that pulls candidate transcriptomics and methylomics datasets from GEO and cross-references them with PubMed and ClinicalTrials IDs, instead of doing that search by hand. Feeds directly into dataset selection for the multi-omics RA project.

Pipeline. GEO query → Scrape metadata → Cross-reference IDs → Structured dataset.

Tools. Python, BeautifulSoup, pandas.

Experiment 6. Synthetic Biomedical Image Generation with GANs

2024 · PhD Research, University of Lethbridge · Machine Learning & Modeling

Question. Can Generative Adversarial Networks produce realistic synthetic biomedical images to augment limited biomedical imaging datasets?

Approach. An exploratory project testing whether GANs could generate realistic synthetic biomedical images to help with small, data-scarce imaging datasets. Early-stage work, not part of the main PhD pipeline.

Pipeline. Data → Generator / discriminator → Adversarial training → Quality evaluation → Synthetic samples.

Tools. Python, TensorFlow/PyTorch, OpenCV.

Experiment 7. RNA-Seq Differential Expression & CRISPR Guide Design

2022 · Bioinformatics Internship, ArrayGene · Genomics & Transcriptomics

Question. Which genes are differentially expressed across experimental conditions, and which are viable CRISPR-Cas9 targets for functional follow-up?

Approach. During my internship at ArrayGene, I ran differential expression analysis on RNA-seq data with DESeq2 and edgeR (negative-binomial modeling, FDR-corrected significance) and designed CRISPR-Cas9 guides with on-target and off-target scoring for client genomics projects. I delivered heatmap, volcano, and PCA summaries alongside the guide candidates.

Pipeline. RNA → QC → DE (edgeR) → Pathway enrichment → CRISPR guide design.

Tools. R, DESeq2, edgeR, samtools, Linux.