CHIRAL Bangladesh CHIRAL Bangladesh
  • About
    • About the centre
    • People
    • Collaborators
    • Support our work
    • Governance & accountability
    • Contact
  • Research
    • Research overview

    • Computational Biology
    • Public Health & Informatics
    • Bio-AI
    • Geospatial Health
  • Publications
    • All research outputs
    • Software & data
  • Training
    • Programmes
    • Learning resources
    • Events & seminars
  • News
  • Join us

Learning resources

The public databases, tools and set-up steps our courses are built on. Everything here is free to use.

Video lectures

Recorded course playlists

Full course recordings on YouTube, free to work through at your own pace.

Statistics · R

R for Research

Video lectures covering R programming fundamentals, data manipulation, statistical testing, and visualization for researchers.

Watch the playlist →
Machine Learning · Bioinformatics

Machine Learning for Bioinformatics

Hands-on video tutorials applying machine learning techniques to biological and genomic datasets.

Watch the playlist →
Genomics · R

RNA-Seq Analysis with R

Step-by-step video tutorials on bulk RNA-seq analysis using R and Bioconductor, from raw counts to biological interpretation.

Watch the playlist →
Machine Learning · Drug Discovery

AI for Drug Discovery

Video series introducing AI-powered approaches to drug discovery, including toxicology modeling and structure-activity analysis.

Watch the playlist →
Cancer Genomics · Bioinformatics

Cancer Bioinformatics

Video tutorials on cancer genomics analysis, covering TCGA data access, mutation analysis, survival modeling, and multi-omics integration.

Watch the playlist →
Academic Writing

Academic Writing

Video lectures on scientific writing, manuscript preparation, and publishing in peer-reviewed journals.

Watch the playlist →
Python · Data Science

Python for Health Data Analytics

Video tutorials on using Python for health data analysis, covering data wrangling, visualization, and statistical methods for public health and clinical datasets.

Watch the playlist →
No matching items
Pipelines

Reusable analysis workflows

Working pipelines from our own projects, with the inputs they expect and the outputs they produce stated up front.

RNA-Seq

End-to-end bulk RNA-seq Quantification Pipeline using Salmon

A complete pipeline for pseudo-alignment and quantification of bulk RNA-seq data using Salmon — from raw sequencing reads to transcript-level abundance estimates ready for downstream differential expression analysis.

In
Raw FASTQ files, reference transcriptome
Out
Transcript/gene-level count matrix, quantification summary
PythonSalmontximetaDESeq2
GitHub →
RNA-Seq

nf-core/rnaseqmeta: A Nextflow Pipeline for Meta-Analysis of RNA-seq Datasets

A Nextflow-based pipeline for reproducible meta-analysis of multiple RNA-seq datasets — automating sample retrieval, quality control, quantification, batch correction, and integrated differential expression analysis across studies.

In
Multiple RNA-seq datasets (FASTQ or SRA accessions), sample sheets
Out
Integrated count matrix, cross-study DE results, batch-corrected expression profiles
Nextflownf-coreSalmonDESeq2
GitHub →
Single-Cell

Fast and efficient preprocessing of scRNA-seq with kallisto | bustools | kb-python

End-to-end pipeline for scRNA-seq preprocessing using kallisto, bustools, and kb-python — from raw FASTQ files to filtered count matrices ready for downstream analysis with scanpy.

In
Raw FASTQ files (10x Chromium or similar)
Out
Filtered cell × gene count matrix, QC metrics
Pythonkallistobustoolskb-pythonscanpy
Google Colab →
Single-Cell

A Practical Guide for Single-Cell Data Analysis with scverse Ecosystem

Comprehensive walkthrough of the scverse single-cell analysis workflow — covering quality control, normalization, dimensionality reduction, clustering, marker gene identification, and cell type annotation using scanpy and related tools.

In
Count matrix (AnnData .h5ad or 10x format)
Out
Annotated cell clusters, UMAP embeddings, marker genes, cell type labels
Pythonscanpyscvi-toolsAnnData
Google Colab →
Protein Structure Prediction

Predicting Protein Structures with ColabFold and AlphaFold2 in Google Colab

Predict 3D protein structures from amino acid sequences using ColabFold's accelerated AlphaFold2 pipeline — with MSA generation via MMseqs2 and interactive structure visualization directly in the browser.

In
Amino acid sequence (FASTA format)
Out
Predicted 3D structures (PDB), confidence scores (pLDDT), PAE plots
PythonColabFoldAlphaFold2MMseqs2py3Dmol
Google Colab →
Protein Structure Prediction

Boltz2-Notebook: Diffusion-Based Protein–Ligand Structure Prediction & Affinity Analysis

Predict protein–ligand complex structures and binding affinities using the Boltz2 diffusion model — enabling rapid in silico docking and interaction analysis without traditional molecular dynamics simulations.

In
Protein sequence and ligand SMILES/SDF
Out
Predicted complex structures (PDB), binding affinity scores, interaction maps
PythonBoltz2RDKitpy3Dmol
Google Colab →
Drug Discovery

AI in Drug Discovery: Molecular Property Prediction and Virtual Screening

End-to-end pipeline for AI-driven drug discovery — covering molecular featurization, toxicity prediction, ADMET property modeling, and virtual screening of compound libraries using machine learning and cheminformatics tools.

In
Compound libraries (SMILES), molecular descriptors
Out
Toxicity predictions, ADMET profiles, ranked hit compounds
PythonRDKitDeepChemscikit-learn
Google Colab →
Drug Discovery

AI in Drug Discovery: In Silico Toxicology Modeling

Build machine learning models to predict compound toxicity from molecular structure — covering molecular fingerprint generation, toxicity endpoint classification, structure–activity relationship analysis, and model interpretation for safety assessment in early-stage drug discovery.

In
Chemical compounds (SMILES), toxicity endpoint labels
Out
Toxicity classification models, SAR insights, safety predictions
PythonRDKitscikit-learnMordred
Google Colab →
No matching items

Data and Tools for Cancer Genomics

  • The Cancer Genome Atlas (TCGA): TCGA is a comprehensive collection of multi-dimensional cancer genomics data covering multiple cancer types.

  • International Cancer Genome Consortium (ICGC): Description: ICGC provides high-quality genomic and clinical data from various cancer projects worldwide.

  • Gene Expression Omnibus (GEO): GEO is a public repository hosted by the National Center for Biotechnology Information (NCBI) containing a vast collection of gene expression data, including cancer datasets.

  • European Genome-phenome Archive (EGA): Description: EGA is a repository for secure storage and sharing of human genetic and phenotypic data, including cancer datasets.

  • National Cancer Institute (NCI) Genomic Data Commons (GDC): Description: GDC is an open-access data portal providing access to a wide range of cancer genomics datasets.

  • OncoLnc: Description: OncoLnc is a web resource that provides survival analysis and expression correlation for genes of interest across multiple cancer datasets.

  • UCSC Cancer Genomics Browser: The UCSC Cancer Genomics Browser offers a comprehensive collection of cancer genomics data integrated with genomic annotations.

  • GREIN : GEO RNA-seq Experiments Interactive Navigator: GREIN is an interactive web platform that provides user-friendly options to explore and analyze GEO RNA-seq data. GREIN is powered by the back-end computational pipeline for uniform processing of RNA-seq data and the large number (>6,000) of already processed datasets. These datasets were retrieved from GEO and reprocessed consistently by the back-end GEO RNA-seq experiments processing pipeline (GREP2).

  • GEPIA2: GEPIA2 is a web-based tool for analyzing gene expression data in cancer. It stands for Gene Expression Profiling Interactive Analysis 2 and is an updated version of the original GEPIA tool. GEPIA2 allows users to explore gene expression patterns, perform survival analyses, and visualize gene expression data across various cancer types.

  • UALCAN: UALCAN is a web-based platform that provides interactive and comprehensive analysis of cancer transcriptome data. It enables users to explore gene expression patterns, perform survival analyses, and compare gene expression between tumor and normal samples across different cancer types. UALCAN utilizes data from The Cancer Genome Atlas (TCGA) to facilitate cancer research and provide insights into tumor biology.

  • cBioPortal for Cancer Genomics:: cBioPortal hosts a large collection of cancer genomics datasets, allowing users to explore and visualize the data.

  • ONCOMINE: ONCOMINE is a powerful web-based platform for the analysis and visualization of cancer transcriptomic data. It provides researchers with access to a vast collection of publicly available gene expression datasets derived from cancer studies. ONCOMINE allows users to explore gene expression patterns, identify potential biomarkers, and compare gene expression between different cancer types or subtypes.

Guideline for Bioconductor Users

Bioconductor is an open-source and open-development software project that provides a comprehensive collection of bioinformatics and computational biology tools in the R programming language. It focuses on the analysis and comprehension of high-throughput genomic data, including DNA sequencing, RNA sequencing, microarray analysis, proteomics, and more.

Required software

  • R: http://www.r-project.org/ (FREE)
  • RStudio (additional libraries required): http://www.rstudio.com/ (FREE)

Prework

Before attending the any workshop please have the following installed and configured on your machine. - Recent version of R - Recent version of RStudio - Most recent release of the Bioconductor and other packages used in courses

Install the latest release of R, then get the latest version of Bioconductor by starting R and entering the commands.

if (!require("BiocManager", quietly = TRUE))
    install.packages("BiocManager")
BiocManager::install(version = "3.16")
  • Ensure you can knit R markdown documents

    • Open RStudio and create a new Rmarkdown document
    • Save the document and check you are able to knit it.

Install Bioconductor Packages

if (!require("BiocManager", quietly = TRUE))
    install.packages("BiocManager")
BiocManager::install()

Install specific packages, e.g., “GenomicFeatures” and “AnnotationDbi”, with

BiocManager::install(c("GenomicFeatures", "AnnotationDbi"))

The install() function (in the BiocManager package) has arguments that change its default behavior; type ?install for further help.

R Packages RNASeq and Single-cell RNA-seq Analysis

  • DESeq2: DESeq2 is a widely used package for differential gene expression analysis in RNA-seq data.
  • edgeR: edgeR is another popular package for differential gene expression analysis in RNA-seq data.
  • limma: limma is a package commonly used for the analysis of microarray and RNA-seq data, particularly for differential expression analysis.
  • Ballgown: Ballgown is a package for differential expression analysis and visualization of transcriptome assembly data.
  • DEXSeq: DEXSeq is specifically designed for the detection of differential exon usage in RNA-seq data.
  • NOISeq: NOISeq is a package for non-parametric analysis of differential expression in RNA-seq data.
  • clusterProfiler: clusterProfiler is a package for functional enrichment analysis of gene clusters derived from RNA-seq data.
  • GenomicFeatures: GenomicFeatures provides tools for working with genomic features, such as gene models, and is useful for annotating RNA-seq data.
  • Seurat: Seurat is a package for single-cell RNA-seq data analysis, allowing exploration and visualization of cellular heterogeneity.

Blogs for R Programming, Statistics, and Data Analyis

  • Programiz - https://www.datamentor.io/r-programming/
  • PennState STAT 484 - https://online.stat.psu.edu/stat484/
  • PennState Topics in R Statistical Language - https://online.stat.psu.edu/stat484/
  • Simply Statistics - https://simplystatistics.org/
  • TutorialPoint - https://www.tutorialspoint.com/r/index.htm
  • R for Biologists - https://www.rforbiologists.org/
  • Computational Genomics with R - https://compgenomr.github.io/book/
  • Stat and R - https://statsandr.com/
  • Rafa Lab - https://rafalab.github.io/pages/harvardx.html
  • University of Florida - https://bolt.mph.ufl.edu/software/r-phc-6055/

Videos

  • July 9, 2023: I’m thrilled to announce that the highly anticipated videos from our workshop on “R for Research: Fundamentals of R - Part 1” are now available for viewing. Whether you missed the live event or want to revisit the valuable insights shared during the session, these videos are your gateway to mastering the basics of R programming language for data analysis and research. Join us as we explore the foundations of R and learn essential skills to enhance your data analysis capabilities. Don’t wait any longer; dive into the videos today and take your research to new heights! Check it out:

  • April 12, 2023: Watch this informative 2-hour workshop on how NASA Earth Observing Data can help improve public health. Discover how we can use these data to monitor our environment and identify potential health risks. Learn about the different ways NASA Earth Observing Data can benefit our communities and keep us safe. Check it out:

 

CHIRAL Bangladesh

কাইরাল বাংলাদেশ

Centre for Health Innovation, Research, Action and Learning. Computational biology, Bio-AI and population health research, based in Dhaka.

Research

  • Overview
  • Computational Biology
  • Public Health & Informatics
  • Bio-AI
  • Geospatial Health

Outputs

  • Publications
  • Software & data
  • Training programmes
  • Learning resources
  • Events

Centre

  • About
  • People
  • Collaborators
  • Join us
  • Support us
  • Governance
  • Contact

Azimpur, Dhaka 1205, Bangladesh · chiralbd@gmail.com © 2025 CHIRAL Bangladesh · Content CC BY 4.0 · Code MIT