India Unfolding
All sectorsVerified data
Genomics & Biotech Research

Genomics & Biotech Research

India's public genomics effort rests on shared research infrastructure: population-scale sequencing, a national repository for biological data, health cohorts with biobanks, and incubators for biotech start-ups. Under GenomeIndia, sanctioned in January 2020, the Department of Biotechnology sequenced the whole genomes of 10,074 healthy individuals from 83 populations and archived the data at the Indian Biological Data Centre (IBDC) in Faridabad. The data was dedicated to researchers on 9 January 2025. CSIR's Phenome India cohort and its National Biobank are collecting genomic, lifestyle and clinical data from 10,000 people, and the DBT–BIRAC network had 94 bioincubators across 25 States and Union Territories by April 2026.

A building on the campus of the National Centre for Biological Sciences, Bengaluru.
A building on the campus of the National Centre for Biological Sciences, Bengaluru. · Rohit Suratekar · CC BY-SA 4.0 · Wikimedia Commons
10,074
Healthy individuals whole-genome sequenced under GenomeIndia (83 populations)
94
DBT–BIRAC bioincubators across 25 States and UTs (April 2026)
13,000
Biotech start-ups in India (2025)
Healthy individuals whole-genome sequenced under GenomeIndia (83 populations, 99 sites)
2020
GenomeIndia sanctioned
The Department of Biotechnology sanctioned GenomeIndia on 16 January 2020 for three years across 20 institutions, with a target of sequencing the whole genomes of 10,000 individuals.
2020
Biological data centre set up
DBT established the Indian Biological Data Centre (IBDC) in March 2020, with a 4 PB parallel file system and 1.5 PB of disk and tape for backup.
2022
IBDC dedicated to the nation
2023
Phenome India cohort launched
2024
Phenome India crosses 10,000 samples
2024
BioE3 policy approved
2025
GenomeIndia data released
2025
National Biobank opens
2026
94 bioincubators
2020
2milestones
20202026
Latest
Biological data centre set up
DBT established the Indian Biological Data Centre (IBDC) in March 2020, with a 4 PB parallel file system and 1.5 PB of disk and tape for backup.

Why it matters Most disease-risk prediction algorithms are based on data from Caucasian populations, and CSIR notes they may not be very accurate for Indians because of differences in genetic make-up and lifestyle. An Indian genome dataset can be used to develop indigenous genomic chips, diagnostics and therapeutics. Shared data centres and incubators give laboratories and start-ups access to computing, storage, laboratory space and regulatory guidance.

  • GenomeIndia: 10,074 healthy individuals from 83 populations at 99 sites; about 36.7% of samples rural, 32.2% urban and 31.1% tribal (DBT, March 2025)
  • GenomeIndia resource at IBDC: FASTQ files of 9,772 samples (~700 TB), gVCFs (~35 TB) and phenotypic data from 9,330 samples (DBT, April 2025)
  • IBDC, set up in March 2020, is mandated to archive all life science data from publicly funded research in India and runs seven data portals
  • INSACOG, set up in December 2020 with 10 laboratories to sequence SARS-CoV-2, had 28 laboratories by July 2021 (Ministry of Health)
  • BIRAC incubation network, December 2025: 75 BioNEST and 19 E-YUVA centres, over 9,00,000 sq. ft., supporting more than 3,000 entrepreneurs and start-ups
  • Biotech start-ups: 5,365 in 2021 and 13,000 in 2025 (PIB backgrounder, September 2025)

History

The Department of Biotechnology sanctioned GenomeIndia: Cataloguing the Genetic Variation in Indians on 16 January 2020, for three years across 20 institutions, with a target of whole genome sequencing of 10,000 individuals. In March 2020 it established the Indian Biological Data Centre (IBDC), and the Biotech-PRIDE Guidelines on data sharing followed in 2021. During the COVID-19 pandemic, 16 bio-repositories for COVID-19 samples were designated in May 2020 under ICMR, DBT and CSIR. The Indian SARS-CoV-2 Genomics Consortium (INSACOG) was set up in December 2020 with 10 laboratories to track variants and had 28 by July 2021.

Notable milestone

On 9 January 2025, at the Genomics Data Conclave, the GenomeIndia data was dedicated to researchers, and the Framework for Exchange of Data (FeED) Protocols and IBDC portals were launched to give access to 10,000 whole genome samples. The dataset covers 10,074 healthy individuals from 83 populations at 99 sites, with about 36.7% of samples from rural, 32.2% from urban and 31.1% from tribal populations. As of April 2025, the resource at IBDC comprised FASTQ files of 9,772 samples (~700 TB), gVCF files (~35 TB) and phenotypic data from 9,330 samples. DBT has invited proposals for translational research using the data, and also accepts independent requests for access.

How it works

The IBDC, at the Regional Centre of Biotechnology in Faridabad with a disaster recovery site at NIC Bhubaneswar, is mandated to archive all life science data generated from publicly funded research in India and runs seven specialised data portals. GenomeIndia data is shared under the Biotech-PRIDE Guidelines and FeED Protocols: researchers get gVCF files, while the ~700 TB of raw FASTQ files are not available for download at present. CSIR's Phenome India cohort, launched on 7 December 2023, collects clinical, lifestyle, imaging, biochemical and molecular data to build India-specific risk models for cardio-metabolic disease. Under BIRAC's BioNEST scheme, bioincubators give start-ups incubation space, access to equipment, mentoring and regulatory guidance.

Outlook

At the January 2025 conclave, the Minister of State for Science and Technology announced a future target of sequencing 10 million genomes. The Phenome India National Biobank, opened at CSIR-IGIB on 6 July 2025, is to track the health of its 10,000 participants over several years. The DBT–BIRAC network reached 94 bioincubators across 25 States and UTs by April 2026; in December 2025 DBT reported 75 BioNEST and 19 E-YUVA centres supporting more than 3,000 entrepreneurs and start-ups. Under the BioE3 policy, 186 projects had been recommended for funding and a network of six biofoundries launched by December 2025; the bio-economy's size is tracked on the Bio-economy · BioE3 page.

By the numbers

10,074 individuals whole-genome sequenced under GenomeIndia (83 populations, 99 sites). ~700 TB of FASTQ files from 9,772 samples archived at IBDC (April 2025). 10,000 participants in the Phenome India cohort and National Biobank. 94 bioincubators across 25 States and UTs (April 2026). Biotech start-ups 5,365 (2021) → 13,000 (2025). Sources: Department of Biotechnology, CSIR, BIRAC and Ministry of Health & Family Welfare via PIB; Lok Sabha reply (19 March 2025); Rajya Sabha reply (2 April 2026); BIRAC.

Data current to: as reported up to 2 April 2026

Source: Department of Biotechnology, CSIR and BIRAC via PIB, and Parliament replies — the timeline dates the sanction of GenomeIndia (January 2020), the Indian Biological Data Centre (set up March 2020, dedicated to the nation November 2022), CSIR's Phenome India cohort (December 2023 and June 2024), the BioE3 policy (August 2024), the release of GenomeIndia data (January 2025), the Phenome India National Biobank (July 2025) and the DBT–BIRAC bioincubator network (April 2026). Headline: Lok Sabha reply published by PIB on 19 March 2025 — whole genome sequencing of 10,074 healthy individuals from 83 populations at 99 sites, archived at IBDC (linked). · link