Illustration by tuput
English
The Department of Biotechnology opened the GenomeIndia data to researchers on 9 January 2025. A preprint posted on 24 March 2026 analyses the 9,768 genomes that passed quality checks and reports 129.93 million variants, 44.03 million of them absent from global databases. It had not been peer reviewed when posted.
The GenomeIndia project sequenced the whole genomes of 10,074 healthy adults from 83 population groups, and on 9 January 2025 the Department of Biotechnology (DBT) opened the data to researchers through the Indian Biological Data Centre (IBDC) in Faridabad. The US National Human Genome Research Institute describes a genome as the full set of DNA letters in a cell, about 6 billion counting both inherited copies, and a variant as a place where those letters differ between people.
The fullest analysis so far is a preprint, a paper posted before peer review. Dated 24 March 2026, it reports 129.93 million variants in the 9,768 people who passed quality checks, and says 44.03 million of them were previously unreported in global databases. The earlier journal item, a Nature Genetics comment of April 2025, gave preliminary counts.
What was released on 9 January 2025
At the Genomics Data Conclave in Vigyan Bhavan, New Delhi, Union Minister Jitendra Singh launched the Framework for Exchange of Data (FeED) protocols and the IBDC portals. The Press Information Bureau release says this made 10,000 whole genome samples accessible to researchers in India and abroad. Singh said the collection “is now made available for research purposes not only within India but globally.”
Prime Minister Narendra Modi sent a video message. He said the project was approved five years earlier and involved more than 20 institutions, among them IISc, the IITs, CSIR and BRIC. His example of what the data could settle was sickle cell anaemia in tribal communities: the problem in one region may differ from another, he said, and only a full genetic study would show how.
The project’s FAQ says it was conceived in late 2017 and launched in January 2020, with the Centre for Brain Research at IISc Bengaluru as coordinator.
Why the number is 83, not 99
Two figures circulate. A Lok Sabha reply released by PIB on 19 March 2025 says the DBT built a resource of 10,074 healthy individuals “from 83 heterogeneous populations from 99 different sites”. The 83 counts groups and the 99 counts places where samples were taken.
The comment explains the design. Of the 83 groups, 30 are tribal and 53 non-tribal. Language is a proxy for genetic diversity in India, so the team made sure the four large language families (Indo-European, Dravidian, Austro-Asiatic and Tibeto-Burman) were represented, and within each region it sampled groups from distinct bio-geographies. It sequenced a median of 159 unrelated people from each non-tribal group and 75 from each tribal group, plus a few parent-and-child trios per group.
Participants were self-declared healthy adults with no diagnosed single-gene disease or chromosomal abnormality. The comment reports 20,459 consenting volunteers, while the preprint and FAQ say 20,195. A genotyping array run on 13,242 samples helped pick people unrelated to each other beyond first cousins, and 10,074 were sequenced, mainly at four centres. After quality control the comment counts 9,772 genomes and the preprint 9,768. The Lok Sabha reply gives the sampling mix as 36.7 percent rural, 32.2 percent urban and 31.1 percent tribal.
What 129.93 million variants means
The preprint reports sequencing at about 30 times coverage, meaning each position was read about 30 times. It finds 129,938,889 high-confidence variants, about 121 million single-letter changes and 8 million short insertions or deletions. About 59 million are singletons, seen once in the whole cohort. Of the 70.74 million variants seen more than once, 11.94 million (9.19 percent) are unreported in global databases.
So the headline 44.03 million includes variants found in a single person. Nature India, reporting on 4 May 2026, says about 44 million are absent from gnomAD (a large public variant database), the 1000 Genomes Project and GenomeAsia. The comment’s earlier figures were about 180 million raw variants before filtering and around 130 million on the autosomes (the non-sex chromosomes) afterwards.
Tribal groups, founder effects and the Nilgiris
In the preprint’s principal component analysis, which sorts people by overall genetic similarity, tribal Austro-Asiatic, Dravidian and Tibeto-Burman groups form tight, separate clusters. Non-tribal Indo-European and Dravidian groups overlap along a rough north-to-south gradient. Longitude alone predicts the first component with an adjusted R-squared of 0.43 and latitude predicts the second with 0.30. Adding language and biogeography brings the model to 84 percent and 78 percent. The analysis finds four broad ancestry components and a fifth for two isolated Dravidian tribal groups from the Nilgiri hills.
Endogamy, marriage within a community, leaves a mark in runs of homozygosity, long stretches where both inherited copies of DNA are identical because of shared recent ancestors. The preprint says the median length of these runs significantly exceeded a 1.5 centimorgan endogamy threshold (a centimorgan is a unit of genetic distance) in 59 of 83 populations. For the two Nilgiri groups the authors say the number of runs ranks among the world’s highest and is more than five times the Ashkenazi Jewish and Finnish levels. The authors write that their findings “demonstrate recent endogamy as a dominant force.”
A founder effect means a small ancestral group gave rise to many descendants, so a variant rare elsewhere can become common. The preprint finds a splice-site variant in the HGD gene, absent from gnomAD, the 1000 Genomes Project and GenomeAsia, at 12.5 percent in one tribal population in the south. Nature India links HGD to alkaptonuria, a rare metabolic disorder. The preprint puts the share of individuals carrying the pathogenic variants it counts at 17.5 percent in tribal populations against 5.5 percent in non-tribal ones. The sickle-cell variant rs334 exceeded 5 percent in several tribal populations.
The FAQ says the project avoids stigma by analysing populations rather than individuals, and that the results do not support simple biological divisions between communities, since most groups share ancestry through past admixture.
What the data are meant to build
The PIB release says the data should enable genotyping chips tailored to Indian populations. Singh listed mRNA vaccines, protein manufacturing and genetic-disorder treatments as areas the data could feed, alongside the drug-making base described in how India became the world’s pharmacy.
The preprint’s concrete product is an imputation panel, a reference of fully sequenced genomes used to fill in untested letters in people who were only run on a cheaper array. Built from 9,768 phased genomes and tested on array data from 7,628 South Asian ancestry participants in the UK Biobank, it improved average imputation quality by 45 percent over the GenomeAsia panel, 20 percent over the Haplotype Reference Consortium panel and 9 percent over TOPMed.
It also tests polygenic scores, which add up many small genetic effects into one risk estimate. Scores built from European-ancestry studies explained 14.8 percent of height variation in GenomeIndia against 11.2 percent in UK Biobank Indians, but only 0.7 percent of BMI variation against 9.7 percent. The authors link the BMI gap to genetic distance and environmental differences. Diabetes risk in India has environmental and early-life threads of its own, covered in the Pune research on diabetes passing between generations.
For drug handling, the preprint reports group-level allele frequencies, for example NUDT15 risk alleles up to 20.8 percent in Austro-Asiatic groups. It tests no individual’s treatment. Meera Purushottam of NIMHANS told Nature India that the effect of such panels on drug choice and dosing is unclear and that adoption should be cautious.
Who can get the data, and in what form
The preprint says allele frequency summary statistics are public through the IBDC. Raw sequencing data are under controlled access because of privacy laws tied to participant consent, and bona fide researchers must send the Data Management Group a request describing the proposed research, and approval is required.
PIB’s 30 April 2025 release, written in reply to newspaper reports, lists what the IBDC holds: FASTQ files (raw reads) for 9,772 samples at about 700 TB, gVCF files (per-person variant calls) for 9,772 samples at about 35 TB, joint call files of about 3.5 TB, and phenotype data for 9,330 samples. Under DBT’s call, FASTQ files are not downloadable at present, the release says, because of their size and because analysing raw files needs two to three times the computing. Phenotype data cover 27 blood measures, such as haemoglobin and cholesterol, plus age, gender, height, weight and body fat. For 442 samples the data were unusable. Access is not limited to DBT’s call for proposals on translational research, and independent requests under the Biotech-PRIDE Guidelines (2021) and FeED protocols are accepted.
The sources do not fully agree. The later preprint describes raw sequencing data as available under controlled access, so the rules for raw reads are worth confirming with the IBDC. The FAQ lists sequencing data for 9,648 samples at the IBDC where PIB lists 9,772.
What the dataset cannot show
The preprint says its 9,768 people do not fully cover India and calls for more sampling of large non-tribal populations. The comment puts India at over 4,600 endogamous groups, of which GenomeIndia covers 83.
The participants were healthy, so Nature India notes that links to disease are inferred, and many loss-of-function variants may be harmless. Sudhakaran Prabakaran of Northeastern University told it the analysis covers protein-coding genes, about 2 percent of the genome. The comment said analysis was ongoing and a detailed manuscript was being prepared. As of 8 October 2026 tuput found no journal version of the atlas.
Analabha Basu, a corresponding author, described the work to Nature India as a starting point and said more is needed to translate it into clinical practice.
Sources & further reading
- PIB (Ministry of Science and Technology), 9 January 2025: India Takes a Giant Leap in Genomics, Launch of Indian Genomic Data Set and IBDC Portals
- PIB (Prime Minister's Office), 9 January 2025: English rendering of PM's remarks at the start of GenomeIndia Project
- PIB, 19 March 2025: Parliament question on the Indian Biological Data Centre (Lok Sabha reply)
- PIB, 30 April 2025: GenomeIndia, access to the national genetic resource
- Bhattacharyya C et al., Mapping genetic diversity with the GenomeIndia project, Nature Genetics 57, 767 to 773 (2025), a Comment by the GenomeIndia Consortium
- NCBS repository record for the Nature Genetics comment (volume, issue, pages, PubMed ID 40200122)
- Subramanian K, Bhattacharyya C et al., An Atlas of Indian Genetic Diversity, medRxiv preprint, posted 24 March 2026 (not certified by peer review)
- GenomeIndia project: Frequently asked questions
- Nature India (Sahana Ghosh), 4 May 2026: India's DNA map uncovers millions of missing genetic variants
- US National Human Genome Research Institute: Genomic data science fact sheet
Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.
Enjoyed this? Get the next one.
One good read at a time, straight to your inbox. No spam, unsubscribe anytime.