SNPs: What a Single Letter Change Does and Does Not Mean

rs numbers, odds ratios, and why 'associated with' is doing far more work than it looks.

A SNP — single nucleotide polymorphism, said "snip" — is one position in the genome where people commonly differ by a single base. One person has a C, another a T. There are on the order of a hundred million known SNPs, and a few million vary commonly enough to be worth genotyping.

rs numbers

Each catalogued SNP has a reference number: rs4988235, rs1815739. The prefix is "reference SNP" and the identifier is stable across databases, which makes it the thing to search when you want to know what a position is actually associated with.

A raw data file from a consumer test is essentially a list of rs numbers with your two alleles at each — one from each parent, so CT means you carry one of each.

Most SNPs do nothing

The majority sit in non-coding regions with no known function. Even inside a gene, the genetic code is redundant — several triplets specify the same amino acid — so a change can be synonymous and produce an identical protein.

The ones that matter fall into a few groups: those changing an amino acid, those creating or destroying a stop signal, those disturbing splicing, and those in regulatory sequence that change how much protein is made rather than its shape. That last group is easy to overlook and frequently the important one.

Linkage disequilibrium, and why the tested SNP is often not the causal one

DNA is inherited in blocks, so nearby variants travel together. A SNP can be strongly associated with a trait while having no effect at all, simply because it sits beside the variant that does.

This is what makes genotyping arrays work — testing one SNP per block captures its neighbours — and it is why a study identifying a SNP has usually identified a region, not a cause. Reading "rs12345 causes X" into a paper reporting association is the single most common misreading in this area.

Reading an effect size honestly

Associations are reported as odds ratios or relative risks, and both are ratios — they say nothing about magnitude on their own.

An odds ratio of 1.3 against a baseline lifetime risk of 2% moves you to roughly 2.6%. Reported as "30% higher risk" that sounds alarming; stated as "from 2 in 100 to about 2.6 in 100" it does not. Always ask what the absolute baseline is, because most common-variant effects are in this range.

The exceptions are real and worth distinguishing: some variants, particularly in BRCA1 and BRCA2, carry large effects. Those are rare, clinically actionable, and not what a typical consumer report is mostly reporting.

Polygenic scores

Most traits are influenced by thousands of variants with individually tiny effects. A polygenic score adds them up, weighted by effect size, into a single number.

Two caveats it is dishonest to omit. Such scores typically explain a modest share of variation even at their best. And they were overwhelmingly derived from studies of people of European ancestry, so they transfer poorly to other populations — a well-documented limitation, not a hypothetical one.

Before acting on any of it

Consumer genotyping arrays have real error rates at individual positions, and a rare-variant call from an array is far less reliable than one from sequencing. Any result with clinical significance should be confirmed by a clinical-grade test and discussed with someone qualified — which is the standard advice precisely because the failure mode is a false positive on a frightening variant.

Reading your raw data · Genotype and phenotype · Back to all articles