Reading a Raw DNA File: The Format and Its Limits
Four columns, several hundred thousand rows, and a set of things the file genuinely cannot tell you.
Consumer testing companies let you download the raw data behind your report. It is a plain text file, usually tab-separated, and its structure is simple enough to read in any text editor. What it supports concluding is narrower than it first appears.
The format
After a block of comment lines beginning with #, each row is one genotyped position:
| rsid | chromosome | position | genotype |
|---|---|---|---|
| rs4988235 | 2 | 136608646 | AG |
| rs1815739 | 11 | 66560624 | CT |
| rs53576 | 3 | 8804371 | GG |
- rsid — the catalogue identifier for that position. Some rows instead carry an internal
i-prefixed ID, which is proprietary to the testing company and not searchable elsewhere. - chromosome — 1–22, then X, Y and MT for mitochondrial.
- position — the coordinate along that chromosome, relative to a specific reference build. Usually GRCh37, sometimes GRCh38, and the same position number means different places in each. The header comments state which; ignoring that is how people look up the wrong variant.
- genotype — your two alleles, one per inherited copy.
--means the position failed to read, which happens on a meaningful fraction of rows.
The strand problem
DNA is double-stranded and the two strands are complementary, so the same variant can be reported as AG or as CT depending on which strand the company reported. Comparing your genotype to a published study without checking strand orientation will sometimes give you the exact opposite of the truth.
Most consumer files report on the plus strand of the reference, but this is a convention rather than a guarantee, and it is the single most common source of confidently wrong home interpretation.
What the file is not
It is not your genome sequence. A genotyping array tests several hundred thousand chosen positions — well under 0.03% of the roughly 3.1 billion base pairs. Everything not on the chip is simply absent, and absence in this file is not evidence of anything.
That matters most for rare variants. A clinically significant mutation that is not on the array will not appear, and a file showing nothing alarming is not a clean bill of health. Screening is what sequencing is for.
Array data has real error rates
Consumer arrays are accurate in aggregate and imperfect at any single position. For common variants that is fine. For rare ones it is a genuine problem: when the true frequency of a variant is very low, even a small false-positive rate means a large share of positive calls are wrong.
This is why third-party interpretation services returning alarming findings from raw data have been a recurring problem, and why any result with clinical weight needs confirmation on a clinical-grade test.
Reasonable things to do with it
- Look up a specific rs number and read what the literature actually reports about it — noting whether the studies are association or causation, and what the absolute risk is.
- Check the effect sizes. Most common-variant associations are small in absolute terms.
- Compare against a relative's file to see inheritance directly.
- Keep a copy. The file is yours, it does not change, and re-downloading later is not always possible.
Privacy is the part worth slowing down on
This file identifies you and, partially, everyone related to you — who never consented to anything. Uploading it to a third-party interpretation service means handing that over, and the terms governing what happens to it if the company is sold or breached are frequently not what people assume. Genetic data cannot be rotated like a password.
None of that argues against downloading and keeping your own file. It argues against uploading it casually.