Reading a Raw DNA File: The Format and Its Limits

Four columns, several hundred thousand rows, and a set of things the file genuinely cannot tell you.

Consumer testing companies let you download the raw data behind your report. It is a plain text file, usually tab-separated, and its structure is simple enough to read in any text editor. What it supports concluding is narrower than it first appears.

The format

After a block of comment lines beginning with #, each row is one genotyped position:

rsidchromosomepositiongenotype
rs49882352136608646AG
rs18157391166560624CT
rs5357638804371GG

The strand problem

DNA is double-stranded and the two strands are complementary, so the same variant can be reported as AG or as CT depending on which strand the company reported. Comparing your genotype to a published study without checking strand orientation will sometimes give you the exact opposite of the truth.

Most consumer files report on the plus strand of the reference, but this is a convention rather than a guarantee, and it is the single most common source of confidently wrong home interpretation.

What the file is not

It is not your genome sequence. A genotyping array tests several hundred thousand chosen positions — well under 0.03% of the roughly 3.1 billion base pairs. Everything not on the chip is simply absent, and absence in this file is not evidence of anything.

That matters most for rare variants. A clinically significant mutation that is not on the array will not appear, and a file showing nothing alarming is not a clean bill of health. Screening is what sequencing is for.

Array data has real error rates

Consumer arrays are accurate in aggregate and imperfect at any single position. For common variants that is fine. For rare ones it is a genuine problem: when the true frequency of a variant is very low, even a small false-positive rate means a large share of positive calls are wrong.

This is why third-party interpretation services returning alarming findings from raw data have been a recurring problem, and why any result with clinical weight needs confirmation on a clinical-grade test.

Reasonable things to do with it

Privacy is the part worth slowing down on

This file identifies you and, partially, everyone related to you — who never consented to anything. Uploading it to a third-party interpretation service means handing that over, and the terms governing what happens to it if the company is sold or breached are frequently not what people assume. Genetic data cannot be rotated like a password.

None of that argues against downloading and keeping your own file. It argues against uploading it casually.

SNPs explained · Genetic deep dive · Back to all articles