The Double Helix Data Problem: Who Really Owns Your Genetic Information After You Spit in That Tube
Photo: DNA genetic testing privacy data security laboratory, via cdn.tuko.co.ke
Somewhere in a climate-controlled data center, a file exists that contains information more uniquely identifying than your fingerprints, more medically revealing than a decade of physician records, and more immutable than any password you will ever set. It was created from a saliva sample you mailed in a small plastic tube, in exchange for a pie chart showing your ancestral origins. You almost certainly agreed to its collection and storage without reading the relevant provisions. You may have no practical mechanism to delete it.
This is the current state of genetic privacy in the United States, and it is considerably more consequential than the consumer DNA testing industry's marketing materials suggest.
The Scale of the Marketplace
Direct-to-consumer (DTC) genetic testing has become a mainstream consumer activity. Companies including AncestryDNA, 23andMe, MyHeritage, and several smaller competitors have collectively processed the DNA of an estimated 30 to 40 million Americans. The appeal is genuine: these platforms offer real genealogical insights, can connect biological relatives who were previously unknown to each other, and in some cases surface medically actionable information about inherited health risks.
The business model underlying these services, however, is not simply the sale of a testing kit. Genetic data — aggregated, anonymized, and licensed — represents a significant revenue stream. Pharmaceutical companies pay handsomely for access to large genetic databases because the research value of correlating genetic variants with health outcomes across millions of individuals is enormous. GlaxoSmithKline's reported $300 million investment in 23andMe, structured around access to the company's genetic database for drug discovery research, offered a rare public glimpse into what that data is actually worth.
What You Agreed To
The terms-of-service and privacy policy documents governing DTC genetic testing platforms are, by any reasonable measure, among the most consequential agreements an ordinary consumer will ever encounter. They are also among the least read.
A typical policy grants the testing company a broad license to use your genetic data for internal research, product development, and — depending on the platform and the consent options you selected — participation in third-party research partnerships. Some platforms offer an explicit opt-in to research data sharing; others structure the opt-out in ways that require affirmative action from users who may not realize a default exists.
Critically, the definition of "anonymized" data in this context warrants scrutiny. Research published in the journal Science demonstrated more than a decade ago that individuals could be re-identified from supposedly anonymous genetic datasets when that data was cross-referenced with other publicly available information. Genetic data is, by its nature, resistant to true anonymization — it is the biological equivalent of a unique identifier that cannot be changed.
Law Enforcement Access: A Legal Gray Zone
Perhaps the most consequential and least publicly understood dimension of genetic privacy involves law enforcement access to genetic databases. Two distinct mechanisms are in operation.
The first is direct legal process. Law enforcement agencies can and do serve subpoenas, court orders, and warrants on DTC genetic testing companies seeking data related to specific individuals. The companies' policies on how they respond to such requests vary. 23andMe, for instance, has historically stated that it requires valid legal process before disclosing user data to law enforcement and notifies affected users when legally permissible. AncestryDNA has published similar commitments. Whether those commitments are consistently honored, and what constitutes "valid legal process" in practice, is not always transparent.
The second mechanism is more novel and far less regulated: investigative genetic genealogy, sometimes called forensic genealogy. This technique involves uploading DNA profiles derived from crime scene evidence to consumer-oriented databases — most prominently GEDmatch and FamilyTreeDNA — and using the resulting relative matches to build family trees that ultimately identify a suspect. The technique gained widespread public attention through its role in identifying the Golden State Killer in 2018.
Forensic genealogy has since been used in hundreds of criminal investigations across the United States. Its efficacy is not in dispute. Its implications for the privacy of individuals who never consented to their DNA being used in criminal investigations — because a relative submitted their data to a consumer platform — are profound and largely unresolved. When you upload your DNA to a consumer database, you are, in effect, making a partial privacy decision on behalf of every biological relative you have.
The 23andMe Bankruptcy: A Case Study in Data Vulnerability
The fragility of corporate data stewardship commitments became starkly apparent in early 2025, when 23andMe filed for bankruptcy protection. The filing immediately raised an urgent question that regulators, privacy advocates, and the company's approximately 15 million customers were forced to confront simultaneously: what happens to a genetic database when the company that holds it becomes insolvent?
Genetic data is a corporate asset. In bankruptcy proceedings, assets are subject to sale. The privacy policies that governed data collection at the time of consent may or may not be binding on a subsequent acquirer. Several state attorneys general, including those in California and New York, issued guidance urging 23andMe customers to delete their data before any acquisition was finalized. The episode illustrated, with unusual clarity, the gap between the privacy assurances consumers receive at the point of data collection and the protections that actually persist over time.
The Regulatory Landscape
Federal law governing genetic privacy is fragmented. The Genetic Information Nondiscrimination Act (GINA) of 2008 prohibits the use of genetic information in health insurance underwriting and employment decisions — but it does not apply to life insurance, disability insurance, or long-term care insurance, sectors where genetic data could plausibly influence underwriting in ways that disadvantage consumers.
The Health Insurance Portability and Accountability Act (HIPAA) governs health data held by covered entities — primarily healthcare providers and their business associates. DTC genetic testing companies, operating outside the clinical healthcare system, are generally not HIPAA-covered entities. The data they hold exists in a regulatory space that is, at present, substantially less protected than the data in your physician's electronic health record.
Several states have enacted their own genetic privacy statutes. California, Texas, and Illinois are among those with laws specifically addressing DTC genetic data. A patchwork of state-level protections, however, offers inconsistent coverage for a national consumer base.
Practical Guidance for the Cautious Consumer
For individuals who have already submitted DNA samples to a consumer testing platform, the most immediate steps are reviewing and adjusting data-sharing consent settings within the platform's privacy controls, understanding the platform's law enforcement response policy, and considering whether to request data deletion if the platform offers that option.
For those considering a test, the decision warrants the same deliberation one might apply to any significant financial or medical choice. Reading the privacy policy — specifically the sections addressing research data sharing, third-party access, and law enforcement requests — before submitting a sample is advisable. Understanding that the decision affects biological relatives who have made no such choice is equally important.
Genetic information is not simply personal data. It is generational data, familial data, and in some respects, data that belongs to people who do not yet exist. The consumer marketplace has moved considerably faster than the legal framework designed to protect it, and the resulting gap is one that Americans are only beginning to fully reckon with.