Atlanta/ Science, Tech & Medicine

Chapel Hill Team Lands $35 Million to Build World's Largest Rare Disease AI Database

AI Assisted Icon
Published on September 01, 2026
Chapel Hill Team Lands $35 Million to Build World's Largest Rare Disease AI DatabaseSource: Wikipedia/Yeungb, CC BY 3.0, via Wikimedia Commons

UNC School of Medicine and Emory University have been awarded up to $35 million to lead a first-of-its-kind effort to build the world's largest data resource for rare disease artificial intelligence, a project designed to shrink diagnostic delays that can stretch on for years and leave patients bouncing between a dozen specialists before getting answers. The federal award, announced August 31, will fund a four-and-a-half-year initiative based in Chapel Hill that combines health records, insurance claims, medical imaging, video technologies and patient surveys into a single searchable resource.

The money comes from the Advanced Research Projects Agency for Health, known as ARPA-H, through its Rare Disease AI/ML for Precision Integrated Diagnostics program, or RAPID, according to UNC Health. ARPA-H is an agency within the U.S. Department of Health and Human Services, established by Congress under Public Law 117-103 in March 2022 to fund high-risk, high-reward biomedical innovation modeled after DARPA, according to the AAMC. The UNC and Emory award is one piece of a larger $98.5 million funding package ARPA-H announced the same day across multiple RAPID contracts.

Rare disease patients in the United States endure an average diagnostic odyssey lasting six to eight years and see between eight and 12 specialists before receiving an accurate diagnosis, according to MedCentral. UNC Health's announcement notes that delayed, incomplete or incorrect diagnoses can lead to inappropriate treatment, irreversible disease progression and excessive medical costs. Rare diseases affect as many as one in 10 Americans and an estimated 350 million people worldwide, spanning more than 10,000 distinct disease types, per the same account.

A Chapel Hill Scientist's Path to This Moment

Lead scientist Dr. Melissa Haendel, the Sarah Graham Kenan Distinguished Professor in UNC School of Medicine's Department of Genetics, joined UNC in May 2024 as Director of Precision Health and Translational Informatics. She previously led development of Mondo, the world's first standardized rare disease coding ontology tool, work highlighted by the UNC School of Medicine Department of Genetics. Haendel said the project aims to create a large-scale dataset spanning roughly 2,700 rare diseases, per the UNC Health report.

Her co-lead, Dr. Richard Moffitt, is an associate professor in Emory University School of Medicine's Department of Hematology and Medical Oncology. Moffitt said rare disease patient information is not centrally available today because each individual condition affects relatively few patients, making it hard for any single hospital system to gather enough data on its own. The two researchers previously co-led electronic health record data harmonization across multi-institution cohorts for the NIH National COVID Cohort Collaborative, giving them an established track record working across disparate hospital systems, according to research posted on medRxiv.

Part of a Larger $98.5 Million Federal Push

UNC and Emory are not working alone. The same RAPID initiative is funding Seattle-based Sage Bionetworks with an award of up to $28.2 million, starting September 1, to build a Rare Disease Data Commons that will evaluate and benchmark new AI diagnostic approaches, according to ARPA-H. San Francisco-based Probably Genetic separately received up to $10 million to gather patient- and caregiver-reported real-world data directly, using online AI algorithms to identify undiagnosed rare disease patients, per BioSpace.

The UNC-Emory project draws on expertise in genetics, biomedical informatics and rare disease research, and its collaborators include AcrossHealthcare, Combined Brain, DartNet, Datavant, EB Research, Global Genes, Johns Hopkins University, Mendelian.co, NORD, Queen Mary University of London, Truveta, UCSF and the University of Iowa. The Monarch Initiative and the All of Us Center for Linkage and Acquisition of Data will also be involved, according to the UNC Health announcement. Additional program support will come from OpenAI, Anthropic, Amazon Web Services and Google.

Why Rare Disease Data Stays Scattered

Rare disease data is often scattered, inconsistent or incomplete, in part because limited access to specialty care, geographic distance from rare disease experts, and insurance barriers all hinder collection of high-quality data, the UNC Health report notes. Fewer than 5% of the 7,000 to 10,000 identified rare diseases currently have an FDA-approved treatment, according to the Orphanet Journal of Rare Diseases. Under the federal Orphan Drug Act passed in 1983, a rare disease is legally defined as any condition affecting fewer than 200,000 people in the United States, per the FDA.

The economic stakes are steep. Rare diseases generate nearly $1 trillion in annual economic burden in the United States, with per-patient medical costs running nearly ten times higher than for non-rare diseases because of delayed diagnoses and emergency interventions, according to research published by Genomes2People. About half of all rare disease patients in the U.S. are children, the same research notes.

What the Resource Is Meant to Do

UNC and Emory teams will acquire data from patient registries and other real-world sources, then strip names and other identifying information before it enters the RAPID resource. The finished dataset is meant to give physicians and researchers new capabilities to identify patterns, support clinical decisions and accelerate rare disease research, while also helping design stronger clinical trials and identify trial participants more efficiently.

Clinical decision support tools built from the resource are intended to help diagnose patients in less sophisticated care settings or move patients along more quickly in their care journey, and the large-scale dataset will train models for deployment in settings that have less extensive data of their own. The initiative also plans to expand access to clinical care more broadly, and qualified physicians and researchers worldwide will eventually be able to apply for access to the resource, with some datasets available publicly and others requiring data use agreements and additional safeguards under a tiered secure-access system based on data sensitivity.