Self-Maintaining, Automated Data Pipeline for Genomic Data at the Stanley Center for Psychiatric Research at the Broad Institute
The Stanley Center for Psychiatric Research at the Broad Institute used Manifold to replace a manual, multi-step genomic data workflow with a self-maintaining pipeline that automatically ingests whole-genome sequencing (WGS) manifests, phenotype files, and cohort metadata into a single queryable dataset. The pipeline publishes per-cohort QC dashboards and AI-ready, fully described tables, so researchers can get real-time answers on cohort status, sample QC, and experiment progress
At A Glance
- The Stanley Center receives large batches of genomic data deliveries (WGS manifest JSONs, phenotype Excel files, and cohort metadata txt files), all landing in S3.
- Because sample QC metrics, sequencing run metadata, and phenotype data sat in different formats and systems, manually bringing them into one queryable view took considerable time and recurring effort.
- Manifold built an automated, end-to-end pipeline that ingests each delivery, organizes it into one queryable dataset, publishes per-cohort QC dashboards, and registers every cohort as a discoverable dataset, all structured so researchers and AI agents can query it directly.
- The team went from manually splitting data on Terra and saving static files to GCP to a self-maintaining, automated data pipeline that delivers a richer joined data model easily queryable in one place.
- A set of AI skills built directly on the consistent and fully described data allow researchers to query the data in plain language and receive real-time answers about cohort status, QC metrics, sample tracking, and experiment progress.
Unifying each delivery was a manual, multi-step process
Each genomic data delivery arrives as a set of related but separate pieces: sequencing and QC metrics as manifest JSONS, phenotype data as spreadsheets, and cohort metadata as txt files. The prior workflow involved processing each type one-by-one, manually splitting the data on Terra and saving static files to GCP.
Such consolidation effort was time- and resource-intensive. Routine questions, like how many samples passed QC or which cohorts were complete, were answered by working through the source files. The process worked, but keeping a unified, updated view took significant time and recurring effort.

Automated end-to-end pipeline turns raw deliveries into agent-ready data
When a new batch of genomic data lands, the agent picks it up on its own, reads every file, and pulls the important details into one organized place. It only touches new files each time, so nothing is reprocessed and nothing is missed.
The agent is also capable of handling a brand-new data type, say from a different assay. Normally, such new data types require hand-building a new transformation pipeline—complex, manual work that demands a deep understanding of where the data comes from, how it's structured, and how it will be used. Instead, the agent reads the new data type, figures out its structure, and builds the pipeline to load it into the unified view.
Everything lands in a single master table that links sample quality scores, sequencing details, phenotype, and cohort information that used to live in separate files.
What makes this more than a tidy database is that every table is labeled so AI agents, not just people who already know the data, can read and query it. Each column carries a plain-language description that travels with the data wherever it goes.
Self-maintaining, unified source of truth, built for researchers and AI agents alike
One unified source of truth
Separated tables and files were pulled into a single organized data pipeline, ready for downstream analysis and for AI agents to access. Every column carries a plain-language description that is preserved everywhere the data goes, the layer of meaning that lets AI agents, not just the people who already know the data, interpret the tables.
Automated data pipeline saving time and money
Per-cohort tables are created and kept current automatically as new cohorts appear, with zero manual splitting. QC dashboards are published per cohort and replace Sigma, making the workflow faster to load with no licensing cost and always remains up to date.
AI-ready and agentic by design
A set of AI skills was built directly on these tables, answering questions about cohort status, QC metrics, sample tracking, and experiment progress. Because the data model is clean, consistent, and fully described, the AI skills query it reliably. And because the pipeline detects and adds new columns on its own as source files change, the data model stays current without engineering work, so the agent skills don't break when Broad adds a new field.
Empowering teams to do the work that matters
In addition to getting cohort raw data, the teams get enriched tables and visual tools for QC review, cohort exploration, and experiment tracking, all maintained automatically with every new delivery. The raw data is transformed into pre-aggregated answers to the questions researchers and stakeholders actually ask and structured specifically to be queried directly in plain language.
“It’s incredibly rewarding to see everything coming together in the thoughtful integration of compliance tracking, sequencing data, and phenotypic information. Pulling these components together has been a challenge for decades, and Manifold has helped unify them into a cohesive, intuitive framework."
— Christine Stevens, Director of Operations & Development, Daly Lab, The Stanley Center for Psychiatric Research at the Broad Institute
What's next
The pipeline already handles new cohorts and new source columns without engineering intervention, so the data model and the agent skills built on it stay current as deliveries evolve.
The next layer of value is in the AI skills built on this foundation: because the tables are clean, consistent, and fully described, the skill set can expand to answer more of the questions researchers and leadership ask without rebuilding the underlying data.