# A synthetic twin of every restricted cancer dataset for code development

Source: https://onco.cc/ideas/idea-data-synthetic-companion-datasets/  
OnCo record `idea-data-synthetic-companion-datasets` (Idea). Data CC BY-NC 4.0, attribute "Data from OnCo (onco.cc)"; commercial use needs a licence.

## TL;DR

Publish a fake but realistic copy of each secure cancer dataset so researchers can write and test their code at home, then run the finished code on the real data.

## Summary

Access to secure datasets typically takes months; analysts then waste TRE time debugging. Synthetic datasets with the same schema and approximate joint distributions (for example the Simulacrum, built from the English cancer registry by Health Data Insight) let code be written and unit-tested outside the enclave. The proposal makes a validated synthetic companion mandatory for every dataset in a national cancer data space, with fidelity and privacy metrics published.

## Fields

- Kind: Idea
- Last checked: 2026-09-08
- Hypothesis: Providing a synthetic companion dataset will reduce median secure-environment compute time per project by a third and the number of failed code submissions by half.
- Rationale: The Simulacrum has been used to develop analyses that later ran unchanged on real registry data; the pattern is proven but not systematic.
- Proposed test: Measure TRE time and code-failure rate for projects with and without access to a synthetic companion in one national TRE over 18 months.
- Maturity: early-clinical
- Actor: data

## Sources

- The Simulacrum (Health Data Insight): https://simulacrum.healthdatainsight.org.uk/

## Connected records

- fronts: [AI & Computation](https://onco.cc/fronts/ai-computation/)
- bottlenecks: [Data silos](https://onco.cc/bottlenecks/b-data-silos/)

---
JSON: https://onco.cc/api/v1/entities/idea-data-synthetic-companion-datasets.json