Before first-patient-in · simulate · check · transform

How dummy data becomes genius data.

Verdatic simulates realistic clinical study data from your study design, checks its quality as it's made, and writes the transformation programs your team will run when the real data arrives. The result: your data pipeline is built, tested, and ready before the study has a single patient in it.

Works entirely on synthetic data — no patient information ever enters Verdatic.

“Agentic AI Meets Data Simulation,” a talk on the approach behind Verdatic, is on the program at the CDISC US Interchange in Denver on October 6. The methodology draft for assessing synthetic data quality will open for comment at clinventive.com in October.

SAS · R · Pythongenerated programs, your choice of language
~5,000curated biomedical concepts & specializations beyond the published standard
45,000+controlled-terminology terms, kept current
3,500+CDISC CORE conformance rules on tap

Find the gaps before the trial does.

Conformance tooling is built to grade data that already exists. Verdatic simulates the study's data before first patient in, so the checks and programs that will run on it can be tested first.

A study can rehearse its data before real data arrive: edit checks on a clean twin and a defective twin, ADaM programs on data shaped like the study, the SDTM programs and Define-XML run against simulated data, vendor deliveries on the vendor's cadence, and unblinding risk with masked exports.

Poor synthetic data fails quietly: the check passes, the program runs, the file loads, and none of it would have happened on real data.

On-the-job training for people is fine, but not for data. Data simulation identifies quality and design gaps before they can negatively impact a live trial.
Verdatic learns what quality means from your own edit checks. Generate clean data for PASS test cases and a dirty twin for FAIL test cases — review both in Verdatic, then load into your CTDB to automate UAT.
Analysis datasets before the study goes live. Statistical programmers can draft and dry-run analysis dataset programs on simulated data shaped like the study, before real data arrive.

One connected workflow, start to finish.

From study design to submission-shaped datasets and the programs that produce them — every step feeds the next, and every decision stays on the record.

01

Start from your design

Import your study design from the file formats you already have — EDC design files or the USDM digital protocol format — or sketch a new design from scratch.

02

Simulate the study

Generate whole synthetic patient populations: treatment arms, visit schedules, labs, vitals, adverse events, and medications that all fit together sensibly.

03

Check quality as you go

Edit checks, quality rules, and the CDISC CORE conformance engine review the data — and the simulator learns from what they find.

04

Map with a helping hand

Smart suggestions link your fields to standard definitions and map columns to their target datasets — you confirm the calls, and the assistant gets better over time.

05

Generate the programs

Get runnable transformation programs in SAS, R, or Python, bundled with a build plan, a Define.xml, and a full record of where every value came from.

What you get.

Everything below ships in the product today.

Verdatic's generated-code panel for the AE dataset, with tabs for SAS data step, Python (pandas) and R (dplyr). The SAS tab shows libname statements, a data step, and attrib statements giving every variable its label.

SAS, R and Python from one specification

Verdatic generates complete, runnable transformation programs in any of the three from the same mapping. Python and R bundles can be run without leaving the app.

Systolic blood pressure for 22 subjects on the active arm across six visits: the whole band of trajectories drifts downward between day −14 and day 198, with a histogram below comparing the observed values against a normal curve.

Data that behaves like real patients

Synthetic data that meets physiological and therapeutic-area expectations.

Blood pressure responds to treatment on the active arm, adverse events start after dosing, and related medications follow, with coded adverse events and medications. You can even ask for deliberately messy data to exercise your edit checks.

The CDISC CORE findings worklist: 315 of 430 rules pass, none open, seven waived, 49 not applicable and 60 conditional, no change from the previous run — and a disposition recorded against each remaining rule.

Conformance issues, handled like real work

Identify, track, and resolve issues using our CDISC CORE Rules engine interface.

Verdatic runs the official CDISC CORE conformance engine and turns its findings into a worklist: see example records, record a decision on each issue, and watch run-over-run deltas show what's resolved and what persists — so nothing gets re-litigated.

A traceability view of the vital-signs dataset: the result variable expanded to show its five per-test source passes, then a table linking each test to a biomedical concept — diastolic blood pressure C25299, pulse C49676 — with every link marked AI proposed and Human confirmed at 92% confidence.

Every decision on the record

End-to-end traceability of specification and mapping decisions.

Each output value traces back through its mapping, its method, and its standard definition to the exact catalog snapshot it came from — with human decisions kept on file right next to the assistant's suggestions.

A build plan for the study: an 18-domain resolved emission order, and a code bundle set to SDTM and SAS with buttons to preview, generate the bundle, open Define.xml and open traceability — the packed ZIP is 481.4 KB with a sha256 checksum.

Programs, packaged properly

Generate, bundle, and version programs with a build plan and Define.xml.

Generated code ships as a complete, versioned bundle: a dependency-aware build plan, shared library functions, a Define.xml, and a data lineage manifest — checksummed and reproducible, ready to hand to another team.

The Global Library asset tree, with folders for ADaM standards, EDC standards, schedules and forms, and a Dictionaries folder holding 16 canonical dictionaries, each versioned and downloadable.

One library, every project

A global library manages canonical assets for your organization — copy and paste between projects.

Forms, dictionaries, methods, and other building blocks live in a shared, versioned library. Save an asset from one project and drop it into the next — no re-inventing, no drift.

Field-link review filtered to the lab form: 143 of 167 fields confirmed, 5 unresolved, 16 unlinked. Each row names the method that proposed the link and its confidence — 0.92 links confirmed, 0.65 and 0.90 links held back as unresolved for a person to settle.

AI suggestions with confidence scores

Assistants propose field links, mappings and draft logic with confidence scores. Field links below the confirmation threshold (0.90 unless you change it) wait for a person, and suggested mappings and drafted logic are never applied without one. Runs with a local, private model by default, so your study designs never have to leave your environment.

The study's visit schedule: six visits from Screening at day −14 to End of Study at day 198, each with its target day and the forms collected at it — vitals, labs, ECG, medical history, adverse events, medications and dispositions.

A sandbox for study ideas

Sandbox your study ideas: trial designs and output formats to see the possibilities.

Clone a design, add an arm, shift the visit schedule, change an output format — then regenerate and compare. Explore what a design choice does to the data before anyone commits to it.

The CDISC Catalog's active snapshot, locked and stamped with a source hash: 6,152 biomedical concepts (1,475 CDISC-published, 4,677 Verdatic-authored), 6,305 SDTM dataset specializations, 1,230 controlled-terminology codelists holding 45,919 terms, and 3,549 CORE conformance rules, over a row of pinned package versions.

A standards catalog that stays current

CDISC Catalog synchronization: ~5,000 user-authored Biomedical Concepts (source-attributed) and Specializations, curated to stay up-to-date with published standards.

The published CDISC library is cached in full — concepts, specializations, terminology, and conformance rules — and extended with thousands of curated, source-attributed additions that retire automatically when an official version supersedes them.

The Import / Create New Study Design screen, with tabs for USDM (JSON), Rave ALS (XLSX), Veeva (CDE JSON) in beta, a demo study and a dataset cloner. The Rave ALS tab is open, offering to build a project with schedules, forms, fields, dictionaries and edit checks from a spreadsheet.

Plays well with your other systems

Interoperable with EDC study design file formats and USDM — import one, export in another.

Bring designs in from Rave ALS workbooks, Veeva casebook design exports (beta), or USDM protocol files; send them back out as a round-tripped ALS or a USDM export. Datasets leave as SAS transport files or Dataset-JSON, spec work as Excel.

Verdatic accelerates the work and keeps the decisions on the record — it doesn't replace the people who own the study. Everything it generates is meant to be reviewed.

Screens from the product.

One synthetic study, end to end.

Two timeline lanes for one synthetic subject on a shared day axis running from D-25 to D209: an Adverse Events lane carrying nausea, dizziness and an ongoing rash, and a Concomitant Medications lane below it carrying metformin, ondansetron, meclozine and cetirizine.
One synthetic patient. Adverse events on the top lane; on the lane below, the medications given to treat them — both on the same day axis, so you can see each medication start on or after the event it answers. The one bar that spans the whole study started before day one: a medication the patient was already taking.
The same subject's visit strip — six visits from D-14 to D198, three of them highlighted in amber where an adverse event starts — above trend charts for diastolic blood pressure, pulse, respiration, systolic blood pressure and temperature.
The same patient's six visits, with the three where an adverse event starts picked out in amber — and the vital signs recorded at each one. Systolic and diastolic blood pressure both fall across the study; pulse, respiration and temperature wander the way real measurements do.
The Generate Simulated Data panel: number of sites and subjects per site, a mode selector offering Clean, Dirty, Paired and Test Case, a categorical output convention, a row sort order, and a global missing rate.
You say how big the run is — and whether the data comes out clean, deliberately dirty, paired, or built from your test cases.
Generated R code for the adverse-events dataset: library(dplyr), a build_ae function piping the input data frame into mutate(), with a comment above each derived column naming the method that produced it.
The generated R program for the adverse-events dataset — each derived column commented with the method behind it.
A code bundle that has been run: a Succeeded badge reading Python 3.12.10 with pandas and pyreadstat, exit 0; a CDISC CORE banner covering 430 rules with 7 issues and define.xml included; chips listing every output file as both .json and .xpt; and a preview grid of the AE dataset.
The bundle runs inside Verdatic and drops out every domain twice — as a SAS transport file and as Dataset-JSON — with the CDISC CORE result attached.

Want to open a full generated study package — datasets, Define.xml, conformance report? Request access and we'll walk you through one.

Realistic, conformant, traceable — what each one actually means here.

Every vendor uses these words. Here's the mechanism behind each one in Verdatic.

Realistic

Lab values sit in physiological ranges; blood pressure responds to treatment on the active arm and not the placebo arm; adverse events start after dosing and pull related medications behind them; terms are coded the way real dictionaries code them.

Conformant

The official CDISC CORE engine runs on the generated data, and findings become a worklist with a decision recorded on each one and deltas between runs.

Traceable

Every output value resolves through its mapping, its method, and its standard definition to the exact catalog snapshot it came from, with human decisions filed next to the assistant's suggestions.

Reproducible

Same inputs, same seed, same datasets, same programs, same Define.xml — checksummed, so another team can re-derive it.

Watch it work.

Six minutes, one connected workflow — from study design to checked synthetic data to the generated programs. It's Verdatic's entry to the 2026 CDISC AI Innovation Challenge.

What the challenge asked for, and what to look for in the video →

Built for the people who build study data.

Statistical programmers

Develop and test SDTM and ADaM conversion programs against realistic data long before first-patient-in — then run the same validated programs on the real thing.

Clinical data managers

Rehearse edit checks and review workflows on data that's realistic — or deliberately dirty on request — without waiting for a live study or touching patient information.

Biostatisticians

See analysis-ready datasets and draft outputs while the protocol is still settling, and test how design choices play out in the numbers.

Standards & governance teams

One curated catalog, source-attributed extensions, versioned snapshots, and an append-only audit trail — governance that's visible instead of assumed.

Prove it on one of your studies.

We work with a small number of sponsors, biotechs and CROs at a time. You bring a study design; we get Verdatic running on it, compare what comes out against how your team does it today, and shape what we build next around what you find. NDA up front. And because Verdatic only ever works on synthetic data, there's no patient information to protect in the first place.

Prefer to start smaller? Guided access — a walkthrough on your design without the benchmark — is the same form below.