Before first-patient-in · simulate · check · transform

How dummy data becomes genius data.

Verdatic simulates realistic clinical study data from your study design, checks its quality as it's made, and writes the transformation programs your team will run when the real data arrives. The result: your data pipeline is built, tested, and ready before the study has a single patient in it.

Works entirely on synthetic data — no patient information ever enters Verdatic.

SAS · R · Pythongenerated programs, your choice of language
~5,000curated biomedical concepts & specializations beyond the published standard
45,000+controlled-terminology terms, kept current
3,500+CDISC CORE conformance rules on tap

Find the gaps before the trial does.

Everyone else automates the work after the data arrives. Verdatic does the work before.

Most study teams can't truly test their data pipeline until real patient data starts arriving — and by then, every surprise is expensive. Verdatic lets you rehearse the whole thing first, on data that looks and behaves like the real study.

On-the-job training for people is fine, but not for data. Data simulation identifies quality and design gaps before they can negatively impact a live trial.
Verdatic learns what quality means from your own edit checks. Generate clean data for PASS test cases and a dirty twin for FAIL test cases — review both in Verdatic, then load into your CTDB to automate UAT.
SAP outputs before the study goes live? See how Verdatic delivers. Because the simulated data is realistic and standards-aligned, statistical programmers can draft and test analysis datasets and outputs months early — so validated conversion programs are ready to run at first-data-in.

One connected workflow, start to finish.

From study design to submission-shaped datasets and the programs that produce them — every step feeds the next, and every decision stays on the record.

01

Start from your design

Import your study design from the file formats you already have — EDC design files or the USDM digital protocol format — or sketch a new design from scratch.

02

Simulate the study

Generate whole synthetic patient populations: treatment arms, visit schedules, labs, vitals, adverse events, and medications that all fit together sensibly.

03

Check quality as you go

Edit checks, quality rules, and the CDISC CORE conformance engine review the data — and the simulator learns from what they find.

04

Map with a helping hand

Smart suggestions link your fields to standard definitions and map columns to their target datasets — you confirm the calls, and the assistant gets better over time.

05

Generate the programs

Get runnable transformation programs in SAS, R, or Python, bundled with a build plan, a Define.xml, and a full record of where every value came from.

What you get.

Everything below ships in the product today — no roadmap promises, no fine print.

Verdatic's generated-code panel for the AE dataset, with tabs for SAS data step, Python (pandas) and R (dplyr). The SAS tab shows libname statements, a data step, and attrib statements giving every variable its label.

All three languages, not a choice of one

Why choose between SAS, R, or Python, when you can have all three?

Verdatic generates complete, runnable transformation programs in any of the three — same logic, same results, your team's preferred language. Python and R bundles can even be executed without leaving the app.

Systolic blood pressure for 22 subjects on the active arm across six visits: the whole band of trajectories drifts downward between day −14 and day 198, with a histogram below comparing the observed values against a normal curve.

Data that behaves like real patients

Synthetic data that meets physiological and therapeutic-area expectations.

Lab values stay in believable ranges, blood pressure responds to treatment on the active arm, adverse events start after dosing, and related medications follow — with medical coding that mirrors real dictionaries. You can even ask for deliberately messy data to exercise your edit checks.

The CDISC CORE findings worklist: 315 of 430 rules pass, none open, seven waived, 49 not applicable and 60 conditional, no change from the previous run — and a disposition recorded against each remaining rule.

Conformance issues, handled like real work

Identify, track, and resolve issues using our CDISC CORE Rules engine interface.

Verdatic runs the official CDISC CORE conformance engine and turns its findings into a worklist: see example records, record a decision on each issue, and watch run-over-run deltas show what's resolved and what persists — so nothing gets re-litigated.

A traceability view of the vital-signs dataset: the result variable expanded to show its five per-test source passes, then a table linking each test to a biomedical concept — diastolic blood pressure C25299, pulse C49676 — with every link marked AI proposed and Human confirmed at 92% confidence.

Every decision on the record

End-to-end traceability of specification and mapping decisions.

Each output value traces back through its mapping, its method, and its standard definition to the exact catalog snapshot it came from — with human decisions kept on file right next to the assistant's suggestions.

A build plan for the study: an 18-domain resolved emission order, and a code bundle set to SDTM and SAS with buttons to preview, generate the bundle, open Define.xml and open traceability — the packed ZIP is 481.4 KB with a sha256 checksum.

Programs, packaged properly

Generate, bundle, and version programs with a build plan and Define.xml.

Generated code ships as a complete, versioned bundle: a dependency-aware build plan, shared library functions, a Define.xml, and a data lineage manifest — checksummed and reproducible, ready to hand to another team.

The Global Library asset tree, with folders for ADaM standards, EDC standards, schedules and forms, and a Dictionaries folder holding 16 canonical dictionaries, each versioned and downloadable.

One library, every project

A global library manages canonical assets for your organization — copy and paste between projects.

Forms, dictionaries, methods, and other building blocks live in a shared, versioned library. Save an asset from one project and drop it into the next — no re-inventing, no drift.

Field-link review filtered to the lab form: 143 of 167 fields confirmed, 5 unresolved, 16 unlinked. Each row names the method that proposed the link and its confidence — 0.92 links confirmed, 0.65 and 0.90 links held back as unresolved for a person to settle.

AI that asks before it acts

AI agents learn as they assist human-in-the-loop decision-making, improving accuracy over time.

Assistants propose field links, mappings, and draft logic with confidence scores; people confirm anything uncertain, and those decisions become training signal. Runs with a local, private model by default — your study designs never have to leave your environment.

The study's visit schedule: six visits from Screening at day −14 to End of Study at day 198, each with its target day and the forms collected at it — vitals, labs, ECG, medical history, adverse events, medications and dispositions.

A sandbox for study ideas

Sandbox your study ideas: trial designs and output formats to see the possibilities.

Clone a design, add an arm, shift the visit schedule, change an output format — then regenerate and compare. Explore what a design choice does to the data before anyone commits to it.

The CDISC Catalog's active snapshot, locked and stamped with a source hash: 6,152 biomedical concepts (1,475 CDISC-published, 4,677 Verdatic-authored), 6,305 SDTM dataset specializations, 1,230 controlled-terminology codelists holding 45,919 terms, and 3,549 CORE conformance rules, over a row of pinned package versions.

A standards catalog that stays current

CDISC Catalog synchronization: ~5,000 user-authored Biomedical Concepts (source-attributed) and Specializations, curated to stay up-to-date with published standards.

The published CDISC library is cached in full — concepts, specializations, terminology, and conformance rules — and extended with thousands of curated, source-attributed additions that retire automatically when an official version supersedes them.

The Import / Create New Study Design screen, with tabs for USDM (JSON), Rave ALS (XLSX), Veeva (CDE JSON) in beta, a demo study and a dataset cloner. The Rave ALS tab is open, offering to build a project with schedules, forms, fields, dictionaries and edit checks from a spreadsheet.

Plays well with your other systems

Interoperable with EDC study design file formats and USDM — import one, export in another.

Bring designs in from Rave ALS workbooks, Veeva casebook design exports, or USDM protocol files; send them back out as a round-tripped ALS or a USDM export. Datasets leave as SAS transport files or Dataset-JSON, spec work as Excel.

Verdatic accelerates the work and keeps the decisions on the record — it doesn't replace the people who own the study. Everything it generates is meant to be reviewed.

See it, not just read about it.

Real screens from the product — one synthetic study, end to end.

Two timeline lanes for one synthetic subject on a shared day axis running from D-25 to D209: an Adverse Events lane carrying nausea, dizziness and an ongoing rash, and a Concomitant Medications lane below it carrying metformin, ondansetron, meclozine and cetirizine.
One synthetic patient. Adverse events on the top lane; on the lane below, the medications given to treat them — both on the same day axis, so you can see each medication start on or after the event it answers. The one bar that spans the whole study started before day one: a medication the patient was already taking.
The same subject's visit strip — six visits from D-14 to D198, three of them highlighted in amber where an adverse event starts — above trend charts for diastolic blood pressure, pulse, respiration, systolic blood pressure and temperature.
The same patient's six visits, with the three where an adverse event starts picked out in amber — and the vital signs recorded at each one. Systolic and diastolic blood pressure both fall across the study; pulse, respiration and temperature wander the way real measurements do.
The Generate Simulated Data panel: number of sites and subjects per site, a mode selector offering Clean, Dirty, Paired and Test Case, a categorical output convention, a row sort order, and a global missing rate.
You say how big the run is — and whether the data comes out clean, deliberately dirty, paired, or built from your test cases.
Generated R code for the adverse-events dataset: library(dplyr), a build_ae function piping the input data frame into mutate(), with a comment above each derived column naming the method that produced it.
The generated R program for the adverse-events dataset — each derived column commented with the method behind it.
A code bundle that has been run: a Succeeded badge reading Python 3.12.10 with pandas and pyreadstat, exit 0; a CDISC CORE banner covering 430 rules with 7 issues and define.xml included; chips listing every output file as both .json and .xpt; and a preview grid of the AE dataset.
The bundle runs inside Verdatic and drops out every domain twice — as a SAS transport file and as Dataset-JSON — with the CDISC CORE result attached.

Want to open a full generated study package — datasets, Define.xml, conformance report? Request access and we'll walk you through one.

Realistic, conformant, traceable — what each one actually means here.

Every vendor uses these words. Here's the mechanism behind each one in Verdatic.

Realistic

Lab values sit in physiological ranges; blood pressure responds to treatment on the active arm and not the placebo arm; adverse events start after dosing and pull related medications behind them; terms are coded the way real dictionaries code them.

Conformant

The official CDISC CORE engine runs on the generated data, and findings become a worklist with a decision recorded on each one and deltas between runs.

Traceable

Every output value resolves through its mapping, its method, and its standard definition to the exact catalog snapshot it came from, with human decisions filed next to the assistant's suggestions.

Reproducible

Same inputs, same seed, same datasets, same programs, same Define.xml — checksummed, so another team can re-derive it.

Watch it work.

Six minutes, one connected workflow — from study design to checked synthetic data to the generated programs. It's Verdatic's entry to the 2026 CDISC AI Innovation Challenge.

What the challenge asked for, and what to look for in the video →

Built for the people who build study data.

Statistical programmers

Develop and test SDTM and ADaM conversion programs against realistic data long before first-patient-in — then run the same validated programs on the real thing.

Clinical data managers

Rehearse edit checks and review workflows on data that's realistic — or deliberately dirty on request — without waiting for a live study or touching patient information.

Biostatisticians

See analysis-ready datasets and draft outputs while the protocol is still settling, and test how design choices play out in the numbers.

Standards & governance teams

One curated catalog, source-attributed extensions, versioned snapshots, and an append-only audit trail — governance that's visible instead of assumed.

The approach behind Verdatic's self-correcting simulation was selected for presentation at the 2026 CDISC Interchange.

Prove it on one of your studies.

We work with a small number of sponsors, biotechs and CROs at a time. You bring a study design; we get Verdatic running on it, compare what comes out against how your team does it today, and shape what we build next around what you find. NDA up front. And because Verdatic only ever works on synthetic data, there's no patient information to protect in the first place.

Prefer to start smaller? Guided access — a walkthrough on your design without the benchmark — is the same form below.