Use case · Blinded data and programs

Masking potentially unblinding data — planned, executed, and provable.

Blinded studies still have to move data: programmers build pipelines, data managers review listings, vendors get extracts. Verdatic masks by declaration — you write down what unblinds and how, and that plan is applied to the datasets, the exports, and the programs themselves, with a record at the end that shows what was done and who accepted what.

The problem

A blinded study still has to move data around. Statistical programmers need to build and test pipelines, data managers need to review listings, vendors need extracts — and none of them should be able to work out who is on active treatment. The variables that would give it away are the potentially unblinding ones, and handling them is usually somebody’s script, written under time pressure, and not quite the same script next time.

  • The obvious columns get masked and the signal survives anyway. Treatment variables are blanked, but the standard-units twin of a masked result still carries the value, or a reference-range flag does, or a supplemental qualifier, or a free-text comment, or a derivation that reaches around the mask entirely.
  • Masking everything destroys the reason for masking. A blinded copy full of nulls can’t exercise the analysis code, so the dry run that justified the exercise gets deferred to real data after all.
  • The programs are a second copy of the problem. Even a properly masked dataset is no use if the transformation programs handed to a partner reconstruct the signal from source.
  • Nobody can show afterwards what was done. Which variables were masked, by which method, what residual risk was accepted and by whom — reconstructed later from memory and file dates.

Verdatic takes the opposite approach: masking by declaration. You write down which variables carry unblinding signal and how each should be handled, and that one declaration is applied everywhere the data goes — the generated datasets, the exports, the transformation programs, and a standalone program you can run on a raw vendor extract with Verdatic nowhere in the loop.

Risk and cost profile

We don’t publish savings percentages, and there is no benchmark study behind this page. One of these rows genuinely resists pricing — and we’d rather say so than invent a figure.

The consequence of a leak isn’t a line item. ICH E9 treats the blind as a design feature to be maintained, with any unblinding documented and explained; in practice a suspected leak means an investigation, a documented deviation, potentially removing people from analysis roles, and a conversation about the integrity of the analysis at the worst possible moment in a study. Nobody can put a defensible average on that, so we won’t.

Cost driverHow it shows upPrice it with your own numbers
A leak The masked copy still carries the signal — in a companion variable, a flag, a comment, a derivation. Not priceable. Investigation, documented deviation, and a question mark over the analysis
Over-masking Nulls everywhere; the blinded copy can’t exercise the analysis it was made for. weeks of analysis programming displaced to after unblinding
One-off masking scripts Written per request by whoever is free; reviewed each time; never quite identical. hours per extract × extracts per study, plus the review of each
Proving it afterwards Reconstructing what was masked, why, and who accepted what. hours to assemble the evidence, per audit or per handover
Outbound extracts Every vendor, partner and reviewer file needs the same treatment, consistently. extracts per month × hours per extract, including the checking

What we won’t claim: that this is a regulatory blind-maintenance system, or that masked output is submission material. It is synthetic development data, the blinding decisions remain yours to own and approve, and Verdatic’s job is to make them explicit, apply them everywhere, and print the record.

The workflow, at a glance

Two things make this different from masking after the fact. Verdatic usually created the signal, so it knows where the signal is — if you built a treatment effect into systolic blood pressure, it can tell you that variable is unblinding rather than leaving you to guess. And it can rebuild the study without the effect: one method re-runs the generator with the treatment difference forced to zero under the same seed lineage, producing a coherent blinded twin rather than scrambled data.

Masking potentially unblinding data, from plan to proof Declare the plan which variables carry the signal Choose a method nine of them, from nulling to a blinded twin Leak analysis six tiers of checks for signal that survived High severity? Outputs data, programs and the report Nothing masked leaves while a high-severity finding stands. You add a rule, or you waive it with a written reason — recorded, revocable, and printed in the report. There is no override switch.
The gate is the whole design. An unblinding risk is either handled or consciously accepted by a named person — there is no third option and no way around it.

The Verdatic workflow

Masking is off until you say the study needs it. Once it’s on, a masked badge follows the project everywhere — on every tab and on its card in the project list — so nobody forgets which study they’re in. The masking screens below all come from one real plan on our own demo study: nine candidates, eight rules, one masked run.

STEP 01

Start from what Verdatic already knows unblinds

Candidates come from two independent places. Structural ones are unblinding by nature in any study — the treatment-arm variables, exposure and drug-accountability domains, arm-varying trial design — and are proposed only for surfaces your study actually has. Overlay candidates come from your own design: any field whose per-arm settings differ from the base is flagged, and the provenance line shows the exact per-arm difference that made it unblinding. They are suggestions, not rules; nothing is persisted until you accept, and dismissing one is recorded with a reason so it doesn’t resurface for the next person.

Systolic blood pressure for 22 subjects on the active arm across six visits: the whole band of trajectories drifts downward between day −14 and day 198, with a histogram below comparing the observed values against a normal curve.
This is the kind of signal masking has to remove: systolic blood pressure falling across the active arm’s 22 subjects. Because Verdatic built that difference, it can tell you the variable is unblinding — you don’t have to work it out from the data.
The potentially-unblinding-data candidate list: nine proposed candidates — seven structural (DM.ARM, DM.ARMCD, DM.ACTARM, DM.ACTARMCD, EX, TA and TV) and two arm-overlay ones on systolic and diastolic blood pressure — each with a provenance line, a method dropdown, and Accept, Accept and edit, and Dismiss buttons.
Nine candidates on this study: seven structural, and two that came out of its own design. The provenance column names the exact per-arm difference that made each one unblinding — Active: drift −3/visit vs base 0 for systolic. They are proposals: nothing is persisted until you accept one.
STEP 02

Write the rule: a scope and a method

Scope narrows in four steps — domain, then variable (or the whole domain), then test code where it applies, then visits. Every picker is filled from your own study’s metadata, so you can’t scope a rule onto something that isn’t there. Where two rules overlap, the more specific one wins, which lets you set a broad default and carve out exceptions.

The Add masking rule dialog with scope narrowed four levels deep — domain VS, variable VSSTRESN, test code SYSBP — and the Visits picker open, listing the study's own six visits from V1 Screening through V6 End of Study.
Scope, four levels deep: domain VS, variable VSSTRESN, test code SYSBP, and the visit list — which is this study’s own schedule, V1 Screening through V6 End of Study. Leave visits blank and the rule covers all of them, or type an expression when a range is easier than a list.
The Accept candidate dialog for the whole-domain EX candidate, opened pre-filled and narrowed to the variable EXDOSE, with a derived-candidate provenance banner above it, the method help for M0 below, and that same provenance carried into the Rationale field.
Accept & edit on the whole-domain EX.* suggestion: it opens pre-filled, and here it has been narrowed to EX.EXDOSE before it becomes a rule. The provenance that made it a candidate is carried into the rationale, so the reason travels with the rule instead of staying in somebody’s head.
The finished masking plan: the candidate list above with COVERED chips on six of its rows, and below it a table of eight active rules with their scope, visits, method, origin and rationale — four M6 rules on the arm variables, one M0 on EX.EXDOSE, and three M4 engine-only rules on the vital-signs results. Header chips read 8 active rules and plan 8c253d327930.
The plan is the rule set: eight active rules, each with its scope, method, origin — auto-derived or hand-written — and its rationale, and the whole plan identified by a hash. Candidates already covered are chipped, so what is still open stays visible. The method column is what the next step is about.
STEP 03

Pick the method that keeps what you still need

Nine methods, because “mask it” means nine different things depending on what has to survive. Blank the value outright. Redraw it from the study’s realism library so it still looks clinically sensible. Redraw from the pooled distribution of what’s actually there. Permute which subject each value belongs to, so aggregates tie out exactly. Quantile-map the arms onto each other so within-subject ordering survives. Recode categoricals to neutral, standards-valid terms so listings and conformance don’t break. Jitter dates by a per-subject offset that is consistent across every dataset. Coarsen a value — round it, threshold it, bin it — when a blinded reader legitimately needs some of it.

And the one no post-hoc tool can do: regenerate the scoped values with the treatment effect forced to zero, under the same seed lineage, so everything not driven by treatment stays identical and only the arm signal disappears. That produces a study a statistician can actually work with — not scrambled data, but a coherent study without a treatment difference. It needs the generative model, so it is the one method that doesn’t work on imported real data, and generated programs refuse it rather than substituting something weaker.

When a method can’t be applied — a stratum too thin, a scope that resolves to nothing — Verdatic falls back to blanking the value and counts it as a downgrade in the report. It never quietly leaves data unmasked.

The method help for M4, the blinded twin: an engine-only badge, what it does — re-simulate the scoped cells with the per-arm drift forced to zero under the same seed lineage — when to use it, and a caveat that emitted program sets refuse a plan carrying an M4 rule rather than substituting a weaker method. Below it, the rationale the author typed for this rule.
Each method carries its own help — what it does, when to reach for it, and what it costs you. M4, the blinded twin, is the one that needs the generative model, which is why an emitted program set refuses a plan carrying it rather than quietly downgrading to something weaker.
The masked run's demographics preview: the ARMCD, ARM, ACTARM and ACTARMCD columns are flagged and every row reads BLINDED, while race, ethnicity and the date columns beside them are untouched. A banner above reads: masked run — values in plan scopes carry no arm signal, 802 cells masked, 0 downgrades, position PostGeneration.
The neutral-categorical method in the data. All four arm variables read BLINDED — a term the standard still accepts, so listings and conformance checks don’t fall over — while race, ethnicity and the dates next to them are untouched. The banner stamps the run with its plan hash, 802 masked cells and 0 downgrades.
The same masked run's vital-signs preview: the systolic and diastolic columns are flagged as masked while pulse beside them is not. The first subject's systolic values read 102.1, 106.14, 109.23, 113.09, 111.81 and 109.52 across visits V1 to V6.
And the blinded twin in the data. Systolic and diastolic were regenerated with the treatment effect forced to zero; pulse beside them was never in scope. The first subject now runs 102.1 → 106.14 → 109.23 → 113.09 → 111.81 → 109.52 across the six visits — the downward drift is gone, but it still reads as a real longitudinal study rather than scrambled numbers.
STEP 04

Run the leak analysis, and be told what you missed

Masking one variable rarely finishes the job. The leak analysis looks for the signal everywhere else it could have survived, in six tiers: companion variables (the standard-units twin, the reference-range flag, the supplemental qualifier), mechanism correlates that move with what you masked, structural exposure of arm identity or arm-varying design, free-text fields, mapped targets that no rule covers, and the source graph — edit checks, derivations and panel siblings that reach around the mask.

High-severity findings block masked output. You cannot export a masked dataset, download a masked program bundle or emit a standalone program while one stands unwaived, and there is no override. You fix it with a rule, or you waive it with a written rationale that is recorded, audited, revocable and printed in the report.

A leak scan on a plan that is still blocked: a banner reading masked deliverables are blocked, 13 unwaived high-severity findings, and that the paired export, the masked bundle and standalone-program emission all refuse until each one is masked by a rule or waived with a rationale. Below it the six tier chips, and the head of the open queue — four tier-one companion-variable findings where the character and standard-units twins reconstruct the masked result.
A live plan, still blocked — which is the gate doing its job, not a finished state. The four findings at the head of the queue are tier one: the character and standard-units twins of the masked systolic and diastolic results reconstruct exactly the value the mask conceals. Each one gets a rule or a written waiver; there is no third door.
STEP 05

Produce the masked artifacts — including the programs

A masked generation run records its plan hash, seed, masked-cell count and downgrade count. A paired export ships both variants in one package, with the unmasked side explicitly labeled. A masked program bundle contains your SAS, R and Python with the masking applied in the programs, alongside the plan and the report — and the ordinary bundle is watermarked as unmasked so the two can never be confused. A standalone program masks a raw extract with Verdatic nowhere in the loop: it logs every downgrade and stops outright if the file it’s given doesn’t match the schema it was built for, rather than passing data through unmasked.

The masking Outputs panel: run #147 stamped MASKED with its plan hash, 802 masked cells, 0 downgrades and position PostGeneration, above a list of earlier unmasked runs each offering a paired masked/unmasked export; a dual code bundles row with a Masked bundle button and an Unmasked plus PUBD warnings button; and a standalone masking program generator.
What one plan produces. A masked run carrying its own hash, cell count and downgrade count. A paired export that ships both variants in a single package. Two labelled program sets from the same specification — masked, and unmasked with the unblinding variables flagged. And a standalone program that masks a raw extract with Verdatic nowhere in the loop.
Verdatic's generated-code panel for the AE dataset, with tabs for SAS data step, Python (pandas) and R (dplyr). The SAS tab shows libname statements, a data step, and attrib statements giving every variable its label.
The programs that carry the masking. In a masked bundle these same SAS, R and Python programs apply the plan themselves — and the R and Python masking libraries were executed against the in-app masker and matched it cell for cell.
STEP 06

Hand over the record, not a reassurance

The Masking Report is the blind-maintenance record for the exercise: the plan, every decision, every waiver with its rationale and author, coverage, and the downgrade counts. It prints. A rule that matches nothing is treated as a problem rather than a no-op — the report shows coverage, so a typo in a scope can’t silently leave data exposed.

You don’t need a simulation to use any of this. Point Verdatic at a raw central-lab extract or an EDC export, let it infer the schema, write rules against those columns, and emit a standalone program. If you later map that study to SDTM, the rules you already wrote carry forward to the mapped targets — and the leak analysis tells you if the mapping created a target that no rule covers.

Executive summary

  • The problem. Blinded studies still move data. Masking is usually a one-off script, the obvious columns get handled, and the signal survives in a companion variable, a comment or a derivation.
  • The cost. A leak isn’t a line item — it’s an investigation and a question over the analysis. Over-masking has a price too: a blinded copy of nulls can’t exercise the analysis it was made for.
  • What Verdatic does. You declare what unblinds and how; that one plan is applied to the data, the exports, the generated programs, and a standalone program for files Verdatic never touched.
  • Nine methods, from blanking through neutral categoricals and date jitter to regenerating the study with the treatment effect forced to zero — a coherent blinded twin a statistician can work with.
  • A gate, not a checkbox. Six tiers of leak analysis; high-severity findings block masked output until they’re fixed or waived with a written reason; every waiver is audited and printed in the Masking Report.

Bring the extract you’re nervous about.

The fastest way to judge this is to point it at a file you already mask by hand and compare what the leak analysis finds against what your script covers.

Request access

Or email hello@verdatic.com. For how access, audit and reproducibility are handled underneath, see Trust & security.

The other three use cases