design-study
Use when checking a radiology or medical AI study design before drafting or submission. Reviews the analysis unit, cohort logic, leakage risks, comparator, validation strategy and reporting-guideline fit. AI-vs-expert benchmarks are /design-ai-benchmarking.
How do I install this agent skill?
npx skills add https://github.com/aperivue/medsci-skills --skill design-studyIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill provides a framework for reviewing medical and radiology AI study designs, identifying validity risks like information leakage and survivorship bias. It includes local Python scripts and shell utilities for causal graph (DAG) analysis to assist with covariate selection. No malicious patterns, obfuscation, or unauthorized network operations were detected.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Design-Study Skill
Core Review Questions
Always inspect:
- The exact research question.
- The analysis unit: patient, lesion, exam, study, phase, report.
- The index date or decision point.
- How inclusion and exclusion criteria are applied.
- Any information leakage.
- The reference standard or endpoint definition.
- A clinically meaningful comparator.
- The validation strategy.
- The uncertainty reporting required.
- The best-fitting reporting guideline.
- Whether exposure/outcome/covariate definitions are literature-grounded or invented ad hoc from the data dictionary. If ad hoc, defer to
/define-variablesbefore drafting Methods.
Standard Output
## Study Design Review
Question: ...
Study type: ...
Analysis unit: ...
Index date / prediction timepoint: ...
### Strengths
- ...
### Major validity risks
1. ...
2. ...
### Minimal fixes
- ...
### Reporting fit
- Recommended guideline: ...
### Decision
- Ready for analysis / Needs redesign / Drafting can proceed with limitations
Cite a reference only with a /search-lit-confirmed DOI or PMID; mark any other [UNVERIFIED - NEEDS MANUAL CHECK]. Never invent a clinical definition, diagnostic criterion, or guideline recommendation: flag an unconfirmed one [VERIFY] and ask the user.
Workflow
Phase 1: Reconstruct the study
Extract from the protocol, draft, slides, tables, or notes: clinical problem, intended use case, population, inputs, outputs, outcome definition, and timing of variable availability.
Gate: Present the reconstructed study summary (question, analysis unit, intended use) to the user and confirm it before proceeding — a wrong reconstruction misdirects the entire validity review.
Phase 2: Check structural validity
Read a reference only when its condition holds:
| File | Read it when |
|---|---|
references/dag_adjustment.md | confounding control needs an explicit adjustment set |
references/target_trial_emulation.md | the design emulates a target trial |
references/venue_accept_recipe.md | a clinical DL / AI-validation study must decide which venue tier the achievable design can be accepted at, and the one design move that reaches the tier above (the bridge into /find-journal); skip when there is no publication-tier decision |
references/combine_models_ablation_design.md | the model is built by combining / adapting / fine-tuning existing models (nnU-Net, TotalSegmentator, SAM/MedSAM, a pretrained backbone) and the comparator must be designed as an ablation; skip for a model trained de novo |
references/multi_model_comparison_design.md | the contribution is comparing several models / architectures head-to-head and the comparison must be fair; skip for a single-model study (one model's ablation → the row above; AI-vs-human → /design-ai-benchmarking) |
references/segmentation_failure_characterization_design.md | the claim is that a segmentation model is clinically usable, not that it scores well; skip when the endpoint is benchmark accuracy (metric choice → /model-assessment; abstention / risk–coverage → /model-assessment) |
A. Analysis unit
Look for mismatches: a patient-level claim from lesion-level analysis; an exam-level split with patient overlap; phase-level samples treated as independent.
B. Leakage
Look for:
- postoperative features used for preoperative prediction
- normalization or thresholding performed before the data split
- repeated exams across train/test
- reader annotations derived from outcome information
- input-text contamination for NLP/LLM extraction tasks: if the input includes clinical history, indication, impression, prior diagnosis, or referral text, confirm it does not name or strongly imply the target label. If it does, the task is information retrieval under label leakage, not phenotype inference: redesign the input mask, report a sensitivity analysis excluding those fields, or reframe.
- construct dependence (a predictor that is a definitional component of the outcome): (i) an input that computes the outcome (fasting insulin and glucose when the outcome is HOMA-IR), or (ii) a near-tautological composite of the outcome's defining components (inflated, near-circular association). Test: "could this predictor be derived, in whole or part, from the outcome's definition or the same measurement?" If yes, exclude it, or keep it only as a labeled calibration probe.
F. Time origin & survivorship (incident / transition models)
For any time-to-event or incident/transition design, check before drafting:
- Time origin per model. Each incident model starts its at-risk clock at the correct origin. Watch for immortal-time bias (a span in which the event cannot occur, misattributed to one group) and left-truncation / delayed entry.
- Mediator-ascertainment-window survivorship. A "progressor" / transition label conditional on surviving to a later ascertainment (a second scan, a follow-up visit) is survivorship-biased; plan a landmark time or an explicit intermediate-state (multistate / illness-death) model.
- Primary-analysis-set selection. If the primary is not the full cohort, pre-specify why, from the missingness mechanism relative to the outcome: complete-case regression is unbiased when missingness is independent of the outcome given the covariates (even with a covariate MNAR) and generally biased when it depends on the outcome, including under MAR, where multiple imputation is the usual primary (Hughes et al., Int J Epidemiol 2019; logistic-regression odds ratios are the exception when missingness depends on the outcome alone). Never pick it for significance.
- A design that cannot yet answer these should say so — but at review time a Methods/Limitations admission that it was "not formally assessed" is escalated to MAJOR by the survival probe (S1).
C. Reference standard
Check who established ground truth, when, whether blinding was possible, and whether only a subset had gold-standard verification. Then:
- Construct ↔ nominal-definition match. Restate each construct's nominal definition and confirm every included case satisfies it. An "incidentaloma" defined as an indeterminate finding must not include frank malignancy reads; a label that overshoots its definition inflates the cohort and breaks the κ.
- Per-flag reference-standard concordance. Report concordance per flag category, not only overall; a construct where most flags (e.g., ~86%) miss the reference standard measures something else.
- Manuscript definition ↔
variable_operationalization.md. Methods definitions must match the/define-variablesoperationalization table verbatim, and a blinded re-classification form must quote the analytic protocol's definition verbatim. A paraphrase or "common-sense extension" in the form is the documented cause of a low κ that is a definition mismatch, not real disagreement.
D. Validation
Classify: apparent only, internal split, cross-validation, temporal, external, or multi-center external.
E. Reader / expert-elicitation studies
When the study elicits expert ratings — a reader study, an annotation panel, an AI-output evaluation —
read references/reader_elicitation_design.md before data collection. The acceptance ceiling of a
perceptual / reader AI study is fixed at design time: no quality of execution lifts a ceiling baked
into the comparator, the estimand, or the reader cohort. For an AI-system-versus-human-expert
benchmark, route to /design-ai-benchmarking (arm definition, LLM-as-judge versus human-as-judge
adjudication, structured export schema).
Phase 3: Clinical framing
Ask whether the comparator and endpoint support the stated claim:
- Is the model better than current practice, or just another model?
- Is the endpoint clinically meaningful?
- Does performance translate to action?
- incremental value: if the model/marker is framed as adding value beyond / on top of /
incremental to an existing tool (a clinical score, a routine test, a baseline model), pre-specify
the baseline comparator built from the in-routine-use predictors and how added value will be
shown: test it with the likelihood-ratio (or Wald) test of the new term in the development data;
report ΔC-index / ΔAUC with a paired CI as the size of the gain, not as a second test (a DeLong
test of nested models fitted and evaluated on the same data is invalid; DeLong is for two models
frozen before a held-out evaluation); prefer the change in decision-curve net benefit at a
pre-specified threshold; NRI only as a pre-specified categorical NRI reported separately for
events and non-events, never the continuous NRI (analyze-stats
incremental_value.md; Kerr et al., Epidemiology 2014). A standalone discrimination number does not support a "beyond X" claim, and the nested-model comparison cannot be added post hoc without the baseline model. - fine-tuning contribution baseline: if an NLP/LLM study claims that fine-tuning, LoRA, prompt
engineering, or a multi-agent wrapper improves extraction/classification, pre-specify a
same-backbone zero-shot or few-shot comparator on the identical input, output schema, and test
split; a weaker or unrelated baseline cannot show the adaptation adds value. For imaging:
- a model built by combining / adapting / fine-tuning existing models → the full ablation ladder
(un-adapted base, best single component, direct-train vs transfer), per
references/combine_models_ablation_design.md; - a head-to-head comparison of several models → comparison fairness: one frozen
split/preprocessing through every model, a strong fairly-tuned baseline, a matched (or disclosed)
compute budget, and a paired delta test, per
references/multi_model_comparison_design.md; - a claim that a segmentation model is clinically usable → a pre-specified failure taxonomy, an
acceptability endpoint with a named judge and adjudication rule, the tail beside the mean, and edit
effort paired against manual-from-scratch, per
references/segmentation_failure_characterization_design.md. A mean DSC cannot be converted into a usability claim after the fact.
- a model built by combining / adapting / fine-tuning existing models → the full ablation ladder
(un-adapted base, best single component, direct-train vs transfer), per
- endpoint↔conclusion scope: decide up front what kind of conclusion the design can support. A
cross-sectional / single-visit / prevalence design cannot support a prognostic or surveillance claim
(rescreen interval, disease progression) — that needs longitudinal follow-up. A binary surrogate
endpoint (present/absent, >0, dichotomized) is risk stratification, not a patient-care directive
(defer/withhold/initiate therapy). At review time
/self-review§D +check_scope_coherence.pyflagCROSS_SECTIONAL_PROGNOSTIC/SURROGATE_CARE_DIRECTIVEagainst the conclusion.
Phase 4: Reporting fit
Recommend one primary guideline — TRIPOD+AI, CLAIM, STARD, STROBE (TARGET for a
target-trial emulation), PRISMA, CARE, or ARRIVE — plus journal-specific additions if needed.
Frequent Failure Modes
Diagnostic AI
No clinically relevant comparator; exam-level instead of patient-level split; unclear reference standard; AUROC-only reporting without threshold metrics.
Prognostic modeling
Unclear time zero; immortal time bias; feature timing mismatch; no calibration.
Retrospective cohort / screening database
- Time zero misalignment: cohort entry ≠ follow-up start → immortal time bias.
- Interval-censored outcomes treated as exact → underestimated event times.
- Unacknowledged healthy-volunteer bias → inflated external-validity claims.
- Surveillance bias from unequal follow-up frequency between groups.
- Map each threat explicitly to one of the three bias classes (Hernán/Robins): selection bias (who enters), information bias (how measured), confounding (what else differs).
- Comparative / causal question → emulate a target trial. For a treatment-vs-treatment,
screening-vs-no-screening, or drug-A-vs-drug-B question on routinely-collected data, specify the
seven target-trial components (eligibility, strategies, assignment, time zero, outcome, causal
contrast, analysis plan) before extraction. This prevents immortal-time, prevalent-user, and
confounding-by-indication bias and turns an association into a defensible causal contrast.
New-user + active-comparator design, grace-period clone-censor-weight, and negative controls are in
references/target_trial_emulation.md. - Confounding completeness: pre-specify the adjustment set from a DAG (not a Table-1 p < 0.05
rule), and plan an extended-adjustment sensitivity model to show whether any measured covariate that
turns out imbalanced by exposure but lies outside the adjustment set leaves the primary estimate
robust. Pre-screen the proposed covariates with
scripts/adjustment_set_helper.py(flags mediator / collider / descendant adjustment and omitted confounders, and proposes a candidate backdoor set), then derive the minimal sufficient set with dagitty — seereferences/dag_adjustment.md. At review time/self-reviewPhase 2.5e and the O-probes inobservational_confounding.md(O1–O18) check this against Table 1 — including O7 over-adjustment, O10 overlapping-subset-gradient discipline, O11 design-based weighting for complex-survey data, O12 data-driven-threshold mining, O13 (a cross-sectional mediation claim cannot order X→M→Y), and O14 (a synergy/joint-effect claim needs the additive interaction scale — RERI/AP/S — not a multiplicative-only test).
Multimodal LLM / report generation
No clear rubric for clinical correctness; benchmark labels derived from noisy reports without adjudication; unsupported claims about safety or workflow benefit; input text containing the target label or diagnosis being predicted; no same-backbone zero-shot/few-shot baseline for a fine-tuning or prompt-engineering claim.
Imaging meta-analysis
Overlapping cohorts; paired modalities analyzed as independent; heterogeneity metrics missing; zero-cell handling unspecified.
Minimal-Fix Principle
Recommend the smallest feasible repair first: clarify the claim, narrow the target population, add a limitation statement, add a clinically relevant baseline, re-run one key sensitivity analysis, or redefine the endpoint more explicitly. Escalate to redesign only when the central claim is not defensible otherwise.
Handoff Rules
- route to
analyze-statswhen the design is basically sound but analysis details need refinement - route to
check-reportingafter the design is locked - route to
self-reviewwhen the user wants a pre-submission quality check on their own manuscript - route back to
write-paperonly after the main validity risks are documented
This skill does not compute statistics, draft full manuscript prose, resolve raw data-engineering issues, or replace a full peer review when journal-facing tone is required.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/aperivue/medsci-skills/design-study">View design-study on skillZs</a>