skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
aperivue/medsci-skills105 installs

self-review

Use when checking your own manuscript before submission from a reviewer's perspective. Returns anticipated Major/Minor comments with fixes, including numerical, citation and leakage checks, with an optional multi-reviewer panel. Someone else's paper is /peer-review.

How do I install this agent skill?

npx skills add https://github.com/aperivue/medsci-skills --skill self-review
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The self-review skill is a specialized tool for auditing scientific manuscripts for integrity, consistency, and compliance. It uses a series of local Python scripts and shell commands to perform deterministic checks on manuscript text, tables, and associated analysis scripts. No malicious behavior was detected.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

Self-Review Skill

Check the user's own manuscript before submission and produce an actionable list of anticipated reviewer comments, each with a specific fix — not a written review.

Optional Flags

  • --fix: after the report, apply fixes for every issue with fixable_by_ai: true, editing the manuscript in place, then report a diff summary. Never fix fixable_by_ai: false issues (missing data, design flaws). Maximum 2 fix-and-re-review iterations; if the score is still below threshold after the second, stop and report what remains as structural (inside /write-paper this routes to Phase 7.4a Audit Recovery).
  • --json: also emit the structured JSON block (Phase 3c). Default when called from /write-paper Phase 7.
  • --panel: run the multi-agent panel review (Phase 2.6) — domain-expert reviewers in parallel plus an editor synthesis — instead of the single-pass review. Opt-in and off by default, because it spawns N reviewer agents + 1 editor and costs several times more tokens; reserve it for a high-stakes final pass on a top-tier target. Do not combine with --fix: a panel diagnoses and prioritizes; run --fix as a separate pass once the author has triaged the panel's findings.

Severity Framing

  • Fatal: a design flaw existing data cannot fix (data leakage that invalidates all results, no reference standard, label–feature circularity). Submission would likely be rejected. Reserve Fatal for true design-level problems.
  • Fixable: significant but addressable with existing data (missing calibration, unclear exclusion criteria, absent CIs, incomplete reporting). Most issues are Fixable.

Two Objectives: the Floor and the Ceiling

The floor (categories A–K and the gates of Phases 2.5–2.5f) minimizes rejection-for-cause — fabricated citations, numbers that do not reconcile, overclaims, missing checklist items, leakage — and much of it works by adding a hedge, caveat, disclosure or audit trail. Iterated, that over-hardens a manuscript: every finding is correct, yet the whole reads as a defensive audit (over-hedged, audit trail in the body, Abstract buried under caveats, the strongest robustness result hidden in Limitations). The ceiling pass (category L / Phase 2.5g) runs after the floor gates, reads the accurate manuscript as a whole, and recommends only SUBTRACTION — REMOVE, MOVE, or TIGHTEN. It is advisory, never blocks, and cannot relax a floor gate. Report its findings as their own block (Phase 3), not folded into the "add this" comments. Phase 2.5i then reads the floor + ceiling state to declare the loop done, including a zero-edit PASS.

Workflow

Phase 1: Intake

  1. Get the manuscript — PDF, Word doc, or pasted text.

  2. Ask the user: target journal (sets reporting standards and scope); manuscript type (original research / review / perspective / technical note / letter / meta-analysis / case report); anything they are already worried about. On an interactive run, offer the --panel review (Phase 2.6) once, in one line, then proceed single-pass unless the user opts in. Never offer or apply the panel under --json or when called from /write-paper.

  3. Read the full manuscript.

  4. SSOT gate — confirm there is one manuscript, not several. Self-review reads a single file, so drift between a legacy working copy and the live submission copy is invisible to it. Before a --panel run or any pre-submission pass, check for multiple copies:

    find . \( -path '*manuscript*' -o -path '*main_document*' \) -name '*.md' | grep -v node_modules
    

    If more than one manuscript-like file exists, confirm which is the SSOT and run /sync-submission's divergence gate — a STALE_COPY (an SSOT numeric claim or heading that did not propagate to the other copy) is a P0 that must clear first:

    python3 "${CLAUDE_SKILL_DIR}/../sync-submission/scripts/detect_copy_divergence.py" \
      --ssot <ssot>.md --copy <other-copy>.md
    

    Review the SSOT copy, never a stale one. Under --panel this is blocking: if find returns more than one file and the SSOT is not pinned (no SSOT.yaml with truth.manuscript_md, no explicit --ssot <path>), STOP before spawning any reviewer and have the user name the SSOT (and clear any STALE_COPY), because a panel on a stale copy wastes the whole pass. Do not auto-pick the longest/newest file. A single-pass review may proceed on the one file it was given.

Phase 2: Systematic Check

Work every category the Research-Type Adaptation table (below) marks as applicable, and for each item decide whether a reviewer would raise it as a Major or Minor comment. The per-item check tables are in references/phases/phase2_systematic_check.md — read it once you have the manuscript and know its type.

CategoryWhat it asks
AStudy Design & Data Integritypatient-level splits, leakage, input-text contamination, analysis unit
BReference Standard & Ground Truthdefinition specificity, timing, annotator independence
CValidation & Statistical ReportingCIs, calibration, comparator, effect size, power-aware nulls, equivalence margins, interaction anchoring
DClinical Framing & Importanceintended use, overclaiming, novelty, endpoint↔conclusion scope
EReproducibilitypreprocessing, model detail, hardware/software, data & code availability
FReporting Completenessabstract↔body consistency, flow diagram, ethics, missing data, word cap
GReporting Guideline Compliancematch the type to its checklist; /check-reporting does the item-level audit
HCircularitylabel–feature overlap, tautological prediction, circular validation
IProtocol Heterogeneitymulti-site acquisition, harmonization, temporal protocol drift
JMethod Transparencymodel provenance, fine-tuning, classical-style body conventions
KReviewer-team consistencySR/MA only — dual-vs-single conjunction, LLM-as-reviewer (both fabrication-grade)
LEditorial impression & defensivenessadvisory, never blocking — the ceiling category: REMOVE / MOVE / TIGHTEN

Run the deterministic gates at Phase 2 entry, on every path:

# D. endpoint↔conclusion scope
python3 "${CLAUDE_SKILL_DIR}/scripts/check_scope_coherence.py" \
  --manuscript manuscript.md --out qc/scope_coherence.json --strict

# J. classical-style body conventions
python3 "${CLAUDE_SKILL_DIR}/scripts/check_classical_style.py" \
  --manuscript manuscript.md --out qc/classical_style.json --strict

# K. reviewer-team consistency (SR/MA only; pass the extraction JSON file or directory)
python "${CLAUDE_SKILL_DIR}/scripts/check_reviewer_team_consistency.py" \
    --manuscript manuscript.md --prospero prospero/record.md \
    --extraction-json extraction/ --out _audit_self/reviewer_team_consistency.md

# J/D. Perspective structure (genre-gated: silent unless article_type is a Perspective).
# Pass the known type via --type; it also self-detects from the front-matter article_type.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_perspective_structure.py" \
  --manuscript manuscript.md --type "${TYPE:-}" --out qc/perspective_structure.json

# Was every reported analysis ever defined (outcome, reference standard)?
python3 "${CLAUDE_SKILL_DIR}/scripts/check_analysis_definitions.py" \
  --manuscript manuscript.md --json --strict > qc/analysis_definitions.json

Verdict mapping: CROSS_SECTIONAL_PROGNOSTIC, SURROGATE_CARE_DIRECTIVE, SECTION_SYMBOL, INBODY_AI_DISCLOSURE, any reviewer-team hit (exit 1), MODEL_OUTCOME_UNDEFINED (a Cox / Fine–Gray / logistic model with no outcome named), MODEL_NOT_IN_METHODS, and REFERENCE_STANDARD_UNDEFINED (discrimination or calibration with nothing to score against) are Anticipated Major Comments. CROSS_SECTIONAL_YIELD_LANGUAGE, ELIGIBILITY_PROSE, DECIMAL_INCONSISTENCY, EM_DASH_OVERUSE, PERSPECTIVE_HEADING_NOT_ASSERTION, PERSPECTIVE_ABSTRACT_NO_AUTHORIAL_MOVE, and TIER_LABEL_UNDEFINED are Minor. The per-verdict rationale and resolution paths are in the Phase 2 reference file.

ANALYSIS_LOAD is informational, never a verdict: many analyses are the cause of omitted definitions, not the defect. Do not cut analyses to satisfy it; restore the definitions they crowded out, and if load is genuinely high, move the defensive analyses to the supplement.

Research-Type Adaptation

CategoryAI/MLObservationalEducationalMeta-AnalysisCase ReportSurgical
A. Study DesignFullFullPartialN/AN/AFull
B. Reference StandardFullFullN/APer-studyPartialFull
C. Validation & StatsFullFullFullSpecial*PartialFull
D. Clinical FramingFullFullFullFullFullFull
E. ReproducibilityFullPartialPartialPartialN/AFull
F. ReportingFullFullFullFullFullFull
G. Guideline ComplianceFullFullFullFullFullFull
H. CircularityFullPartialN/AN/AN/APartial
I. Protocol HeterogeneityFullFullN/APer-studyN/AFull
J. Method TransparencyFullPartialPartialN/AN/APartial
K. Reviewer-team consistencyN/AN/AN/AFullN/AN/A
L. Editorial impressionFullFullFullFullFullFull

*Meta-analysis: replace C with heterogeneity (I², prediction intervals), publication bias (funnel plot, Egger), and sensitivity/subgroup analyses.

Type-specific additional checks:

  • Observational: confounding (DAG or adjustment strategy), selection bias, exposure measurement validity. Run Phase 2.5e (Confounding Completeness), then apply the O-probes in references/domain-probes/observational_confounding.md — O1 (covariate imbalanced by exposure in Table 1 yet absent from the adjustment set) and O8 (records > subjects with the analysis unit undisclosed; check_cohort_arithmetic.py --id-col) are the deterministic ones, and O7 (adjusting for a mediator/consequence of the outcome) is their opposite-direction twin. For a clinical prediction model (TRIPOD / TRIPOD+AI, nested predictor-set comparison), also apply the CP-probes in references/domain-probes/clinical_prediction_model.md.
  • Educational: learning-outcome measurement validity, Kirkpatrick level, control-group adequacy, curriculum fidelity.
  • Meta-analyses: search comprehensiveness (2+ databases), screening reproducibility (2 reviewers), per-study RoB, GRADE certainty.
  • Case reports: diagnostic-reasoning transparency, timeline completeness, informed consent, generalizability disclaimer.
  • Surgical: learning curve, surgeon volume/experience, complication grading (Clavien-Dindo), operative detail.

Domain probe modules — the same probes /peer-review uses, vendored here. Load every module whose row matches:

Manuscript type / signalProbe module
Systematic Review / Meta-Analysisreferences/domain-probes/sr_ma.md (P0–P19)
Time-to-event / survival / prognostic model (Cox, Fine-Gray, DeepSurv, nomogram, risk-stratification cutoff)references/domain-probes/survival_prognostic.md (S1–S9)
Radiomic feature reproducibility / acquisition-parameter sweep / reliability-based feature filteringreferences/domain-probes/radiomics.md (R1–R4)
Cross-modality image synthesis (MRI→PET / MRI→CT / non-contrast→contrast / low-dose→full-dose) claiming functional/molecular information or target-modality substitutionreferences/domain-probes/image_synthesis.md (IS1–IS4)
Narrative / review article / primer / state-of-the-artreferences/domain-probes/narrative_review.md (RV1–RV9)
Perspective / opinion / viewpoint (npj DM long-essay, Lancet Comment, NEJM AI / RYAI short-structured)references/domain-probes/narrative_review.md (RV1–RV9) + the check_perspective_structure.py gate above
AI/ML primary study with a clinical claim (generalizable / outperforms clinicians / deployment-ready / can replace a reader)references/domain-probes/ai_overclaiming.md (AO0–AO7)
Engineer-built medical-imaging model (segmentation / classification / detection) being validated — partition/leakage, seed & run variance, metric selection, reproducibility, reference-standard quality; plus saliency faithfulness, uncertainty/OOD/abstention and deployment feasibility when a clinical-use claim is madereferences/domain-probes/model_development.md (MD0–MD11)
LLM / MLLM evaluated on a clinical task (report generation, VQA, clinical text extraction/classification; closed API or open weights)references/domain-probes/mllm_evaluation.md (ME0–ME8)
Randomised controlled trial (parallel / crossover / cluster / stepped-wedge)references/domain-probes/rct_trial.md (RC0–RC7)
Diagnostic test accuracy (DTA) primary study / multi-reader multi-case (MRMC) reader study (AI-vs-reader, AI-assisted reading, modality comparison)references/domain-probes/diagnostic_accuracy.md (D1–D12)
Case report / case series (incl. adverse-event/pharmacovigilance and imaging-led radiology/nuclear-medicine/IR reports)references/domain-probes/case_report.md (CR1–CR9)
AI/ML, prediction, or diagnostic study claiming cross-population performance, or presenting subgroup analyses as a fairness/equity argumentreferences/domain-probes/equity_fairness.md (EQ0–EQ6)
Mendelian randomization (two-sample, one-sample, multivariable, drug-target / cis-MR, non-linear MR)references/domain-probes/mendelian_randomization.md (MR1–MR8)
Polygenic risk score (PRS / PGS) developed, validated, or applied as a predictor or risk-stratifierreferences/domain-probes/polygenic_risk_score.md (PG1–PG8)
Network meta-analysis (≥3 interventions via direct + indirect evidence, treatment ranking, incl. component NMA)references/domain-probes/network_meta_analysis.md (NM1–NM8)
Health economic evaluation (cost-effectiveness / cost-utility / cost-benefit / budget-impact; trial- or model-based — decision tree, Markov, DES)references/domain-probes/health_economic_evaluation.md (HE1–HE8)
Observational study using routinely-collected health data (claims / EHR / registry / health-checkup DB, linked or not)references/domain-probes/record_routinely_collected_data.md (RD1–RD8)
Self-report survey / questionnaire study (KAP, physician/patient survey, web/e-survey)references/domain-probes/survey_research.md (SV1–SV8)
Scoping review (maps breadth of evidence; PCC framing, charting — not a focused effectiveness/accuracy question)references/domain-probes/scoping_review.md (SC1–SC8)
Qualitative study (interviews, focus groups, ethnography, grounded theory, phenomenology, document analysis)references/domain-probes/qualitative_research.md (QL1–QL8)
Self-improving / self-evaluating system (an agent that critiques and rewrites its own output; training on model-generated data; an LLM judge scoring the training signal; "self-evolving" clinical agents)references/domain-probes/self_improving_system.md (SI1–SI7) + ${CLAUDE_SKILL_DIR}/../peer-review/scripts/check_self_improvement_claims.py

Apply each probe as an additional source of comments, complementing (not replacing) categories A–K: a conclusion-threatening or design-level finding becomes a Fatal Anticipated Major Comment, a reporting-level finding a Fixable Anticipated Minor Comment, each tagged with the closest category letter (A–K).

For a classifier / NLP / tabular ML manuscript, also run the feature-selection-leakage gate — a data-driven selection (feature selection, univariate filtering, vocabulary, a threshold) fit on the FULL dataset before cross-validation inflates the CV metric:

python3 "${CLAUDE_SKILL_DIR}/scripts/check_cv_leakage.py" \
  --manuscript manuscript.md --out qc/cv_leakage.json

CV_SELECTION_LEAKAGE (Major) fires when a selection token co-occurs with cross-validation and no fold-nesting is disclosed ("within each fold" / "nested CV" suppresses it). Patient-vs-image split leakage is a different check (model-assessment/check_split_leakage.py).

Phase 2.5: Numerical Cross-Verification (Internal)

Before the report, verify internal consistency:

  1. Abstract vs Body: every Abstract number matches Results and Tables.
  2. Table vs Text: sample sizes, primary outcomes and p-values agree between tables and narrative.
  3. Figure vs Text: figure legends match the data described in Results.
  4. Percentage arithmetic: n/N percentages are correct (23/150 = 15.3%, not 15.0%). The gate holds each cell to its printed precision (half a unit in the last printed place).
  5. CI plausibility: confidence intervals are reasonable for the sample sizes.
  6. Rate back-calculation: every rate inverts to its own numerator/denominator (incidence rate ≈ events / person-years × scale, ±rounding). A rate that does not recompute, or implies more events than the cohort can supply, is a Major.
  7. Exclusion-cascade and complete-case arithmetic (cohort/observational): start N − Σ(exclusions) == final analytic N, and total − missing == complete. A footnote N that does not equal the subtraction is a Major.

For cohort/observational manuscripts, run the gate instead of eyeballing it (it parses prose equations + GFM tables, and recomputes from a committed CSV when given one):

python3 "${CLAUDE_SKILL_DIR}/scripts/check_cohort_arithmetic.py" \
  --manuscript manuscript.md --data analysis/cohort.csv --id-col mockid \
  --out qc/cohort_arithmetic.json --strict

RATE_BACKCALC / CASCADE_SUM / PARTITION_OVERLAP are Anticipated Major Comments (category A). On health-screening / EMR / registry data pass --id-col (or let it auto-detect a subject-ID column) so it also checks the analysis unit: records > unique subjects with neither the unit nor a one-record-per-subject sensitivity stated emits ANALYSIS_UNIT_UNDISCLOSED (Major — non-independent observations give anti-conservative CIs; probe O8). Flag any remaining internal-consistency discrepancies as Anticipated Minor Comments (category F).

Then recompute what a reviewer recomputes by hand:

# Every "n (%)" in a table, recomputed against its own denominator.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_table_percentages.py" \
  --manuscript manuscript.md --out qc/table_percentages.json --strict

# Every reported P beside a 2×2 (or r×c) count, recomputed from the counts themselves.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_reported_p_from_counts.py" \
  --manuscript manuscript.md --json --strict > qc/reported_p.json

# Every t(df) / F(df1, df2) / χ2(df) / z with a P in its sentence, P recomputed from the statistic.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_test_statistic_p.py" \
  --manuscript manuscript.md --json --strict > qc/test_statistic_p.json   # add --grim grim.json

# Diagnostic-accuracy only: sensitivity/specificity against the reference-standard denominators.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_dta_denominators.py" \
  --manuscript manuscript.md --json --strict > qc/dta_denominators.json

PERCENT_MISMATCH, P_NOT_REPRODUCIBLE, and DTA_DENOMINATOR_MISMATCH / STAGE_ROWSUM are P0 Major — not a rounding disagreement: one of the two numbers is wrong. Run the first two on every manuscript with a table; the third only on diagnostic-accuracy work.

check_test_statistic_p.py reads the statistic and the P at their printed precision: a P no statistic in the rounding interval can give is P_STAT_INCONSISTENT (Minor), and Major P_STAT_DECISION_ERROR when the two also fall on opposite sides of alpha (--alpha). Adjusted or non-standard P values (corrected, exact, Welch, …) and results that match only the unstated one-sided P are never Major. --grim takes declared means ([{"label", "mean", "n", "items"?, "decimals"?}]): GRIM_INCONSISTENT (Major) when no integer sum gives the mean; GRIM_NOT_ASSESSED when n·items >= 10^decimals.

To check each P against its own test, declare the tests as p_tests.json (copy ${CLAUDE_SKILL_DIR}/templates/p_tests.json; schema in references/p_tests_schema.md) and add --tests p_tests.json. The row's P is recomputed with the declared test (fisher, chi2, chi2_yates). A reported P whose whole printed-precision interval lies on the other side of alpha from the recomputed P is P_ALPHA_CROSSING (Major). An adjusted, paired or other: test, a row whose printed percentage is not count/n (another denominator), or a table of 3+ groups (after dropping a Total column) is P_NOT_ASSESSED (Minor).

Known limits: without --tests, check_reported_p_from_counts.py flags a row only when its reported P differs from the crude 2x2 P by more than one order of magnitude under every test family, and only in tables with at least two count rows. It does not flag a reported < bound (<0.001) that the crude P exceeds, or a P on the other side of alpha, by less than that, nor a lone count row (open finding SR-01 in prose-only mode): the table cannot show whether the P is adjusted, paired or on another denominator. Check such P values against the stated test by hand, or declare them. check_table_percentages.py reads one denominator per column. A cell that misses only at its printed precision (within 0.5 pp) on a footnoted row (Current smoker^a | 23 (15.4%) under n = 150 with 149 known) is a Minor PERCENT_PRECISION_NOTE; an unfootnoted row, or a larger miss, is a PERCENT_MISMATCH. Confirm footnoted rows against the footnote.

Phase 2.5a: Numerical Source-Fidelity Audit (External)

Numbers can be self-consistent everywhere and still wrong at the source; only a traversal back to the primary source catches that. First run the displayed-arithmetic gate — a stated difference must equal the subtraction of its two displayed components at the same precision:

python3 "${CLAUDE_SKILL_DIR}/scripts/check_rounded_delta.py" \
  --manuscript manuscript.md --out qc/rounded_delta.json

ROUNDED_DELTA_MISMATCH (Minor) fires when AUCs shown as 0.79 and 0.82 are reported with a difference of 0.02. A higher-precision pair (0.794 vs 0.816) with a 2-dp delta is legitimate and not flagged.

When to run the external audit: MA revisions, submissions, or when the user says "check against the source", "verify extraction", or "random sample". Skip otherwise.

The audit: draw a stratified sample of 5 numerical claims — always including one comparative-arm value and one revision-introduced number — and trace each through three layers (manuscript → extraction CSV → primary-source page; plus analysis script → CSV where a script produced it). Any mismatch is a Major Comment; one that reverses a direction or crosses a significance boundary is a P0 blocker. Every [VERIFY-CSV] tag is a mandatory audit item regardless of sample size.

Read references/phases/phase2_5a_source_fidelity.md when running the external audit — it has the traversal procedure, recording table, sampling strata, and four prose-judgement rules (hand-entered analysis-script inputs; prose↔table statistic-type mismatch, e.g. a median in the text against a mean in Table 1; stale derived CSVs after a model/adjustment-set change; direction reversals internal consistency cannot see).

Phase 2.5a-2: Design & Power Statistic Provenance

Applies only when the manuscript states a sample-size calculation, a power figure, or a detectable-effect claim. These are computed, not copied, so Phase 2.5a cannot check them: re-derive each from the manuscript's own inputs. A value not reproduced by committed code, or reproducible only by a method the committed script does not implement, is a Major Comment (P0 if a headline claim). Read references/phases/phase2_5a2_design_power.md when the manuscript reports a sample-size / power / MDE calculation.

Phase 2.5b: Screening-Count Reconciliation from ID Sets (SR/MA + observational tier/stratum)

When to run: any SR/MA manuscript revision, at any stage (before Phase 3); or any observational manuscript presenting an ordinal tier / mutually-exclusive stratum split. Skip otherwise. A wrong prose total survives every other pass because Abstract, Methods, Results, captions and supplement all cite it back to each other; only a recount from the ID sets catches it.

A. SR/MA — recount from the ID sets. Derive every study count from the screening TSV and the consensus sheet, never from prose, and list the narrative-only IDs explicitly (turning "10 narrative-only studies" into "2 (IDs 120, 474)"). A derived total that disagrees with the Abstract, Methods, Results, Figure 1 caption or Limitations is a P0 Major, blocking submission; an N → M transition claim not backed by an enumerable ID addition/subtraction set is a Major.

B. Observational tier/stratum. A disjoint partition must satisfy Σ(stratum N) == unique total and Σ(stratum events) == total events. Denominators summing above the unique cohort double-count subjects; every stratum n equal to the grand total is a mis-entry. Confirm the reference (baseline) row of any stratified hazard/odds table is present and labelled.

python3 "${CLAUDE_SKILL_DIR}/scripts/check_cohort_arithmetic.py" \
  --manuscript manuscript.md --data analysis/strata.csv --strict

C. Cross-script cut-point consistency. When one cohort is re-stratified in more than one script, the derived categorical must use one cut definition (same breaks, same right= closure, same labels) — otherwise per-stratum Ns drift while the grand total still reconciles. The same gate covers a derived 0/1 composite rebuilt in a second script with a clause dropped.

python3 "${CLAUDE_SKILL_DIR}/scripts/check_binning_consistency.py" \
  --root analysis --root scripts --strict

PARTITION_OVERLAP, BINNING_DRIFT, and DERIVED_DEF_DRIFT are all P0 Major. Read references/phases/phase2_5b_screening_counts.md when doing the SR/MA ID-set recount or a stratified-cohort recount (set definitions, derivation formulas, reconciliation-block template).

Phase 2.5c: Reference Scans (hallucination + adequacy)

2.5c catches a citation that does not exist or has an invented first author; 2.5c-2 catches a claim with no citation. Both need a bibliography — skip them if there is no refs.bib and no reference list. Run /verify-refs --strict first (these scans read its audit), then the adequacy checker:

python3 "${CLAUDE_SKILL_DIR}/scripts/check_reference_adequacy.py" \
  --manuscript manuscript/manuscript.md --bib "$BIB" \
  --article-type "$TYPE" ${CAP:+--journal-cap "$CAP"} \
  --out qc/reference_adequacy.json --strict

A FABRICATED record or any duplicate_findings[] entry in qc/reference_audit.json is a P0 Major Comment that blocks submission. Read references/phases/phase2_5c_reference_scans.md when the manuscript has a bibliography and you are auditing citations.

Phase 2.5d: Cross-Reference QC (Manuscript ↔ rendered DOCX)

In-text Table/Figure citations can resolve to a different caption in the rendered DOCX when the build script carries its own legacy SSOT; no other phase sees this.

Markdown stage (always). Every captioned Figure N. / Table N. must be cited elsewhere in the body:

python3 "${CLAUDE_SKILL_DIR}/scripts/check_figure_citation.py" \
  --manuscript manuscript.md --out qc/figure_citation.json

FIGURE_ORPHAN / TABLE_ORPHAN (Minor): a float with a legend but no in-text citation.

DOCX stage (only when a rendered DOCX exists — circulation drafts, post-build checks):

python3 "${CLAUDE_SKILL_DIR}/../manage-refs/scripts/check_xref.py" \
  --md manuscript/manuscript.md --docx manuscript/manuscript_final.docx \
  --out qc/xref_audit.json [--allow-separate-attachments]

MISMATCH is always Major (P0). MISSING_DOCX and MISSING_BODY are Major (P0) by default. For journals that take figures/tables as separate attachments (European Radiology, Radiology, AJR), pass --allow-separate-attachments: it downgrades MISSING_DOCX to Minor, and MISSING_BODY to Minor only when no --docx was supplied — nothing was checked, so treat those rows as unverified (summary.downgraded_unchecked) and re-run with --docx before submission. MISSING_BODY for a float that IS in the rendered DOCX stays P0 (SSOT drift). UNCITED is Minor.

Do NOT auto-fix cross-reference defects in --fix mode, because rewriting a body caption without re-running the DOCX build only moves the mismatch. Emit each P0 row as its own M-numbered Major Comment with category: "F" and fixable_by_ai: false, and route the user to /write-paper Step 7.6a. Read references/phases/phase2_5d_xref_qc.md when the xref gate fired and you are writing up the reconciliation.

Phase 2.5e: Confounding Completeness (observational only)

When to run: observational manuscripts (cohort, case-control, cross-sectional, health-screening registry) whose central claim is an adjusted exposure–outcome association. Skip for RCTs, diagnostic-accuracy, SR/MA, and descriptive studies.

A covariate that is measured, imbalanced across exposure groups in Table 1, and absent from the adjustment set (probe O1) is invisible to a prose pass. Run the gate and treat each UNADJUSTED_IMBALANCED covariate as an Anticipated Major Comment (category A):

python3 "${CLAUDE_SKILL_DIR}/scripts/check_confounding_completeness.py" \
  --table1 table1_by_<exposure>.csv \
  --adjusted-list "age, sex, BMI, hypertension, diabetes" \
  --exposure-defining-list "body mass index, waist, fasting glucose, triglycerides, HDL cholesterol" \
  --out qc/confounding_completeness.json --strict

For observational manuscripts, read references/phases/confounding_completeness.md for the full procedure: the --exposure-defining-list over-adjustment exemption for guideline-defined exposures (MASLD / metabolic syndrome / CKM / sarcopenia / frailty), the SMD-from-mean ± SD fallback, the extended-adjustment sensitivity model (refit the unadjusted estimate on the reduced complete-case frame, not the full frame), and probes O2–O10 of references/domain-probes/observational_confounding.md.

Phase 2.5f: Claim-vs-Artifact Cross-Check

This phase checks claims against the artifacts they should trace to — the pre-registration, the protocol, the analysis outputs — where the prose is internally consistent yet disagrees with them. Run the gates (pass the supplement so the corpus is complete):

# 1. claims ↔ pre-registration/protocol: estimand provenance + E-value arithmetic
python3 "${CLAUDE_SKILL_DIR}/scripts/check_claim_artifact.py" \
  --manuscript manuscript.md --prereg prereg.md \
  --out qc/claim_artifact.json --strict   # + --evalues evalues.json (templates/) to recompute declared E-values

# 2. Methods ↔ Results ↔ disk coverage (both directions: promised-absent AND run-but-unreported)
python3 "${CLAUDE_SKILL_DIR}/scripts/check_artifact_coverage.py" \
  --manuscript manuscript.md --supplement supplement.md --analysis-dir output/analysis \
  --out qc/artifact_coverage.json --strict

# 3. reader-facing residue in EVERY rendered artifact, not just the body
python3 "${CLAUDE_SKILL_DIR}/scripts/check_supplement_hygiene.py" \
  --supplement supplement.md --supplement tables.md --supplement captions.md \
  --manuscript manuscript.md --out qc/supplement_hygiene.json --strict

# 4. float AND in-text reference-number ([N]) citation order — a desk-reject item the hygiene gate does not cover
python3 "${CLAUDE_SKILL_DIR}/scripts/check_citation_order.py" \
  --manuscript manuscript.md --out qc/citation_order.json --strict

# 5. a headline null is uninterpretable without a precision statement
python3 "${CLAUDE_SKILL_DIR}/scripts/check_null_calibration.py" \
  --manuscript manuscript.md --out qc/null_calibration.json --strict

# 5b. a headline OR/HR/RR whose 95% CI spans an order of magnitude (a direction, not a magnitude), or events/covariates < 10 (EPV)
python3 "${CLAUDE_SKILL_DIR}/scripts/check_effect_stability.py" \
  --manuscript manuscript.md --out qc/effect_stability.json --strict

# 5c. incorporation bias — a trajectory-defined reference standard with a trajectory predictor (growth) reported as associated with the outcome
python3 "${CLAUDE_SKILL_DIR}/scripts/check_incorporation_bias.py" \
  --manuscript manuscript.md --out qc/incorporation_bias.json --strict

# 6. reader/observer study only — prove the (call × confidence) → score encoding is strictly
#    monotonic; a folded score silently mis-estimates the AUC and no prose review can see it
python3 "${CLAUDE_SKILL_DIR}/../analyze-stats/scripts/rating_monotonicity.py" \
  --encoding score_def.json
VerdictSeverity
PRIMARY_REASSIGNEDMajor — the primary was re-designated after results were known
EVALUE_ARITHMETICMajor — recompute for the declared primary estimate; EVALUE_NON_PRIMARY is an advisory flag (check which estimate the E-value bounds)
PROMISED_ABSENT, DISK_UNREPORTED, PROMISED_STAT_NO_VALUEMajor
SUPP_INTERNAL_LABEL, SUPP_PLACEHOLDER, SUPP_BUILD_MARKER, SUPP_RESPONSE_FRAMING, SUPP_PLANNING_RESIDUE, SUPP_XREF_UNRESOLVEDMajor — a slip in a supplement is as fatal at a technical check as one in the body
CITATION_ORDERMajor; CITATION_GAP Minor
CONFIRM_NULL_NO_MDEMajor
ESTIMAND_DRIFT, PRIMARY_DISCLOSURE_NOTEAdvisory Minor — never a blocker. The provenance match is fuzzy (token overlap); confirm against the actual registration first. PRIMARY_DISCLOSURE_NOTE flags an honest disclosure the guidance recommends — do not penalise it.

With --evalues (schema references/evalues_schema.md), each declared E-value is recomputed from its declared RR and CI (VanderWeele–Ding; the CI E-value from the limit nearest 1) over the printed precision of every number: EVALUE_DECLARED_MISMATCH is Major, EVALUE_DECLARED_NOT_IN_TEXT Minor; an OR/HR must be converted to an RR or declared other: (Minor, not recomputed). Known limits: in prose-only mode the E-value check splits sentences at every '.', so the decimal in "HR 1.52" can cut the estimate out of its window (EVALUE_UNVERIFIABLE); "E-value = 3.10", "(E-value 3.10)" and the plural are not read, and a CI-limit E-value is not recognised (open finding SR-02 without --evalues). Check by hand, or declare them.

Checks no script makes (prose judgement):

  1. Primary-change guard — two models for one contrast, one significant and one null, the significant one foregrounded: confirm which was pre-specified.
  2. Headline vs own-sensitivity direction — a headline claim pointing the opposite way from the authors' own sensitivity estimate means the paper contradicts its own robustness check (Major).
  3. Figure-embedded numbers are grep-blind — every numeric audit above misses numbers inside a rasterised figure. Read each figure page visually before submission.

Re-run /sync-submission's check_cross_artifact_stale.py after any reframe, not just once at the start. For time-to-event manuscripts, apply probe S8 (estimand provenance) of references/domain-probes/survival_prognostic.md. Read references/phases/phase2_5f_claim_artifact.md when a gate above fired and you need the rationale and resolution path, or there is a pre-registration to reconcile.

Phase 2.5g: Editorial-Impression / Defensiveness Scan (the ceiling pass)

Run this after the floor gates (Phases 2.5–2.5f): it reads the accurate manuscript and recommends what to take back out. It is advisory and non-blocking — it never produces a Major and never gates submission.

python3 "${CLAUDE_SKILL_DIR}/scripts/check_editorial_impression.py" \
  --manuscript manuscript.md --out qc/editorial_impression.json

It exits 0 even under --strict and emits up to six verdicts, each tagged with a SUBTRACTION action:

VerdictReads asAction
HEDGE_DENSITYdefensive-caveat tokens per 1,000 narrative words over thresholdTIGHTEN
HEDGE_REPEATone caveat motif repeated across body + AbstractTIGHTEN
AUDIT_IN_BODYSHA / commit / unit-test / post-lock / manifest / seed in the narrativeMOVE (→ Methods/supplement)
LIMITATIONS_VOLUMEa long enumerated Limitations listTIGHTEN (consolidate)
ABSTRACT_CAVEAT_LOADseveral caveat clauses in the AbstractTIGHTEN
BURIED_DEFENSEstrong numeric robustness result only in Limitations/supplementMOVE (→ Results)

Each finding becomes a Minor issues[] entry with category: "L", category_name: "Editorial impression", issue_type: "editorial_impression", subtype: <verdict>, and action: "REMOVE" | "MOVE" | "TIGHTEN", reported in the Phase 3 "Editorial-Impression Risks" block. Mark them fixable_by_ai: false — tightening a hedge or moving a result is the author's voice-and-judgment edit — except a clearly redundant HEDGE_REPEAT, which --fix may collapse to one statement.

When an earlier phase recommends adding a caveat or disclosure, weigh it against L: an integrity-critical disclosure is a must, stated once and crisply; a defensive over-disclosure is a cut or move. Place it once and point to the supplement rather than repeating it at every claim site.

Phase 2.5h: Baseline Drift (anchor to the last human-approved version)

Run after the ceiling pass and before the loop controller (Phase 2.5i). Each refine pass takes the previous AI output as its baseline, so framing bias compounds unseen; this gate compares the manuscript against the last human-approved version (the frozen v_N circulated to senior authors/co-authors — not the last AI output). With no baseline (a first draft), skip it.

python3 "${CLAUDE_SKILL_DIR}/scripts/check_baseline_drift.py" \
  --manuscript manuscript.md --baseline "$BASELINE_MD" \
  --out qc/baseline_drift.json
VerdictSignal (baseline → current)Fold into report as
STRENGTH_INFLATIONcertainty markers up while hedges fallMinor — tone back to the approved strength
SIGNIFICANCE_INFLATION_DRIFTnovel/pivotal/unprecedented tokens addedMinor — remove the inflation
SCOPE_INFLATION_DRIFTnew generalization phrases ("in clinical practice")Minor — the estimand did not widen; re-scope
HEDGE_ACCRETIONhedge/caveat density upMinor — cumulative over-hardening; TIGHTEN

Every finding is Minor and advisory; the gate never blocks. Treat drift as a prompt to review against the approved anchor, not an instruction to revert — new analysis can justify a stronger claim, but the author should confirm it. qc/baseline_drift.json feeds the loop controller, so a drifted draft does not read as a zero-edit PASS.

Phase 2.5i: Refinement Terminal-State (the loop controller)

Run this last, after the floor gates and the ceiling pass: it reads their qc/*.json artifacts and classifies whether the review → revise → review loop is done, so an accurate manuscript is not over-hardened by a pass it does not need. It is advisory and never blocks; do not treat it as a second gate on the floor detectors, which already fail under --strict on their own Majors.

python3 "${CLAUDE_SKILL_DIR}/scripts/refinement_stop.py" \
  --qc-dir qc --out qc/refinement_stop.json
VerdictMeaningWhat the harness must do
CONTINUEa floor gate still reports a Majorgenuine work remains — keep going
STOP_OVERHARDENINGfloor clean, ceiling flags accumulationSTOP adding; only optional SUBTRACTION (REMOVE/MOVE/TIGHTEN) remains — do not run another additive pass
STOP_MINOR_OPTIONALfloor clean, only optional Minor polish leftstop the required-work loop; present the Minor items as an optional menu, do not loop for them
STOP_ZERO_EDITfloor at fixed point, ceiling cleanthe manuscript is submission-ready as-is — NO EDITS REQUIRED. Do not manufacture changes. Report the zero-edit PASS as a first-class outcome
INDETERMINATEno gate artifacts yet, no floor gate parsed (e.g. only the ceiling ran), or an empty / invalid qc/*.json (a gate that crashed under --json > file)run, or re-run, the floor + ceiling gates first

Known limits: a detector-keyed artifact whose findings are not under claims / findings (for example /verify-refs' qc/reference_audit.json) is listed as Unparsed with a WARNING but does not by itself block a STOP_*; read its own verdict before acting on the stop signal.

Once the verdict is any STOP_*, stop the additive cycle: surface the terminal state in the Phase 3 report and do not re-run self-review to find "one more thing". A zero-edit or minor-optional result is a legitimate outcome, not a failure to try harder.

Phase 2.5j: Refinement Regression (fixed vs broke, across runs)

Run each round, after the loop controller. A revision that resolves finding X can introduce finding Y, and the pass-rate hides it; this step reads a run-history ledger (one line per run, the verdict@where fingerprints of its findings) and reports what the revision fixed vs what it broke.

python3 "${CLAUDE_SKILL_DIR}/scripts/refinement_regression.py" \
  --qc-dir qc --ledger qc/refinement_ledger.jsonl --append \
  --out qc/refinement_regression.json

Use --append on a real run so the current findings become the next entry; omit it to classify without recording.

VerdictMeaningWhat the harness must do
PROGRESSINGfindings resolved, none newcontinue
REGRESSIONthe revision introduced new finding(s)review the new findings before accepting the fix — the pass-rate went up but something broke
CHURNINGa resolved finding reappeared (Mirror Loop)stop revising and re-anchor — more passes re-derive, they do not converge
CONVERGEDnothing new, nothing carriedthe loop is done
INDETERMINATEfirst run, no prior entryre-run after a revision

It is advisory and never blocks. Report both axes in Phase 3: a revision is an improvement only if it resolved findings and the new/churn columns are empty.

Phase 2.6: Multi-Agent Panel Review (--panel, opt-in)

Run this phase only when --panel is passed, after the numerical audits (Phases 2.5–2.5d) so the reviewers see source-verified numbers, and before the Phase 3 report, which it feeds. Two things bind before you spawn anything: the SSOT must be singular (the Phase 1 step 4 gate — halt and ask if more than one manuscript-like .md is unpinned), and the roster must not be a substrate monoculture (a panel sharing the drafter's model inherits its blind spots; route at least one lens to Codex or a human co-author). check_panel_diversity.py --strict enforces the second, and fires PANEL_UNDERRETURN when fewer reviewers returned than were spawned — a panel with <2 returned reviews is a failed run, not a thin one.

Read ${CLAUDE_SKILL_DIR}/references/phases/phase2_6_panel.md when --panel is passed — it has the reviewer-set table, roster manifest, editor synthesis and lens-diversity gate.

Phase 3: Report

Before writing the comments, skim references/exemplar_findings/ for the finding at hand (cohort-arithmetic mismatch, unadjusted confounder, cross-sectional scope overreach, post-hoc primary / estimand drift). Each models the full shape — gate fired, the comment in a reviewer's words, Fatal/Fixable, category letter, fix, fixable_by_ai, R0-ready line. Match the structure, not the wording; they are synthetic.

Write the report to qc/self_review.md with this structure:

# Self-Review Report: {manuscript title}

**Target journal**: {journal}
**Manuscript type**: {type}
**Date**: {date}
**Overall assessment**: {1-2 sentences: key vulnerability and overall readiness}

## Anticipated Major Comments (fix before submission)

M1. **{Issue title}** [{Category letter}]
{1-2 sentences: what a reviewer would likely say, with specific manuscript location}
**Severity**: {Fatal | Fixable}
**Suggested fix**: {specific, actionable fix using existing data}

M2. ...

## Anticipated Minor Comments (address proactively)

m1. **{Issue}** [{Category}]: {1 sentence with location + fix}
m2. ...

## Editorial-Impression Risks (REMOVE / MOVE / TIGHTEN)

*The subtraction axis — what to take out, move, or tighten so the accurate manuscript reads
confidently. Advisory and non-blocking; from Phase 2.5g / category L. Omit this block only if the
scan returned nothing.*

L1. **{Issue}** [{REMOVE | MOVE | TIGHTEN}]: {1 sentence — what reads as over-defensive and where, with the subtraction to make}
L2. ...

## Strengths (emphasize in cover letter)

- {Specific strength 1}
- {Specific strength 2}
- ...

Keep the ADD / FIX axis (Major / Minor Comments) and the SUBTRACTION axis (Editorial-Impression Risks) visually separate; never fold L items into the Minor Comments.

In suggested fixes and in any text drafted in Phase 4:

  • References only from /search-lit with a confirmed DOI or PMID; mark any other reference [UNVERIFIED - NEEDS MANUAL CHECK] (Phase 2.5c blocks a FABRICATED one).
  • Clinical definitions, diagnostic criteria and guideline recommendations you cannot verify: flag with [VERIFY] and ask the user; never invent them.

Conciseness: each comment says what a reviewer would raise, where, and the fix: a Major in a few lines, a Minor in a sentence or two, an Editorial-Impression Risk in one sentence. List every Major the gates and the review produced (each P0 gate row is its own Major); never merge or drop one to fit a count. Strengths: the few worth repeating in the cover letter.

Phase 3b: R0 Numbering (Optional)

If the user plans to use /revise once real reviews arrive, offer to append R0-numbered output so they can later tell anticipated (R0) from novel (R1-only) comments:

## R0 Pre-Submission Findings (for /revise cross-reference)

R0-1 [MAJ] {mapped from M1}: {issue title}
R0-2 [MAJ] {mapped from M2}: {issue title}
R0-3 [MIN] {mapped from m1}: {issue title}
...

Phase 3c: Structured JSON Output (--json)

Emit machine-readable JSON only when --json is passed (or another skill consumes this run): append the block to the report and also write it to qc/self_review.json. Read ${CLAUDE_SKILL_DIR}/references/phases/phase3c_json_output.md for the schema, field semantics and worked example when --json was passed, or a downstream skill consumes this run. The JSON carries a coverage ledger (status per category A–L and per loaded probe module), and PASS is not allowed while any applicable entry is not_assessed. Finalise it with scripts/check_review_coverage.py --review qc/self_review.json --out qc/review_coverage.json --strict (exit 1 = PASS over an unworked category; fix the JSON before any consumer reads it).

Phase 4: Fix Support (on request)

The review ends at Phase 3. Read ${CLAUDE_SKILL_DIR}/references/phases/phase4_fix_support.md when --fix was passed or the user asks you to apply or draft fixes for the findings.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/aperivue/medsci-skills/self-review">View self-review on skillZs</a>