check-reporting
Use when auditing a manuscript item by item against a reporting guideline or risk-of-bias tool. Covers 49 reporting guidelines and risk-of-bias tools (STROBE, CONSORT, STARD, TRIPOD+AI, PRISMA, QUADAS and more), marking each item PRESENT, PARTIAL or MISSING. Not a reviewer critique (/self-review).
How do I install this agent skill?
npx skills add https://github.com/aperivue/medsci-skills --skill check-reportingIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill audits medical research manuscripts for compliance with reporting guidelines using a series of specialized Python scripts. While functionally safe, the skill is susceptible to indirect prompt injection from the manuscripts it processes, which is standard for agents ingesting untrusted text.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Check-Reporting Skill
Reference Files
Checklists are vendored under ${CLAUDE_SKILL_DIR}/references/checklists/, one file per
instrument. Each file's header gives its version, source citation and licence, and says when its
item text is an own-words summary rather than the published wording.
STROBE.md-- STROBESTROBE_MR.md-- STROBE-MR 2021RECORD.md-- RECORD 2015 (RECORD-PE for drug studies)REMARK.md-- REMARKTARGET.md-- TARGET 2025GATHER.md-- GATHER 2016CHEERS_2022.md-- CHEERS 2022CROSS.md-- CROSS 2021 + CHERRIES (internet surveys)SRQR.md-- SRQR 2014 (all qualitative approaches)COREQ.md-- COREQ 2007 (interviews and focus groups)STARD.md-- STARD 2015STARD_AI.md-- STARD-AI 2025TRIPOD.md-- TRIPOD 2015TRIPOD_AI.md-- TRIPOD+AI 2024TRIPOD_LLM.md-- TRIPOD-LLM 2025PGS_RS.md-- PGS-RS / PRS-RS 2021CONSORT.md-- CONSORT 2025CONSORT_AI.md-- CONSORT-AI 2020SPIRIT.md-- SPIRIT 2025SPIRIT_AI.md-- SPIRIT-AI 2020CLAIM_2024.md-- CLAIM 2024DECIDE_AI.md-- DECIDE-AI 2022MI_CLEAR_LLM.md-- MI-CLEAR-LLMCLEAR.md-- CLEAR (radiomics)ARRIVE_2.md-- ARRIVE 2.0CARE.md-- CARE 2013SQUIRE_2.md-- SQUIRE 2.0GRRAS.md-- GRRASPRISMA_2020.md-- PRISMA 2020PRISMA_2020_Abstracts.md-- PRISMA 2020 for Abstracts, 12 items. A separate instrument, not a subset of the 27-item checklist (main item 2 defers to it): score it with its own denominator.PRISMA_DTA.md-- PRISMA-DTAPRISMA_P.md-- PRISMA-PPRISMA_ScR.md-- PRISMA-ScRMOOSE.md-- MOOSESWiM.md-- SWiMAMSTAR2.md-- AMSTAR 2QUADAS3.md-- QUADAS-3QUADAS2.md-- QUADAS-2QUADAS_C.md-- QUADAS-CRoB2.md-- RoB 2ROBINS_I.md-- ROBINS-IROBINS_E.md-- ROBINS-EROBIS.md-- ROBISROB_ME.md-- ROB-MERoB_NMA.md-- RoB NMAPROBAST.md-- PROBASTPROBAST_AI.md-- PROBAST+AINOS.md-- NOSCOSMIN_RoB.md-- COSMIN RoB
Workflow
Step 0: Existing-checklist staleness pre-check
If a checklist already exists for this project (qc/reporting_checklist.json or a prior .md
report), verify it targets the current manuscript before reusing it — one generated against an
older version carries stale section/line references and a stale version label:
python3 "${CLAUDE_SKILL_DIR}/scripts/check_checklist_version.py" \
--checklist qc/reporting_checklist.json --manuscript manuscript_v8.md
A non-zero exit means the checklist is stale (older target_version, changed source_sha256,
different target_manuscript) or pre-dates the version contract: regenerate it against the current
manuscript (Steps 1–5) rather than reusing it. Every report you generate carries the
target_manuscript / target_version / source_sha256 fields (Part A header + Part D JSON) so
this check works next round.
Step 1: Select Guideline
Use the guideline the user names; otherwise auto-detect it from the study type. If the study type is ambiguous, ask the user to confirm before selecting a guideline.
Auto-detection mapping:
| Study Type | Primary Guideline | AI Extension |
|---|---|---|
| Observational study | STROBE | -- |
| Mendelian randomization study | STROBE-MR (base STROBE + MR extension) | -- |
| Health economic evaluation (cost-effectiveness / cost-utility / cost-benefit / budget-impact) | CHEERS 2022 | -- |
| Observational study using routinely-collected data (claims / EHR / registry / health-checkup DB) | RECORD (base STROBE + RECORD extension; RECORD-PE for drug studies) | -- |
| Survey / questionnaire study (KAP, physician/patient, cross-sectional, e-survey) | CROSS (+ CHERRIES for internet surveys) | -- |
| Scoping review (maps breadth/nature of evidence, clarifies concepts, identifies gaps — not a focused effectiveness/accuracy question) | PRISMA-ScR (base PRISMA + scoping-review extension) | -- |
| Qualitative study (interviews, focus groups, ethnography, grounded theory, phenomenology, document analysis) | SRQR (all qualitative approaches); COREQ (interviews/focus groups specifically) | -- |
| Randomized controlled trial | CONSORT 2025 | CONSORT-AI |
| Diagnostic accuracy study | STARD 2015 | STARD-AI |
| Prediction model (development/validation) | TRIPOD | TRIPOD+AI |
| Polygenic (risk) score prediction study | PGS-RS (with TRIPOD / TRIPOD+AI) | -- |
| Prognostic tumor-marker / biomarker study (single or multiple markers; e.g., ctDNA / molecular residual disease) | REMARK (pair with STROBE for the observational-design items; TRIPOD / TRIPOD+AI if a prognostic model is developed) | -- |
| Causal / comparative-effectiveness question emulated on observational data (treatment vs treatment, screening vs none, drug A vs B on registry / EHR / claims data) | TARGET (pair with the /design-study target-trial-emulation module for design; RECORD / STROBE for the routinely-collected-data items) | -- |
| Health-estimate / burden-of-disease modeling study (GBD or GBD-satellite, comparative-risk / population-attributable-fraction, cause-of-death or prevalence/incidence estimation, with or without forecasts) | GATHER (pair with /analyze-stats burden-decomposition-forecasting guide for the analytic layer) | -- |
| Systematic review / meta-analysis | PRISMA 2020 | PRISMA 2020 for Abstracts (run on the abstract, scored separately) |
| DTA systematic review / meta-analysis | PRISMA-DTA | PRISMA 2020 for Abstracts (run on the abstract, scored separately) |
| Meta-analysis of observational studies | MOOSE | PRISMA 2020 (use both) |
| Risk of bias (DTA studies) | QUADAS-3 (current recommended version) | QUADAS-2 only when appraising or reproducing a review that used it |
| Risk of bias (RCTs) | RoB 2 | -- |
| Risk of bias (non-randomised intervention studies) | ROBINS-I | -- |
| Risk of bias (non-randomised exposure studies) | ROBINS-E | -- |
| Risk of bias (comparative DTA studies) | QUADAS-C | QUADAS-3 (use both; apply the E&E's adaptation — see Using QUADAS-C with QUADAS-3 in QUADAS3.md) |
| Risk of bias (prediction models) | PROBAST | PROBAST+AI |
| Risk of bias (systematic reviews) | ROBIS | AMSTAR 2 |
| Risk of bias (missing evidence in MA) | ROB-ME | -- |
| Risk of bias (network meta-analysis) | RoB NMA | -- |
| Risk of bias (measurement properties) | COSMIN RoB | -- |
| Quality assessment (observational) | NOS | -- |
| Case report | CARE | -- |
| Study protocol | SPIRIT 2025 | SPIRIT-AI |
| Animal study | ARRIVE 2.0 | -- |
| AI/ML study in clinical imaging | CLAIM 2024 | -- |
| Study using a large language model (develop/fine-tune/prompt/evaluate an LLM) | TRIPOD-LLM | MI-CLEAR-LLM (use alongside when LLM accuracy is an outcome) |
| Early-stage / live clinical evaluation of an AI decision-support system (human factors, workflow, safety) | DECIDE-AI | -- |
| LLM accuracy evaluation in healthcare | MI-CLEAR-LLM | STARD-AI or CLAIM 2024 (use alongside) |
| Reliability / agreement study | GRRAS | -- |
| SR protocol | PRISMA-P | -- |
| Synthesis without meta-analysis | SWiM | PRISMA 2020 (use both) |
| Quality of systematic reviews | AMSTAR 2 | ROBIS |
| Radiomics study | CLEAR | CLAIM 2024 (if deep learning component) |
| Educational / QI study | SQUIRE 2.0 | -- |
| Generative AI images ARE the study object (realism / real-vs-synthetic reader study / model-vs-model quality) | (no single guideline -- assemble) | see decision aid below |
QUADAS-3 has two protocol-stage phases, and this skill usually runs too late for them. Phase 1 (state the synthesis question) and phase 2 (define the ideal test accuracy trial each judgement is made against) are review-level and belong in the protocol, with the review-specific guidance for answering each signalling question. Reaching them for the first time during manuscript QC means writing the comparator after seeing the results. If they are missing, say so as a limitation rather than reconstructing them, and route the protocol work to
/meta-analysisPhase 1. Phases 3–6 are what a QC pass can genuinely run.
Rules:
- A guideline the user names is the one you score. The rules below choose the guideline only when
the user names none. If they call for a different or additional instrument (STARD-AI for an AI index
test, TRIPOD+AI for an ML prediction model, an AI extension), score the named guideline and put one
line at the top of the report saying which instrument fits better and the command to rerun, e.g.
/check-reporting STARD-AI. Do not switch or add instruments on your own: the user maps the report to the checklist form the journal asked for. - If the study involves AI/ML, apply the AI version of the guideline. CONSORT-AI and SPIRIT-AI are scored alongside their base guideline; TRIPOD+AI, STARD-AI and TRIPOD-LLM replace theirs (below).
- Exception — TRIPOD: TRIPOD+AI 2024 (Collins et al., BMJ 2024) is a complete rewrite, not an addendum to TRIPOD 2015 (Moons et al., Ann Intern Med 2015). For non-AI prediction models, use TRIPOD 2015 only. For AI/ML prediction models, use TRIPOD+AI 2024 only. Do NOT apply both simultaneously.
- STARD-AI (Sounderajah et al., Nat Med 2025) extends STARD 2015 with 14 new and 4 modified items (40 total) and incorporates all STARD 2015 items. For AI diagnostic accuracy studies use STARD-AI only — do NOT apply STARD 2015 and STARD-AI simultaneously.
- TRIPOD-LLM (Gallifant et al., Nat Med 2025) is the reporting guideline for studies that develop, fine-tune, prompt, or evaluate a large language model for a clinical/biomedical task. It extends the TRIPOD family (TRIPOD 2015 → TRIPOD+AI 2024 → TRIPOD-LLM 2025); name the base instrument and the extension and cite each. It is modular — task-specific items (Annotation, Prompting, Summarization, Instruction-tuning) are N/A when that component is absent. Use TRIPOD-LLM for LLM studies in place of TRIPOD+AI; pair with MI-CLEAR-LLM when LLM accuracy is an evaluated outcome. The vendored checklist is an educational summary (own-words paraphrase of item intent); complete the official instrument for a submission checklist.
- MI-CLEAR-LLM is a supplementary checklist (8 item categories in the 2025 update; the 2024 original had 6), not a standalone reporting guideline. Always pair it with the study's primary guideline (e.g., STARD-AI for AI diagnostic accuracy, CLAIM for imaging AI). Apply it whenever the study evaluates LLM accuracy as an outcome — do NOT apply it merely because the manuscript was written with LLM assistance. Its scope is LLM accuracy studies (including VLMs interpreting images); it does not apply at study level when a generative model produces the images under study (next bullet).
- Generative-AI images as the study object (a generative model synthesizes images and the study evaluates their realism, controllability, real-vs-synthetic distinguishability, or model-vs-model quality) has no single dominant checklist. Assemble: CLAIM 2024 (imaging-AI umbrella; model-development items N/A when commercial models are used as-is) + FUTURE-AI traceability + MI-CLEAR-LLM transparency items only (prompt/model/version/params/runs — for generation provenance, not study-level compliance) on the generator side; STARD-AI (for real-vs-synthetic detection) + GRRAS (reader reliability) + MRMC reporting on the evaluation side. Map applicable items and cite base + extension; never claim wholesale compliance. Full decision aid:
${CLAUDE_SKILL_DIR}/references/genai_image_study_object_decision_aid.md. - If multiple guidelines apply (e.g., a diagnostic accuracy study that is also an AI study), check against all relevant guidelines and merge into one report.
Step 2: Load Checklist
-
Run the fail-fast guard first for every guideline you intend to apply:
python "${CLAUDE_SKILL_DIR}/scripts/check_checklist_exists.py" --guideline "STARD-AI"- Exit 0 → the vendored checklist exists; read it from
${CLAUDE_SKILL_DIR}/references/checklists/and proceed. - Exit 1 (
MISSING_CHECKLIST_CONTRACT_VIOLATION) → the guideline is routed but no checklist file is vendored. Do not construct items from memory. Halt, report the violation to the user, and stop unless they explicitly opt in (next bullet). - Exit 2 (
UNKNOWN_GUIDELINE) → the name is not recognised; confirm the correct guideline with the user.
- Exit 0 → the vendored checklist exists; read it from
-
No silent fallback. A from-memory checklist is permitted only when the user explicitly accepts it — re-run the guard with
--allow-from-memory(exit 0 + a NON-AUTHORITATIVE warning). The report MUST then carry a prominent banner that the assessment was constructed from model knowledge and is not backed by a vendored checklist, andsubmission_safemust not be asserted on its basis.
Step 3: Scan Manuscript
Read the whole manuscript before assessing any item — including tables, figures and their captions, supplementary material, and the reference list (registration numbers and protocol references often sit there).
Step 4: Assess Each Item
Items most often missing in medical manuscripts — look for these first, whichever guideline applies: registration number and registration/amendment date consistency (run Step 4c), sample-size justification, missing-data handling, blinding, funding and conflicts of interest, ethics approval with committee name and approval number, and a data availability statement; for AI studies, the training/validation/test split, model architecture and hyperparameters, failure-mode analysis, fairness/bias assessment, and commercial interests with data/code availability.
For every checklist item, determine:
| Status | Criteria |
|---|---|
| PRESENT | The item is fully addressed with sufficient detail. |
| PARTIAL | The item is mentioned or partially addressed but lacks required detail. |
| MISSING | The item is not found anywhere in the manuscript. |
| N/A | The item does not apply to this particular study (justify why). |
For each item, record:
- Status: PRESENT / PARTIAL / MISSING / N/A
- Location: Section name and paragraph or approximate position (e.g., "Methods, paragraph 3")
- Notes: What was found (if PRESENT/PARTIAL) or what should be added (if MISSING)
Be strict. PARTIAL means the item is mentioned but lacks specificity; a vague reference does not count as PRESENT — the detail level must match what the guideline expects. "We used appropriate statistical tests" = PARTIAL (which tests?); "We used the Mann-Whitney U test for continuous variables and Fisher's exact test for categorical variables" = PRESENT. If an item is genuinely unclear in its applicability, mark it N/A with justification.
Two gaps that are easy to pass:
- Power-aware framing of a null result (STROBE 16a / 18 / 20) — for an observational study whose headline is a non-significant association, a flat "X was not associated with Y" overreads the data when the analysis is not powered to exclude a clinically meaningful effect. Mark item 18/20 PARTIAL unless the manuscript states the precision as an exclusion (e.g., "the 95% CI excluded an eGFR difference larger than ~1.7") or reports a minimum detectable effect — "no effect" vs "could not exclude an effect of size X" are different claims, and a negative conclusion needs the latter.
- Confounder-selection rationale, not "adjust for everything that differs" (STROBE 16a explicitly asks which confounders were adjusted for and why) — flag a kitchen-sink adjustment set chosen because variables differ in Table 1. The Methods must give a causal rationale (DAG / prior literature) and must not adjust for a mediator or consequence of the outcome (over-adjustment, e.g. serum uric acid in an eGFR model); both an unjustified inclusion and an unjustified omission are item-16a gaps.
What is appraised is the source paper's reporting — never your convenience in using it. This holds for every instrument here, reporting checklists and risk-of-bias / quality tools alike, and is easiest to lose in a systematic review, where you read each paper in order to extract from it. An item asking "are the results clearly reported?" is not asking "were they reported in the unit my pool needs". If a downgrade's stated reason turns on a denominator, an analysis unit, a subgroup you needed and they did not report separately, or a format you could not parse, it is an extraction note, not a scoring reason: record it in a separate extraction-note column and restore the score. An extraction limitation often belongs in your limitations paragraph; folded into the score it makes the appraisal unreproducible, because another assessor with a different pool would score the same paper differently.
Step 4b: Section Boundary Check
In addition to checklist items, verify that:
- Results section contains only factual findings: no interpretation, no "why" explanations, no prior literature comparisons, no evaluative adjectives without numbers.
- Discussion section does not introduce new data not presented in Results.
- Flag any boundary violation as a separate finding in Part C Action Items with the label
[BOUNDARY].
Step 4c: Registration / Protocol Timing Consistency Check
Applies to: systematic reviews, meta-analyses, and intervention studies with prospective registration (PRISMA 2020, PRISMA-DTA, PRISMA-P, MOOSE, CONSORT, SPIRIT). The registration identifier is a single checklist item and can pass Step 4 while the manuscript is inconsistent about when registration and amendments happened relative to the analysis.
Read ${CLAUDE_SKILL_DIR}/references/step4c_registration_timing.md (item-by-item procedure,
JSON schema, flagging edge cases) and run its five checks: (1) registration identifier present in
Methods, Abstract, and cover letter; (2) initial registration date precedes — or is explicitly
disclosed as post-dating — the extraction milestone; (3) amendment dates appear in Methods, the
described change is visible in Methods, analysis was re-run if the amendment post-dates the lock,
and no amendment post-dates submission; (4) Methods agree with the registry record (PROSPERO PDF,
ClinicalTrials.gov export) — a silent discrepancy is a finding; (5) a retrospective-registration
disclosure paragraph when evidence suggests post-extraction filing.
Flagging: any failure is logged in Part C Action Items with label [REGISTRATION-TIMING].
fixable_by_ai: false when reconciliation requires an external amendment filing; true only when
the fix is a Methods-text insertion of a date already disclosed elsewhere. Part D JSON includes a
registration_timing object (registry, id, initial_registration_date, amendments[],
timing_consistency, findings[]).
Registration-ID format gate: a PROSPERO ID is CRD42 + 9 digits = 14 characters
(^CRD42\d{9}$, e.g. CRD42024500001). Run grep -oE 'CRD42[0-9]+' manuscript.md and
assert each match is 14 characters long; a 15-character ID (a stray inserted digit) is a
transcription error logged as [REGISTRATION-TIMING] (fixable_by_ai: false — verify against
the live PROSPERO record, do not guess the correct digit).
Step 4d: PRISMA Figure 1 Arithmetic & Cross-Reference Audit
Applies to: systematic reviews and meta-analyses using PRISMA 2020 / PRISMA-DTA / PRISMA-P. Triggers when Item 16a (flow diagram) is PRESENT — the diagram can pass Step 4 while its numbers do not add up or disagree with the text.
-
Choose the Figure 1 source, in this order: (a)
analysis/figures/Figure1_PRISMA.mdmarkdown manifest, (b) caption text inmanuscript.md, (c) PPTX text run if a.pptxexists, (d) manual entry from PNG/SVG. -
Run the audit. It checks the five subtractions (screened = identified − duplicates; sought-for-retrieval = screened − excluded at screening; retrieved = sought − not retrieved; assessed for eligibility = sought − not retrieved; included = assessed for eligibility − excluded with reasons) and that the body-text PRISMA numbers match the Figure 1 boxes 1:1, and writes
qc/prisma_figure_audit.json:python3 ${CLAUDE_SKILL_DIR}/scripts/check_prisma_figure.py \ --md <manuscript.md> --figure <Figure 1 source: .md manifest / caption / text export> \ --out qc/prisma_figure_audit.jsonExit
1= an arithmetic or cross-reference MISMATCH; exit2= missing/unparsable input. When the Figure 1 numbers exist only in a PNG/SVG, transcribe them by hand and run the same checks manually as${CLAUDE_SKILL_DIR}/references/step4d_prisma_figure_audit.mdspecifies (regex set, JSON schema, and edge cases: duplicates across databases, the citation-searching strand, dual-reviewer screening). -
By hand (the script does not do this): the reasons for exclusion in Methods and the Figure legend must agree on counts and category names, and when identification is split across sources (databases, registers, other methods) the per-source counts must sum to the
identifiedtotal — the script reads only the firstidentifiedcount. -
If
analysis/figures/_figure_manifest.md(from/make-figures) exists, verify that the row whoseType = prisma(orType = prisma-dta) points at the same file used as the audit source, and that itsCriticfield isyesorpartial(notno). A missing row, mismatched path, orCritic = nologs[MANIFEST-XREF](advisory); the arithmetic check still runs.
Flagging: any MISMATCH or arithmetic failure logs a Part C Action Item with label
[PRISMA-FIGURE], fixable_by_ai: false (the author must reconcile the numbers).
PRISMA Cascade Arithmetic Auto-Verify
When PRISMA 2020 or PRISMA-DTA is selected and round-by-round screening TSV artifacts are
available, recompute the screening cascade from the raw decisions. Off-by-one errors in the prose
cascade are a high-frequency reviewer red flag (e.g., 151 + 108 + 39 + 1 + 1 + 4 = 304
followed by a prose summary "305" four lines later).
python "${CLAUDE_SKILL_DIR}/scripts/prisma_cascade_check.py" \
--round1 2_Screening/round1.tsv \
--round2 2_Screening/round2.tsv \
--round3 2_Screening/round3_adjudication.tsv \
--manuscript manuscript.md \
--out qc/prisma_cascade.json --strict
The script counts INCLUDE / EXCLUDE / MAYBE decisions per round, computes the cascade, and
reports per-stage drift where the manuscript's stage counts disagree. Treat any
manuscript_drift entry as a P0 blocker — fix the prose to match the computed cascade and re-run.
A --manuscript path that is not a file exits 2. When the manuscript is read but none of the
script's stage phrases is found in it, manuscript_check is "unverifiable",
submission_safe is false, an UNVERIFIED line is printed, and --strict exits 1: the
drift check did not run, so compare the prose stage counts to stage_counts by hand.
stages_compared lists the stages that were actually checked.
Step 4e: Reporting-Framework Naming Audit
Applies to: any manuscript that invokes an AI/extension reporting framework
(PROBAST+AI, STARD-AI, TRIPOD+AI, TRIPOD-LLM, CONSORT-AI, SPIRIT-AI, PRISMA-DTA, QUADAS-C).
A base reporting tool and its extension are distinct instruments with separate citations, and
Step 1 does not police how the framework is named in prose. The recurring failures: invoking an
extension without ever naming or citing the base instrument it extends; mixing +AI and -AI
hyphenation for one family within a single document; coining item labels like "12-AI"; and waving
at "recent guidance" instead of naming the framework.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_framework_naming.py" \
--manuscript manuscript.md --out qc/framework_naming.json --strict
Verdicts: BASE_MISSING (extension used, base instrument never named standalone) is a
Major and logs [FRAMEWORK-NAMING] in Part C with fixable_by_ai: true (insert the base
name + its citation). HYPHEN_MIX, CITE_MISSING, SELF_COINED_LABEL, and VAGUE_GUIDANCE
are Minor (fixable_by_ai: true). Part D JSON includes a framework_naming object mirroring
the script's claims[].
Step 4f: Critical-item floor cross-check
Applies to: every guideline assessment for which the floor defines a row (load and
check only those; do not invent a floor for an unlisted guideline). After the item-by-item
table, read ${CLAUDE_SKILL_DIR}/references/critical_item_floor.md and check the small set of
non-waivable items for this study type — and, for AI/ML and radiomics manuscripts, the
methodological-quality / risk-of-bias instrument (PROBAST+AI, METRICS/RQS, APPRAISE-AI) and
concerns it lists. A MISSING critical item is surfaced as a Critical gap and becomes the
report's headline regardless of the overall percentage — a high percentage with a missing
critical item (undefined reference standard, no leakage-controlled partition, calibration absent
for a prediction model, an unreconciled flow diagram) is not "broadly acceptable."
Step 5: Generate Report
This report is an internal working audit — it carries auto-fix annotations, a
machine-readable JSON block (compliance_pct, fixable_by_ai, …), and Action Items. It is
NOT the official reporting checklist a journal expects (that is the blank guideline form with
Item | Recommendation | Reported in page/section, which the authors fill in). Never submit
this report as the submission checklist. So that the file is self-identifying and cannot be
reused by filename into a later submission package, the report MUST begin with this banner as
its very first line:
<!-- INTERNAL AUDIT — NOT FOR SUBMISSION. This is the /check-reporting working
report, not the official journal checklist. Do not upload to a submission portal. -->
Write the checklist content and report in English (matching the guideline originals), whatever
language you use with the user. Save it as qc/reporting_checklist.md and the Part D JSON as
qc/reporting_checklist.json.
Read ${CLAUDE_SKILL_DIR}/references/report_templates.md when you write the report — it holds the
literal templates for the four parts:
- Part A — Summary. Header (manuscript file, version token, guideline, date), the
PRESENT/PARTIAL/MISSING/N-A count table, and overall compliance. The headline is the critical
items (Step 4f), not the percentage: report
{present}/{total}and name every missing critical item with the section it belongs in. - Part B — Item-by-item checklist. One row per item:
# | Section | Item | Status | Location | Notes. - Part C — Action items (MISSING and PARTIAL only), ordered by: items most journals enforce strictly (ethics approval, registration, sample size) → items in Methods (easiest to fix) → everything else.
- Part D — Machine-readable JSON, appended as a fenced block. MUST be present under
--jsonor when called from/write-paperPhase 7, which parses it.
JSON field contract (the part other skills depend on — get these right):
compliance_pct—present / (total_items - na) * 100, one decimal;nullwhen every item is N/A (total_items == na).action_items— MISSING and PARTIAL only; PRESENT and N/A are excluded.fixable_by_ai—truewhen the fix inserts or expands text using information already in the manuscript or inferable from it;falsewhen it needs external facts the author alone holds (registration number, IRB approval number, protocol details, sample-size rationale and target).suggested_fix— concrete draft text, insertable as written. A fix that still contains a bracketed placeholder ([N],[rationale]) is not insertable and isfixable_by_ai: false.source_sha256— first 12 hex chars of the SHA-256 of the manuscript bytes, so a stale report cannot be silently attributed to a newer manuscript.
In notes and suggested fixes, cite a reference only with a /search-lit-confirmed DOI or PMID and
mark any other [UNVERIFIED - NEEDS MANUAL CHECK]; never invent clinical definitions, diagnostic
criteria, or guideline recommendations — flag anything uncertain [VERIFY] and ask the user.
When the user asks for the filled checklist a journal requires, build it from the Submission
checklist export template in ${CLAUDE_SKILL_DIR}/references/report_templates.md.
Skill Interactions
| When | Call | Purpose |
|---|---|---|
| During manuscript writing | /write-paper Phase 7 | Final compliance check |
| Need to add Methods text | /write-paper Phase 3 | Draft missing Methods content |
| Need statistical details | /analyze-stats | Generate missing statistical reporting |
| Need flow diagram | /make-figures | Generate CONSORT/STARD/PRISMA diagram |
Gates
| Gate | Severity | Trigger | Action on fail |
|---|---|---|---|
| Mandatory items present | ENFORCED at submission | < 100% of guideline-mandatory items marked PRESENT | Auto-fix MISSING items where text exists; otherwise route to /write-paper Phase 7 for re-draft |
| Step 4d PRISMA Figure 1 arithmetic & cross-reference audit (PRISMA / PRISMA-DTA only) | ENFORCED for SR/MA | flow numbers don't sum (e.g., screened ≠ included + excluded), or in-text counts mismatch flow diagram | HALT; reconcile against extraction artifacts |
| Optional items (e.g., supplementary AI declarations) | ADVISORY | < 80% of optional items present | warn; user accepts |
| Cross-reporting-guideline routing (study type → guideline) | ENFORCED | study type undeclared or guideline missing | Ask user; do not silently default |
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/aperivue/medsci-skills/check-reporting">View check-reporting on skillZs</a>