skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
aperivue/medsci-skills106 installs

calc-sample-size

Use when planning how many patients or cases a study needs before data collection (power analysis, IRB justification). Walks a decision tree to the right test and returns reproducible R/Python code and IRB-ready justification text. Analyzing collected data is /analyze-stats.

How do I install this agent skill?

npx skills add https://github.com/aperivue/medsci-skills --skill calc-sample-size
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The skill is a specialized statistical calculator for medical research that guides users through decision trees and generates reproducible R and Python scripts for sample size and power calculations. While the skill employs dynamic code generation and command-line execution, these functions are central to its purpose and utilize standard scientific libraries.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

Calc-Sample-Size Skill

Decision Tree

Walk the user through this tree one question at a time; do not assume answers. Reader studies, segmentation and model-comparison designs sit outside the tree: see Tests 14–17.

What is your primary outcome?
|
+-- Binary (yes/no, positive/negative)
|   |
|   +-- Paired data (same subjects, two methods)?
|   |   +-- YES --> [5] McNemar test
|   |   +-- NO  --> How many groups?
|   |       +-- 2 groups, superiority     --> [4] Two-proportion comparison (chi-square)
|   |       +-- 2 groups, non-inferiority --> [10] Non-inferiority / equivalence
|   |       +-- Multivariable model       --> single-predictor hypothesis test? --> [9] Logistic regression
|   |                                     --> clinical prediction / AI model for use?
|   |                                         +-- developing the model  --> [12] Prediction-model development (Riley)
|   |                                         +-- externally validating  --> [13] External-validation (Riley)
|   |
+-- Continuous (measurement, score)
|   |
|   +-- How many groups?
|       +-- 2 groups  --> [6] Independent t-test
|       +-- 3+ groups --> [8] One-way ANOVA
|
+-- Time-to-event (survival, recurrence)
|   |
|   +-- Two groups, unadjusted      --> [7] Log-rank test
|   +-- Multivariable / adjusted HR  --> [7] Log-rank (Schoenfeld) + [11] Cox EPV
|
+-- Agreement (inter-rater, reproducibility)
|   |
|   +-- Continuous measurements --> [2] ICC
|   +-- Categorical ratings     --> [3] Kappa
|
+-- Diagnostic accuracy (Se, Sp, AUC precision)
    |
    +--> [1] Diagnostic accuracy (precision-based)

Tests 1–11

Once the test is chosen, read ${CLAUDE_SKILL_DIR}/references/formulas.md § Test N — the parameter table with defaults, the effect-size interpretation, the formula, R/Python code and the methodological reference.

#TestUse when
1Diagnostic accuracy — Se/Sp precisiondesired 95% CI half-width for sensitivity or specificity
2ICC agreement (Walter 1998 test; Bonett 2002 CI width)inter-/intra-rater agreement on continuous measurements (tumor size, angle)
3Kappa agreement (Donner & Eliasziw 1992; needs the trait prevalence)agreement on categorical ratings (BI-RADS category, lesion present/absent)
4Two-proportion comparison (chi-square)two independent groups (AI vs conventional detection rate)
5McNemar (paired proportions)paired binary outcomes (two readers on the same cases, before/after)
6Independent t-testmeans in two independent groups (lesion size, malignant vs benign)
7Survival / log-rank (Schoenfeld events, then patients)time-to-event between two groups
8One-way ANOVAmeans across 3+ independent groups
9Logistic regression (Peduzzi EPV + Hsieh 1998, continuous or binary predictor)multivariable binary outcome, single-predictor hypothesis test
10Non-inferiority / equivalencenew method not worse than standard by more than a pre-specified margin, or equivalent within it
11Cox regression EPVmultivariable Cox model — enough events for stable estimates
  • Test 9: Peduzzi EPV ≥ 10 is a minimum baseline for a single-predictor hypothesis test only. For a clinical prediction / medical-AI model intended for use, EPV-10 is outdated and reviewer-vulnerable — use Test 12 (development) / Test 13 (validation). Always report both Peduzzi and Hsieh and recommend the larger N.
  • Test 10: NI alpha is one-sided (typically 0.025); orient the difference so > 0 favours the new method. Equivalence is TOST, powered jointly. The margin must be clinically justified (see formulas.md § Margin selection).
  • Test 11: EPV ensures model stability, not power for a specific HR. If an HR is available, also run Test 7 (Schoenfeld) and recommend the larger N.

Tests 12–17 (specialised designs)

Each has its own reference file (parameters, method, reporting); read it once the test is chosen.

  • Test 12 — Prediction-model development (Riley). Developing a clinical prediction / classification model (including a medical-AI model evaluated as one) for use. EPV-10 does not apply: N is the largest satisfying all Riley criteria — three for a binary or time-to-event outcome, four for a continuous one (R pmsampsize).
  • Test 13 — External validation (Riley). Validating an existing prediction/AI model: size for the CI width of the C-statistic, calibration slope, O:E and (if claimed) net benefit (R pmvalsampsize); ≥ 100 events and ≥ 100 non-events is only a floor. For Tests 12–13 read ${CLAUDE_SKILL_DIR}/references/prediction_model_sample_size.md.
  • Test 14 — MRMC reader study (Obuchowski–Rockette). Readers with vs without the AI, or AI non-inferior to readers. Test 1 under-sizes it because readers are a random effect: size readers J × cases from pilot/literature variance components with RJafroc / MRMCaov / iMRMC (do not hand-roll the OR algebra) and report the J × N power grid. Read ${CLAUDE_SKILL_DIR}/references/mrmc_reader_study_sample_size.md.
  • Test 15 — Segmentation-metric precision (Dice / HD95 / NSD). The outcome is a per-case score, not a proportion: n ≈ (1.96·SD/δ)² from the pilot SD of per-case Dice, sized on the worst structure; CI by patient-level bootstrap (BCa). This is precision; a comparison is Test 16. Read ${CLAUDE_SKILL_DIR}/references/segmentation_metric_sample_size.md.
  • Test 16 — Between-model comparison. Model A beats B, C, …: power the paired per-case difference, n = ((z₁₋α/₂ + z₁₋β)·SD_Δ/Δ)² (a CI sized to just exclude zero has ~50% power), not each model's precision; for > 2 models pre-specify one primary contrast or pay the family-wise correction; a ranking claim needs multiple seeds. Read ${CLAUDE_SKILL_DIR}/references/multi_model_comparison_sample_size.md.
  • Test 17 — Segmentation usability. Clinicians can use it (acceptability rate, catastrophic-failure bound, edit time): acceptability is a proportion sized per structure class; DE ≈ 1 + (m−1)ρ holds only when each case has its own readers — the same readers on every case (crossed) add a reader term more cases cannot shrink; bounding failures at ≤ 1% needs ~300 clean cases (rule of three). Read ${CLAUDE_SKILL_DIR}/references/segmentation_acceptability_sample_size.md.

Out of Scope

Do not compute adaptive trials (group-sequential, sample size re-estimation), cluster-randomized trials (design effect, ICC-based inflation), Bayesian sample size determination, crossover designs, or multi-endpoint correction (mention Bonferroni if asked, but do not compute corrected sample sizes). Say the design is beyond this skill and point to G*Power (free, https://www.psychologie.hhu.de/gpower), PASS, or a biostatistician.

Workflow

Phase 1: Understand the Study

  1. Ask the user to describe the study briefly (design, primary outcome, groups).
  2. Walk the decision tree to a test, and confirm it with the user before proceeding.

Phase 2: Collect Parameters

  1. Present the selected test's parameter table.
  2. For each parameter without a user-provided value, explain it and offer the default.
  3. Estimate effect sizes from prior literature (ask for the references) or pilot data; use Cohen's conventions only as a last resort, noting that convention-based estimates are less precise.

Phase 2b: Retrospective Studies

When the dataset already exists, formal power analysis is often impractical. Offer:

  • Fixed extract: read ${CLAUDE_SKILL_DIR}/references/observational_cohort.md and report event budget / confidence-interval precision instead of forcing a prospective recruitment-style power calculation.
  • Experience-based justification (acceptable for IRB and many journals):
    • Institution volume: total exams in period × prevalence × (1 − exclusion rate) = expected N. Ask for the annual exam volume for the modality, study period, prevalence and exclusion rate. This gives a realistic upper bound for N.
    • Prior studies: report the N of 3–5 comparable published studies and cite them; the user's N should be in the same range or larger.
    • IRB templates for both: ${CLAUDE_SKILL_DIR}/references/justification_examples.md § Retrospective.

Use a formal calculation (Phase 3) even for a retrospective study when a subset is enrolled prospectively, the primary analysis tests a hypothesis (not just estimation), the journal's Instructions for Authors require a power analysis, or the IRB requires it.

Phase 3: Calculate and Report

  1. Generate R code (primary) and Python code (alternative) from the reference formula.
  2. Run the R code via Bash; the reported N is the number it prints, not a hand calculation.
  3. Present the result in the Output Format below. Cite methodological sources only from formulas.md or the test's reference file; any other reference needs a DOI/PMID confirmed via /search-lit, otherwise mark it [UNVERIFIED - NEEDS MANUAL CHECK]. Mark an effect size, clinical definition or threshold you could not confirm [VERIFY].
  4. In a project, save the IRB text as protocol/sample_size_justification.md and the scripts as protocol/sample_size_calc.R / .py: /write-protocol and /write-paper embed that text verbatim, so the numbers are never retyped.

Phase 4: Sensitivity Analysis (Optional)

If a parameter is uncertain or the effect-size estimate is vague, flag it and offer a table of N across plausible values (e.g., varying effect size, or power from 0.80 to 0.90).

Output Format

Always structure the final output as follows:

## Sample Size Calculation Report

### Study Design
[1-2 sentence summary of the design and test selected]

### Parameters
| Parameter | Value | Source |
|-----------|-------|--------|
| ... | ... | user / literature / convention |

### Result
- **Required sample size**: N = [value]
- **With [X]% attrition adjustment**: N_adj = [value]

### R Code (Reproducible)
```r
# [complete, self-contained R script]
# Dependencies: [list packages]
# Run: Rscript sample_size_calc.R

Python Code (Alternative)

# [complete, self-contained Python script]
# Dependencies: [list packages]
# Run: python sample_size_calc.py

IRB Justification Text

A sample of [N] participants is required to detect [effect description] with [power]% power at a [one/two]-sided significance level of [alpha], assuming [key assumptions]. Accounting for an estimated [X]% attrition rate, we plan to enroll [N_adj] participants. This calculation is based on [formula/method reference].

Effect Size Interpretation

[Cohen's benchmark classification + clinical meaning in the context of this study]


The IRB text must state N; name the test and its formula source; give every assumed parameter
(effect size, alpha, power); state the attrition adjustment and final enrollment target; cite the
methodological reference (e.g., "Schoenfeld, 1981"); and use formal, third-person language. Read
`${CLAUDE_SKILL_DIR}/references/justification_examples.md` for per-design exemplars when writing it.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/aperivue/medsci-skills/calc-sample-size">View calc-sample-size on skillZs</a>