skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
aperivue/medsci-skills95 installs

cross-national

Use when comparing an exposure-outcome association across countries with parallel national surveys (KNHANES, NHANES, CHNS). Harmonizes variables, runs parallel weighted analyses and builds comparison tables for 2-country (KR+US) or 3-country (KR+US+CN) designs.

How do I install this agent skill?

npx skills add https://github.com/aperivue/medsci-skills --skill cross-national
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The skill is a medical research tool for cross-national statistical analysis. It carries a low security risk because it processes external survey data (CSVs) that could theoretically contain malicious instructions (indirect prompt injection) and generates R scripts to perform the analysis (dynamic execution).

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

Cross-National Comparison Study Skill

Inputs

  1. Research question: exposure → outcome association to compare across countries
  2. Korean data path: KNHANES CSV file
  3. US data path: NHANES CSV directory (multiple tables to merge)
  4. Harmonization table (optional): CSV mapping variables across surveys
    • Default: /replicate-study's references/harmonization_knhanes_nhanes.csv

Reference Files

  • /write-paper's references/paper_types/cross_national.md — writing template
  • /analyze-stats's references/analysis_guides/survey_weighted.md — survey-weighted analysis guide
  • references/chns_coding.md — CHNS files, merge keys, coding and warnings; read it when the design includes China (3-country design)
  • references/additional_variables.md — asthma, sleep, physical activity, diet, treatment and non-HDL-cholesterol coding, plus composite-score (LE8) warnings; read it when the study uses any of them

Workflow

Phase 1: Study Definition

  1. Confirm research question: Exposure → Outcome
  2. Define variable coding for both countries:
    • Exposure: PHQ-9, BMI category, smoking, etc.
    • Outcome: diabetes, hypertension, mortality, etc.
    • Covariates: age, sex, education, income, smoking, alcohol, obesity, CVD
  3. Check harmonization table for variable availability
  4. Output: study protocol summary for user approval

Phase 2: Data Preparation

Never guess a variable name, dataset column name, or variable coding. If a mapping is uncertain, output [VERIFY: variable_name] and ask the user to confirm it against the data dictionary.

KNHANES (single CSV):

  1. Load CSV and keep every row — the age ≥20 (or per-protocol) restriction is applied to the design in step 3

  2. Derive variables using KNHANES coding:

    VariableRaw VarCoding
    SmokingBS3_11,2=Current; 3=Former; 8=Never
    AlcoholBD1_112-6=Frequent (current drinker); 1=Occasional (past-year abstainer); 8=Never
    ObesityHE_obe1-3=Normal; 4-6=Obesity (BMI≥25, Asian cutoff)
    DepressionBP_PHQ_1~9Sum ≥10 = depression
    DiabetesHE_glu, HE_HbA1c, DE1_dgFPG≥126 or HbA1c≥6.5 or DE1_dg=1
    CVDDI4_dg, DI5_dg, DI6_dgAny = 1 → CVD yes
    Educationedu1-3=Non-college; 4=College
    IncomeincmQuartile (1=lowest … 4=highest): 1-3=Bottom 75%; 4=Top quartile
  3. Set survey design on the full file, then restrict to the analytic domain: des <- svydesign(id=~psu, strata=~kstrata, weights=~wt_itvex, nest=TRUE, data=df); des_ad <- subset(des, age >= 20). Never filter rows before svydesign() — dropping them changes the standard errors (see survey_weighted.md, subpopulation analysis)

NHANES (multiple CSVs):

  1. Load and merge tables by SEQN (DEMO_J, DPQ_J, GHB_J, GLU_J, BMX_J, SMQ_J, ALQ_J, DIQ_J, MCQ_J, BPQ_J, BPXO_J)

  2. Derive variables using NHANES coding. CRITICAL: NHANES data downloaded via R nhanesA package uses TEXT LABELS, not numeric codes.

    VariableRaw VarCoding (text labels)
    SexRIAGENDR"Male" / "Female" (NOT 1/2)
    SmokingSMQ020 + SMQ040100 cigs (SMQ020 "Yes" / "No") + now smoke (SMQ040 "Every day" / "Some days" / "Not at all")
    AlcoholALQ121 + ALQ111Frequent (current drinker): any ALQ121 frequency except "Never in the last year"; Occasional (past-year abstainer): "Never in the last year"; Never (lifetime non-drinker): ALQ111 == "No" (ALQ121 will be NA)
    ObesityBMXBMI (BMX_J, kg/m²)≥30 (WHO cutoff, NOT Asian)
    PHQ-9DPQ010~DPQ090"Not at all"→0, "Several days"→1, "More than half the days"→2, "Nearly every day"→3; sum ≥10 = depression
    DiabetesLBXGLU (GLU_J, fasting subsample, mg/dL), LBXGH (GHB_J, %), DIQ010LBXGLU≥126 | LBXGH≥6.5 | DIQ010=="Yes" (DIQ010: "Yes" / "No" / "Borderline"). CRITICAL: fasting glucose is GLU_J LBXGLU, not BIOPRO_J LBXSGL — CDC says the serum LBXSGL should not be used to determine undiagnosed diabetes; an analysis that uses LBXGLU needs the fasting-subsample weight (step 3). LBXGLU is NA for non-fasters, and in R NA | TRUE is TRUE but NA | FALSE is NA, so this composite goes missing only for non-fasters who would be non-diabetic; dropping those NAs under WTMEC2YR inflates prevalence. Either analyse the FPG composite only in the fasting domain with the fasting weight, or use the non-fasting variant LBXGH>=6.5 | DIQ010=="Yes" on WTMEC2YR, and say which in variable_mapping.csv
    CVDMCQ160B/C/D/EMCQ160B=="Yes" (CHF) | MCQ160C=="Yes" (CHD) | MCQ160D=="Yes" (angina) | MCQ160E=="Yes" (MI); labels "Yes" / "No" / "Don't know"
    HTNBPXOSY2+3, BPXODI2+3, BPQ020mean(BPXOSY2, BPXOSY3)≥140 | mean(BPXODI2, BPXODI3)≥90 | BPQ020=="Yes" (BPXOSY3 is the 3rd reading, not an average; the mean of the 2nd and 3rd matches KNHANES)
    EducationDMDEDUC25 text levels
  3. Set survey design on the full file, then subset() the design object to the analytic domain (e.g. RIDAGEYR >= 20): svydesign(id=~SDMVPSU, strata=~SDMVSTRA, weights=~WTMEC2YR, nest=TRUE). The weight follows the files: the single-cycle _J tables above take WTMEC2YR; WTMECPRP goes only with the pre-pandemic P_ files (P_DEMO, P_BMX, ...). A variable from the fasting subsample (GLU_J LBXGLU, TRIGLY_J LBXTR/LBDLDL) takes the fasting weight instead: WTSAF2YR (_J) or WTSAFPRP (P_); a variable that combines a fasting-subsample component with full-sample ones (diabetes above) is defined only in that fasting domain, so it cannot enter a WTMEC2YR model. Pooling cycles follows the NCHS rules: divide each cycle's weight by the number of cycles pooled (1999–2002 has its own 4-year weights), and to combine 2015–2016 with 2017–March 2020 use 2/5.2 × WTMEC2YR and 3.2/5.2 × WTMECPRP.

CHNS (3-country design): read references/chns_coding.md before preparing China data.

Phase 3: Parallel Analysis

For EACH country independently:

  1. Table 1: Baseline characteristics by exposure (weighted counts + percentages)
  2. Main analysis: Sequential logistic regression models
    • Model 1 (unadjusted)
    • Model 2 (age + sex)
    • Model 3 (fully adjusted: + education, income, smoking, alcohol, obesity, CVD)
  3. Subgroup analyses: By sex, age group, education, income, alcohol, smoking, CVD, obesity. Whether the association differs between subgroups is tested with an exposure × subgroup interaction term fitted on the full design (svyglm on the whole sample), not by comparing the subgroups' P values.
  4. Dose-response (if applicable): RCS with 3 knots

Phase 4: Cross-National Comparison Table

Generate a side-by-side comparison:

AnalysisKorea wOR (95% CI)US wOR (95% CI)Ratio of wORs (95% CI); P
Overall (fully adjusted).........
Male.........
Female.........
............

Compare the countries with the ratio of their odds ratios, not with whether the directions agree: two estimates in the same direction can differ, and opposite directions can be compatible. With log odds ratios b₁, b₂ and their design-based standard errors SE₁, SE₂ from the two independent surveys, the ratio is exp(b₁ − b₂) with 95% CI exp(b₁ − b₂ ± 1.96·√(SE₁² + SE₂²)) and z = (b₁ − b₂)/√(SE₁² + SE₂²) (Altman & Bland, BMJ 2003;326:219).

Every number comes from executed code output (analysis_korea.R, analysis_us.R) — never an invented p-value, effect size, confidence interval, or sample size.

Phase 5: Output Files

{working_dir}/
├── cross_national_report.md    — Study summary + comparison tables
├── variable_mapping.csv        — Variable mapping with match status
├── analysis_korea.R            — KNHANES analysis (self-contained)
├── analysis_us.R               — NHANES analysis (self-contained)
├── results/
│   ├── table1_korea.csv
│   ├── table1_us.csv
│   ├── main_results_comparison.csv
│   └── subgroup_comparison.csv
└── manuscript_draft/           — Optional: Methods + Results draft
    ├── methods_draft.md
    └── results_draft.md

Take every citation in the report or draft from /search-lit (confirmed DOI/PMID); mark any other [UNVERIFIED - NEEDS MANUAL CHECK], and never generate references from memory.

Critical Rules

  1. NEVER pool data across countries. Each country analyzed with its own survey design.
  2. Country-specific BMI cutoffs: Korea ≥25 (Asian), US ≥30 (WHO).
  3. Country-specific income: KNHANES quartile, NHANES PIR → harmonize to binary.
  4. Weighted analysis mandatory: Both KNHANES and NHANES are complex surveys. CHNS has no survey weights — analyse it unweighted (see references/chns_coding.md).
  5. Document all harmonization decisions: What matches, what needed recoding, what differs.
  6. Same analytic approach: Identical model specifications for both countries for fair comparison.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/aperivue/medsci-skills/cross-national">View cross-national on skillZs</a>