schema-unify
Migrate a brain from gbrain-base (or any pack) to gbrain-base-v2's 14-canonical-type taxonomy via gbrain onboard --check + the unify-types Minion handler. Collapses 94 noisy types to 15 canonical with subtypes, alias rows, and link rows. Triggers when an agent notices pack_upgrade_available, type_proliferation, or asks "what is the canonical taxonomy / how do I clean up my page types".
How do I install this agent skill?
npx skills add https://github.com/garrytan/gbrain --skill schema-unifyIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill provides a playbook for migrating database schemas between specific versions of the gbrain tool. It describes a multi-phase workflow involving discovery, preview, application, and verification using the gbrain CLI. All external references are to the author's own official documentation on GitHub. No malicious patterns or security risks were identified.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Schema Unification (gbrain-base → gbrain-base-v2)
v0.41.22 ships gbrain-base-v2 — a 15-type DRY/MECE taxonomy (14 canonical + note catch-all) — as the install default for new brains. Existing brains on gbrain-base can opt in via the pack_upgrade_available onboard finding + the unify-types PROTECTED Minion handler.
This skill is the playbook for that migration.
brain_first: exempt
This skill is ABOUT the brain's shape — it can't depend on the brain it's reshaping. No gbrain search lookup first; jump straight to onboard.
When this skill fires
- Agent runs
gbrain onboard --checkand seespack_upgrade_availableortype_proliferationwarnings - User asks "what is the canonical taxonomy / how do I clean up my page types / migrate to v2"
- A
dangling_aliasesfinding surfaces (post-unify GC) - An agent ingesting from a custom pack wants to consult the v2 taxonomy as a reference
Mental model (one paragraph)
A production gbrain brain accreted 94 distinct pages.type values over years of ingestion: tweet / tweet-thread / tweet-bundle / tweet-single / media/x-tweet/bundle / tweet-stub all coexisting; 5.5K concept-redirect pages; atom-partner-link pages that should be links; civic / framework / insight / memo / anecdote one-offs. The cure: collapse to 15 canonical types (person, company, media, tweet, social-digest, analysis, atom, concept, source, deal, email, slack, writing, project, note) with subtypes/format/origin pushed to frontmatter, alias-rows for redirects, real link-rows for edge-shaped pages, and a catch-all that bins long-tail unknowns to note with frontmatter.legacy_type = <original> for rollback.
Workflow
Phase 1: Discovery
Confirm the brain is actually on gbrain-base (not already on v2).
gbrain schema active --json | jq -r '.identity'
Expected: gbrain-base@1.0.0+<sha>. If you see gbrain-base-v2@..., the brain is already on v2 — skip the migration.
Then run onboard to see what would change:
gbrain onboard --check
Look for the pack_upgrade_available finding. If it's ok, there's no successor declared for the active pack — done.
Phase 2: Preview
Run the per-cluster narrative:
gbrain onboard --check --explain
This invokes the unify-types handler in dry-run mode and prints:
- How many pages would retype per cluster (tweets, articles, companies, etc.)
- How many concept-redirect pages would become alias rows
- How many edge-shaped pages would convert to real links
- The synthesized catch-all rules for unknown types
Review the output. If the proposed changes look wrong, don't proceed — file an issue or write a custom pack with adjusted mapping_rules.
Phase 3: Apply
The handler is PROTECTED (manual_only) — autopilot will never auto-fire it. Submit explicitly:
gbrain jobs submit unify-types \
--params '{"target_pack":"gbrain-base-v2","apply":true}'
On PGLite (the install default), or on any setup without a running gbrain jobs work worker or supervisor daemon, add --follow so the job executes inline:
gbrain jobs submit unify-types \
--follow \
--params '{"target_pack":"gbrain-base-v2","apply":true}'
The persistent worker daemon is Postgres-only. Without --follow on PGLite, the job sits queued forever and the migration never runs.
apply defaults to false (dry-run) per the handler contract, so
"apply":true is required here or the job reports success having retyped
nothing and left the active pack unflipped. Omit it to preview.
Watch progress per phase (worker-daemon runs; with --follow the same progress streams inline):
gbrain jobs get <job_id> # one job: status, progress, result
gbrain jobs watch --follow # live dashboard of the whole queue
A job that stays queued here means no worker is running; resubmit with --follow to execute it inline.
On a 186K-page brain expect ~10 minutes. The handler runs:
- Preflight (validate target pack has
mapping_rules:) - Stats snapshot (pre-state for celebration summary)
- Acquire
gbrain-unifydb-lock (60min TTL) - Apply phases:
- Explicit retype rules (tweets, articles, companies, etc.)
- Catch-all retype (unknown types → note with legacy_type)
- Page-to-link rules (atom-partner-link, symlink)
- Page-to-alias rules (concept-redirect)
- Final sync (untyped rows by path-prefix)
- Flip active pack to gbrain-base-v2
- Verify + celebration summary
Phase 4: Verify
gbrain onboard --check
gbrain schema stats
Expected:
pack_upgrade_available→ok(active pack is now v2)type_proliferation→ok(≤16 distinct typed values)dangling_aliases→ok(slug_aliases all point at active canonicals)gbrain schema statsshows ≤16 distinct types
Phase 5: Post-migration
Search and query --type article keep returning those pages post-unify: media declares article as an alias, so the type filter expands through the active pack's alias closure (the results also include other media pages). Direct SQL against pages.type needs updating to the canonical types.
Search queries get a small ranking signal: pages reached via slug_aliases (canonicals of one or more aliases) get a 1.05x boost. Visible via gbrain search --explain.
Rollback
Every retyped page preserves frontmatter.legacy_type = <original>.
Restore types in bulk (Postgres/Supabase deployments only; requires direct DB access):
UPDATE pages SET type = frontmatter->>'legacy_type'
WHERE source_id = 'default' AND frontmatter->>'legacy_type' IS NOT NULL;
On PGLite there is no SQL shell, so use the CLI surface instead: frontmatter.legacy_type persists per page, so individual retypes can be reverted through the normal put_page/CLI surface, and the soft-delete restore and pack-flip revert below work on every engine.
Page-to-alias and page-to-link source pages soft-delete with 72h TTL. Restore within that window:
gbrain restore <slug>
Revert the active pack flip:
gbrain schema use gbrain-base
Anti-patterns
- Don't run unify-types under autopilot. It's manual_only by design. Autopilot remediation should never silently change your taxonomy.
- Don't expect mapping_rules to cover every legacy type explicitly. Use the catch-all (
*unknown*) for the long tail. Pages get retyped tonotewithlegacy_typepreserved. - Don't rewrite body-text wikilinks. The slug_aliases table IS the resolver.
[[old-redirect-slug]]keeps working viaengine.resolveSlugWithAliasshort-circuit. - Don't bypass the dry-run. Always run
--explainbefore applying. The trust delta is real. - Don't run two unify jobs concurrently. The
gbrain-unifydb-lock serializes them; the second submission rejects with "already in progress."
Decision tree
Active pack already gbrain-base-v2?
→ Skip migration.
Custom pack with own mapping_rules?
→ Run --check --explain to see if your pack declares migration_from
for the active pack. If yes, target_pack = your pack name.
Brain has many custom types not covered by gbrain-base-v2 mapping_rules?
→ The catch-all retype binds them to `note` with legacy_type preserved.
Review by inspecting frontmatter.legacy_type after the migration.
Federated brain (multiple sources)?
→ Add --params source_id to scope the migration per-source. Each
source can be migrated independently.
Worried about a specific cluster's mapping?
→ Fork gbrain-base-v2 (`gbrain schema fork gbrain-base-v2 my-pack`),
edit mapping_rules in your fork, then target the fork.
Contract
Inputs:
- A brain on
gbrain-base(or any pack withmigration_from: gbrain-base-v2). - Trusted local CLI access on the brain host:
gbrain jobs submitgrants the PROTECTED-handler opt-in itself for protected names; the remote MCPsubmit_jobop cannot. - ~10 min wallclock on a 186K-page brain.
Outputs:
- Pages retyped to canonical types with
frontmatter.legacy_typepreserved (per-page rollback signal). slug_aliasesrows for concept-redirect pages (alias table IS the resolver — no link rewrite).- Real
linksrows for edge-shaped pages (atom-partner-link,symlink, etc.). - Active pack flipped to
gbrain-base-v2atomically at end of successful run.
Side effects:
- Source pages soft-deleted with 72h restore TTL (
gbrain restore <slug>). - One-time cache invalidation on KNOBS_HASH_VERSION bump (5→6); self-healing in
cache.ttl_seconds. - Search/query
--type Xexpands through the active pack's alias closure (back-compat).
Failure modes:
- Concurrent submission rejected by the
gbrain-unifydb-lock; second call exits gracefully. - Catch-all retype excludes
page_to_link+page_to_aliassource types (caught in E2E pre-merge). - Phase failures abort the run before
active_pack_flipped; partial state restorable via op_checkpoint resume.
When it fails
Follow the agent operator protocol for any gbrain error code, exit code, [AGENT] block or notice block. Specific to this skill:
- A second unify submission is rejected because the
gbrain-unifylock is held ("already in progress"): wait for the running job (gbrain jobs get <id>); do not resubmit. - A phase fails before
active_pack_flipped: the pack did not change; resume from the checkpoint rather than restarting from scratch. - The run reports a cost line: retyping can call a model, so confirm the budget with the user before submitting on a large brain.
Anti-Patterns
DON'T:
- Submit
unify-typesvia the remote MCPsubmit_jobop. PROTECTED handlers require trusted local callers (gbrain jobs submiton the brain host); remote MCP rejection is the intentional trust boundary. - Edit
mapping_rulesingbrain-base-v2.yamlto skip clusters you don't trust. Fork the pack instead (gbrain schema fork) so the source-of-truth migration stays consistent across brains. - Run
unify-typesfrom inside an autopilot tick. The check ismanual_only— autopilot deliberately never auto-fires it because pack upgrades are one-time consenting taxonomy decisions. - Hard-delete soft-deleted source pages before the 72h restore window. Use
gbrain restore <slug>first if rollback is needed. - Assume
frontmatter.legacy_typesurvives every roundtrip. The marker is canonical for the immediate post-migration window; downstream re-imports may overwrite it.
Output Format
Per phase, the handler emits to stderr:
[unify-types] phase=retype-explicit applied=N skipped=M cost=USD ttl=Ns
[unify-types] phase=retype-catch-all applied=N
[unify-types] phase=page-to-link converted=N pages soft-deleted
[unify-types] phase=page-to-alias aliased=N pages soft-deleted
[unify-types] phase=sync residual=N
[unify-types] active_pack flipped from gbrain-base to gbrain-base-v2
Final celebration summary to stderr:
═══════════════════════════════════════════════════════════
gbrain-base-v2 migration complete
═══════════════════════════════════════════════════════════
Before: 94 distinct page types
After: 15 canonical types
Retyped: 25,632 pages
Aliased: 5,521 redirects → slug_aliases table
Linkified: 65 ghost pages → real link rows
Soft-deleted: 5,586 pages (restorable for 72h)
═══════════════════════════════════════════════════════════
For structured JSON, gbrain call get_job '{"id": <id>}' returns the job row; its result field carries the UnifyTypesResult shape with per_phase, pack_identity_after, active_pack_flipped (gbrain jobs get <id> prints the same result inline).
Reference
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/garrytan/gbrain/schema-unify">View schema-unify on skillZs</a>