backlink
Use for backlink work - finding, qualifying, submitting, analyzing or verifying backlinks, directory and blog-comment placements, competitor link sources, anchors, toxic links and disavow, outreach templates, and reading Ahrefs, Semrush or Similarweb backlink and traffic reports (外链、反链、去哪发、提交目录、评论外链、发出去没有、毒性、disavow、竞品外链、竞品导流、数据面板). This Skill owns what to read from those reports and which placements to pursue. Not for generic logged-in browser driving, sessions or doctor faults (use opencli), and not for sitemap, IndexNow or Search Console indexing operations (use rankup); index-submission here is only a reference kept apart from backlinks.
How do I install this agent skill?
npx skills add https://github.com/yan-labs/yan-skills --skill backlinkIs this agent skill safe to install?
- Gen Agent Trust Hubfail
The skill is a comprehensive toolkit for managing backlink campaigns, including discovery, qualification, and form submission. It uses the OpenCLI browser tool to interact with SEO dashboards and automate web forms. Security measures like token redaction are implemented to protect user credentials. Automated scan flags on data files and target URLs appear to be contextually benign or false positives related to the nature of SEO data.
- Socketwarn
6 alerts: gptSecurity, gptAnomaly
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Two former Skills were merged in on 2026-08-16 and deleted: backlink-analyzer
(analysis templates, toxicity rubric, outreach — now in three references under
its original Apache-2.0 licence) and browser-harvest (pulling tables out of
logged-in dashboards — now <ref file="references/harvest.md"/>). The harvest
knowledge is general-purpose: ad platforms, e-commerce backends, any no-API
SaaS report. When a harvesting task has nothing to do with links, load this
Skill anyway and read that one reference.
</mission>
| Ask | Start |
|---|---|
| 去哪发、批量提交 | node scripts/targets-select.mjs --stats; <ref file="references/submission-lanes.md"/>; 100+ targets: <ref file="references/batch-campaign.md"/> |
| 免费免注册、即时发布 | data/free-channels.json (account:none, status:live); <ref file="references/instant-publish.md"/> |
| 竞品外链、新机会 | <ref file="references/discovery-loop.md"/> and Semrush platform manual; write leads back to registry |
| 质量、毒性、要不要 disavow | <ref file="references/link-quality-rubric.md"/> and data/network-fingerprints.json |
| 发出去了没有 | <workflow-ref id="verify"/>; inspect exact public anchor and rel |
| 索引提交 | <ref file="references/index-submission.md"/>; it publishes no backlink |
| Semrush/Similarweb 报表、流量 | <ref file="references/authorized-data-sources.md"/>; platform manuals under ../platforms/ |
| 没有 API 的后台表格 | <ref file="references/harvest.md"/> |
Query the database instead of guessing from examples: <cmd><![CDATA[ node scripts/targets-select.mjs --stats node scripts/paid-platform-registry.mjs list --min-sites 2 ]]></cmd> </routing>
<browser-runtime> <summary>Use OpenCLI with the owner's authorized Chrome. Run `node scripts/health.mjs` before browser work; read <ref file="references/browser-runtime.md"/>. The full measurements, failure cases, and commands are in <ref file="references/browser-runtime-detail.md"/>. For dated Semrush/Similarweb route findings read <ref file="references/platform-capabilities.md"/>.</summary> <law id="one-session-one-tab">One descriptive session name owns one tab; give parallel pages distinct sessions and close each after use.</law> <law id="tools-share-is-a-global-mutex">Hold the shared dashboard tool lock for a whole collection; use the existing launcher and its session.</law> <law id="no-multi-tab-api">Do not use `tab new`, `tab select`, or `open --tab` to hold several pages under one session name; use separate sessions.</law> <law id="no-literal-session-name">Never use a shared literal default session name; use `defaultSession(base)` or an explicit descriptive, unique name.</law> <law id="claim-handles-first">Claim all needed session handles before starting a parallel browser loop.</law> <law id="background-by-default">Use background for routine page work; report routes that need visible hydration use their script's virtual-display/active path.</law> <law id="hidden-tabs-do-not-hydrate">A hidden, empty report is inconclusive. Check actual visibility and the value-bearing region; follow the route-specific readiness rule before calling it empty.</law> <law id="readiness-must-bind-to-this-query">Bind a report to the requested route, target, scope, and content before classifying values. Headers, skeletons, or a prior target's numbers do not prove readiness.</law> <law id="every-measurement-needs-two-witnesses">For unfamiliar report routes use `scripts/ground-truth.mjs`: pair a deep DOM census with a screenshot; scripts collect and the AI judges.</law> <law id="one-collector-per-quota-tool">One collector at a time for a shared quota tool; keep one session through the task and budget quota before the batch.</law> <law id="scripts-collect-ai-judges">Treat machine classifications and empty states as suggestions until the evidence scene supports the verdict.</law> </browser-runtime> <data-sources> <terminology lang="zh"> **当用户说「数据面板」「数据勘测」「查一下数据」「用 Similarweb 看看」「Semrush 拉一下」, 指的都是同一件事:走那个共享账号的代理面板,用 Similarweb 或 Semrush 查。** 这两个产品是这里唯一的第三方数据源,没有别的候选,不需要反问用户指的是哪个平台。 </terminology> <division lang="zh"> 分工固定,按问题类型选,一次只开一个:| 问题 | 用哪个 | 拿得到什么 |
|---|---|---|
| 这个站多大、流量从哪来、还有哪些同类站 | Similarweb | 总访问量(含直接/推荐)、渠道构成、相似站、地理分布 |
| 这个词多少量、多难、谁在排、它的外链长什么样 | Semrush | 分国家搜索量与 KD、关键词全库导出、自然排名、主要页面、引荐域名与反链 |
两边的「流量」口径不同,对不上很正常。 Semrush 域名概览给的是自然搜索流量估算,
Similarweb 给的是总访问量。同一个站两边差三倍以上是常态,写结论时必须标明口径,
否则会得出「竞品比想象中弱」这种错误判断。绝不放进同一列。
</division>
<panel-launch>
Use node scripts/tools-share-open.mjs --tool semrush or --tool similarweb; reuse the same session for a task and check the landed route, target and quota. Semrush organic estimates and Similarweb total visits have different denominators. Panel launch, node quotas, and the Semrush overview criteria: <ref file="references/data-source-operations.md"/>. Tool account and source details: <ref file="references/authorized-data-sources.md"/>.
</panel-launch>
</data-sources>
bulk: feed an authorized referring-domains export straight in.
Edges are typed refdomain — do NOT route these through import-commenters,
which would record a commenter relationship nobody observed.
node scripts/discovery-queue.mjs import-refdomains --file .backlink/discovery.json
--source competitor.com --input .backlink/competitor-refdomains.csv
node scripts/harvest-commenters.mjs --session "discovery-commenters" --url https://example.com/article --out .backlink/commenters.json
node scripts/discovery-queue.mjs import-commenters --file .backlink/discovery.json --input .backlink/commenters.json
node scripts/discovery-queue.mjs next --file .backlink/discovery.json --limit 10
]]></cmd>
<footprint>
A second, non-recursive lane: search-operator footprints instead of competitor
backlink rows. footprint-discover.mjs prefers Serper.dev's API (a real
Google SERP over HTTP, no browser, no CAPTCHA) when SERPER_API_KEY is set —
its free tier caps operator queries at 10 results. Without a key it falls
back to the owner's own logged-in Chrome via OpenCLI, where a session hits a
CAPTCHA / "unusual traffic" wall after roughly 4 operator queries, and a
flagged exit IP gets blocked even from a fresh independent profile
(2026-09-12 agent-browser finding). General search APIs and Bing/DuckDuckGo
still do not execute inurl:/intitle: operators either way. See
references/discovery-loop.md § "Footprint discovery" for the effective/noisy
footprint table and the CAPTCHA policy before running a real sweep.
<cmd><![CDATA[
node scripts/health.mjs # confirm opencli before opening a browser
node scripts/footprint-discover.mjs --keyword "browser games" --preset submit
--num 20 --out .backlink/footprint-browser-games.jsonl
writes JSONL incrementally; stops and leaves a scene under
.backlink/footprint-browser-games.jsonl.evidence/ on any CAPTCHA signal.
--resume picks a stopped run back up later without re-running done queries.
]]></cmd>
</footprint>
<recon>
Domain overview is one page; semrush-report.mjs covers the other eight, which
have no export button and are where competitor recon actually happens. Five of
those eight (organic-overview/organic-positions/organic-pages/keyword-magic/
keyword-overview) have no worldwide option and land on an unpredictable country
without --db — pass it explicitly or the script exits with an error; the
other three (backlinks-list/referring-domains/backlinks-overview) aren't
country-scoped. Pass the same
--session across the whole recon — the panel launch costs 20–40s and a
login, the report itself ~15s, and semrush-report.mjs skips the launch when
the session is already parked on the tool origin (sessionReused: true says
which happened).
<cmd><![CDATA[
Semrush is a quota site: the script resolves the session to the fixed
semrush-nav itself, so do NOT pass --session. Passing one is ignored with a
warning; the fixed name is what serialises concurrent callers into one tab.
node scripts/semrush-report.mjs --report keyword-overview --keyword 'grid maker' --db us node scripts/semrush-report.mjs --report backlinks-overview --domain rival.com node scripts/semrush-report.mjs --report organic-positions --domain rival.com --db us opencli browser semrush-nav close ]]></cmd> <note> This is the one place a session legitimately handles several reports — it is still one page at a time, navigated in sequence, which is what <law-ref id="one-session-one-tab"/> allows. Holding them open simultaneously would need N session names. </note> </recon> <caution> These metrics help discover and prioritize candidates. They never prove a backlink is public, indexed, followable, or causally producing traffic. The parsing traps that make a report silently return zeros are documented in <ref file="references/authorized-data-sources.md"/> — read it before writing any new reader, especially the rule that a readiness predicate must key on a data row, never on a tab name, column header, or filter chip.
Then check the parser against itself. A ready page and a correct parse are different claims, and the second one fails silently. One live run under-reported all five domains it touched — the worst lost 91 rows of 93, and the one that looked healthiest still lost 49 — with no error anywhere and a wrong written conclusion on top.
The check is two comparisons, and conflating them produces false alarms:
<check level="1" compares="rawText vs parsed.rows.length"> Count the record-shaped lines in `rawText`, compare with `parsed.rows.length`. A gap here means **your regex has a blind spot** — the rows arrived and you dropped them. This is the silent, dangerous one. Fix the parser. </check> <check level="2" compares="the page's own headline count vs rawText"> Semrush prints its own total (`自然搜索排名: N`). If that exceeds what `rawText` even contains, the rows **never reached you**: these tables are virtual-scroll and only mount a fraction at a time, so a full pull needs the export, which costs quota. This is a known ceiling, not a bug — say so rather than "fixed the parser". </check>A live re-run shows both at once: three domains matched their headline exactly (14/14, 22/22, 5/5) while one read 91 against a claimed 430. The first three prove the parser; the fourth is level 2 and needs no fix. </caution> </workflow>
<workflow id="screen" when="before filling anything, always"> <statement> The qualifying test is real traffic (`>= 100` monthly visits), never DR, and it runs BEFORE the form does. The division of labour is fixed: **scripts collect evidence, the AI reads the evidence and judges, a human can re-check both.** The batch scripts emit measured raw values plus a parse status, a raw-text excerpt, a screenshot in `<out>.jsonl.evidence/`, and a `stopReason` — never a pass/fail verdict. `apply-traffic-screen` copies numbers and evidence paths into the table; the threshold is computed at query time by `targets-select`. </statement> <read><ref file="references/traffic-screen.md"/></read> <cmd><![CDATA[ node scripts/similarweb-batch.mjs --domains-file domains.txt --out sw.jsonl # rows: {totalVisits, parse, stopReason, rawExcerpt, evidence:{screenshot,raw}} node scripts/apply-traffic-screen.mjs --in sw.jsonl --source similarweb node scripts/targets-select.mjs --cohort open --min-traffic 100 ]]></cmd> <caution> **A null value is NOT low traffic — it is an unmeasured or failed capture, and treating it as a conclusion is banned.** `stopReason` tells you which: `stable`/`empty-state` mean the capture completed (and `empty-state` means the source itself printed a no-data sentence — whether that means "too small to measure" is the AI's call, made against the screenshot and raw text); `unstable`/`timeout`/`exception` mean *this check did not finish* — resume retries them, and applying such a row clears any stale same-source measurement instead of writing one. Rows without a number never fall into the "unqualified" bucket: `targets-select --min-traffic` lists them separately and `--unmeasured` queues them. </caution> <headline> Measuring a domain costs one query; filling its form costs two orders of magnitude more. One run filled every form across a 73-domain family and only then sampled five for traffic — every filled form was discarded. </headline> </workflow> <workflow id="submit" when="a route exists and the target passed the screen"> <read><ref file="references/submission-lanes.md"/></read> <inspect> Inspect every target independently. Never infer a form from a sibling site. `inspect-page.mjs` outputs a **full form census** — every form, every field (visible and hidden) with its semantic tuple and stable marker — plus a captureScene pair (piercing census + screenshot) in the evidence dir. Its `fillable` / `blocker` / `reason` / `selectedForm` fields are **heuristic suggestions** (see the `suggested` note in the output), not verdicts: the AI judges "can this page be filled, and what gates it" from the census and the screenshot, and may overrule the suggestion. The mechanical rule safe-fill still enforces: it only proceeds on one unambiguous qualifying form. A CAPTCHA page may be staged only when the owner explicitly accepts normal human completion; never bypass or solve it by an external CAPTCHA service. <cmd><![CDATA[ node scripts/inspect-page.mjs --session "inspect-comment-scan" --mode comment \ --url https://example.com/article --out .backlink/scan.json ]]></cmd> Modes are `comment`, `directory`, or `auto`. Evidence lands in `.backlink/scan.json.evidence/` (override with `--evidence-dir`). </inspect> <payload> Create a reviewed JSON payload with truthful values. For comment mode, `description` is the comment body. <cmd><![CDATA[ { "url": "https://owned.example/relevant-page", "name": "Real owner or product name", "email": "owner@example.com", "description": "A page-specific, useful comment or truthful listing description" } ]]></cmd> </payload> <fill> <cmd><![CDATA[ node scripts/safe-fill.mjs --session "fill-submission" \ --scan .backlink/scan.json --payload .backlink/payload.json ]]></cmd> It revalidates the URL, form identity, field semantics, login state, and CAPTCHA state, installs a submit guard, and never submits. The human reviews the rendered page and performs final submission. Only after the user explicitly authorizes one exact reviewed submission may the agent run `release-submit-guard.mjs` — and releasing the guard still does not click Submit. When the owner has explicitly accepted normal CAPTCHA completion, add `--allow-captcha`; this only permits guarded filling and leaves the CAPTCHA and final submission untouched. CAPTCHA routes are always handoff-only. Once the filled form is visibly ready for the owner, release only the guard with `release-submit-guard.mjs --human-handoff`, then stop. The agent must not solve the challenge or click Submit, even when a batch authorization already exists. </fill> <staged-queue> Lane B leaves forms on screen for the owner to finish. **One session name per staged site** — a session owns one tab, so reusing one session overwrites the previous staged form while the report still says N staged. `adapter-phpld.mjs` carries the reference implementation. </staged-queue> </workflow> <workflow id="analyze" when="the user has exported backlink data already"> <statement> Analyze referring-domain quality and topical relevance; suspicious networks, sitewide links, and toxic patterns; anchor and target-page diversity; follow/nofollow/UGC/sponsored distribution **when observed**; competitor gaps and prioritized next opportunities. </statement> <read> <ref file="references/link-quality-rubric.md"/> — scoring, toxicity, disavow. <ref file="references/analysis-templates.md"/> — report shapes. <ref file="references/outreach-templates.md"/> — frameworks; sending needs the user's explicit approval per message. </read> <hard-limit> These templates assume you already have the data. They do not fetch it. **A report built from templates alone, with no observed rows behind it, is fabrication.** Do not disavow links, contact site owners, or change production sites unless the user separately asks. Treat third-party authority and traffic estimates as directional and time-sensitive. </hard-limit> </workflow> <workflow id="harvest" when="the numbers are visible in a logged-in dashboard with no API"> <read> <ref file="references/harvest.md"/> before writing any scraping loop. It documents failures that produce **plausible, silently wrong output**: virtual scroll tables that are not `<table>` and drop rows without erroring, long URLs that make whole rows vanish, execution-channel timeouts that look like failure while the page loop is still running, and Chrome's intensive throttling stretching a four-second loop into twenty-five minutes. </read> <cmd><![CDATA[ sh scripts/harvest-collect.sh # wait for downloads to settle, then collect node scripts/harvest-merge.mjs # merge by field shape, refuse duplicate files ]]></cmd> <note> `scripts/harvest.browser.js` is the in-page collector. Its output arrives via a Blob download rather than a return value, because the execution channel truncates at roughly 1 KB.Reach for it only when you need the whole table as a file. For reading a report — "what does this page say", "is there data here at all" — use <ref file="scripts/ground-truth.mjs"/> instead: it is the collector <law-ref id="every-measurement-needs-two-witnesses"/> names, and harvest.browser.js has exactly one witness (the DOM), no manifest, and no screenshot to contradict it. When you do use harvest.browser.js, take the screenshots yourself. </note> </workflow>
<workflow id="verify" when="closing the loop on any placement"> <states>candidate → qualified → drafted → filled → submitted → public → indexed → rel_verified</states> <cmd><![CDATA[ # Track a submission node scripts/ledger.mjs upsert --file .backlink/ledger.json --url https://target.example/page node scripts/ledger.mjs transition --file .backlink/ledger.json \ --url https://target.example/page --state public \ --evidence "Observed the exact public anchor on 2026-07-30"Per-project progress: what have I submitted vs what's left?
node scripts/ledger.mjs stats --file .backlink/ledger.json node scripts/ledger.mjs remaining --file .backlink/ledger.json --min-traffic 100 node scripts/ledger.mjs remaining --file .backlink/ledger.json --cohort open --free-only
Select next batch — reads .backlink/ledger.json by default (relative to
cwd) and excludes submitted-or-later AND rejected domains with no flag needed
node scripts/targets-select.mjs --cohort open --min-traffic 100
]]></cmd>
<evidence-bar>
submitted, public, indexed, and rel_verified each require an evidence
note. Never promote a record from a filled form, a pending notice, or a
historical assumption. indexed must name the engine — indexed@google,
indexed@brave. An unqualified "indexed" is a claim about the whole web built
from one crawler's opinion.
</evidence-bar>
</workflow>
</workflows>
opencli comes first. Every browser action in this Skill runs through it, and
it carries the rules this Skill only summarises.
It also requires the OpenCLI binary and browser extension from
yan-labs/OpenCLI releases —
not the Chrome Web Store build. The store build defaults to foreground: it raises
a window and steals the tab the person is reading. That failure is silent — commands
still succeed, only the behaviour is wrong — so opencli doctor flags an extension
older than 1.0.32 explicitly. When it does, act on it rather than working around it.
</install>
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/yan-labs/yan-skills/backlink">View backlink on skillZs</a>