browser-intent
Execute a natural-language browser intent via page-agent (browser_act) when the target is easier to describe than to select — degrades gracefully when page-agent or an OpenAI-compatible LLM provider isn't configured
How do I install this agent skill?
npx skills add https://github.com/ruvnet/ruflo --skill browser-intentIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill enables natural-language browser automation using a third-party agent. It includes built-in safety filters and secure credential proxying via a loopback proxy, but its autonomous nature and interaction with external web content create a surface for indirect prompt injection.
- Socketwarn
1 alert: gptAnomaly
- Snykwarn
Risk: MEDIUM · 1 issue
What does this agent skill do?
Browser Intent
Natural-language layer on top of the low-level browser_* selector tools. Where browser-extract and browser-form-fill compose selector-based primitives (browser_click, browser_fill, browser_snapshot), browser-intent lets the caller say what they want ("Click the login button", "Fill the search box with cats and submit") and delegates execution to page-agent — in-page injected JS that turns the DOM into text and drives an LLM tool-call loop against it.
When to use
- The target element is easier to describe in words than to select reliably (dynamic class names, ambiguous structure, A/B-tested markup).
- A one-shot interaction where writing out a selector chain isn't worth it.
- Prefer
browser_click/browser_fill/browser_snapshotdirectly when you already know the exact selector or ref (@e1) —browser_actadds LLM latency + cost that a direct selector call doesn't.
Steps
- Call
browser_actwith ataskstring, and optionallyurl(navigates first) andsession(default"default"):mcp__plugin_ruflo-core_ruflo__browser_act({ task: "Click the login button", url: "https://example.com/account", session: "my-session" }) - Read the response contract:
{ success: true, result, steps, history, contentFlagged, llmSource }— the intent executed.resultis the AIDefence-gated final text page-agent produced;historyis the full step trace (reflection + action + tool result per step);stepsishistory.length.{ success: true, degraded: true, reason, hint }— page-agent isn't installed, or no OpenAI-compatible LLM provider is configured. Never treatdegraded: trueas an error to retry — surface thehintand fall back to selector-basedbrowser_*tools instead.{ success: false, error, ... }— a real failure (browser open failed, injection failed, execution timed out, or page-agent's ownexecute()reportedsuccess:false).
- On
contentFlagged: true, the returnedresulthas already been redacted by AIDefence (PII or a prompt-injection/threat pattern was detected in the page-agent output) — do not attempt to recover the original text. - Prefer a recorded session (
browser-record) when the interaction matters enough to replay later;browser_actitself does not open an RVF container — it operates on whatever session id you pass (or"default").
Provider requirements (why this degrades so often)
page-agent calls its LLM directly from the browser page context via a plain OpenAI-compatible POST {baseURL}/chat/completions. That means:
- A bare
ANTHROPIC_API_KEYis not sufficient — Anthropic's native API is a different shape (/v1/messages). - Configure one of:
OPENROUTER_API_KEY(OpenRouter, OpenAI-compatible),OLLAMA_API_KEY(Ollama Cloud, OpenAI-compatible), orCLAUDE_FLOW_PAGE_AGENT_BASE_URL+CLAUDE_FLOW_PAGE_AGENT_API_KEYfor a custom OpenAI-compatible endpoint. - The real provider key never enters the page:
browser_actstarts a short-lived loopback HTTP proxy that holds the key server-side and injects the realAuthorizationheader itself. The page only ever sees a127.0.0.1URL and a placeholder key string.
Caveats
page-agentis anoptionalDependenciesentry (npm i page-agentif the doctor/degraded hint asks for it) — this plugin stays fully operational without it; you simply lose the natural-language layer and fall back to selector-based tools.- The npm bundle's demo auto-init tail (which would otherwise construct a second
PageAgentinstance against Alibaba's public test endpoint) is stripped before injection — you should never see traffic to apage-ag-testing-*host from this tool. - Every successful
browser_actcall best-effort records the intent + resulting trajectory into thebrowsermemory namespace (ADR-174 distillation loop). This is fire-and-forget — a memory-store failure never fails the tool call. timeoutMs(default 120000) bounds how longbrowser_actpolls forexecute()to settle; a slow multi-step intent may need a higher value.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/ruvnet/ruflo/browser-intent">View browser-intent on skillZs</a>