debug-plan-cache-misses
Find and debug Chalk query plan cache misses from engine logs. Use when a query is replanning on every request, an "ad-hoc query request missed the plan cache" warning shows up, latency spikes from repeated planning, or someone asks why two seemingly identical queries don't share a cached plan. Drives the `chalk logs` CLI search.
How do I install this agent skill?
npx skills add https://github.com/chalk-ai/chalk-ai-plugins --skill debug-plan-cache-missesIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill provides instructions for identifying and debugging Chalk query plan cache misses using the official chalk CLI to search and compare engine logs. It is a legitimate diagnostic utility for the Chalk platform with no malicious behavior detected.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Debug Chalk query plan cache misses
Chalk caches a compiled query plan the first time it plans a query, keyed by a
normalized BatchQuery. A cache miss forces the engine to replan on the
hot path, which is slow. Misses are expected the first time a query shape is
seen; they are a problem when a query that "looks identical" to a previous
one keeps missing — that means some field in the request is varying between
calls.
This skill: (1) finds misses in the logs, and (2) diffs the planned queries to identify exactly which field differs and forces the miss.
What the cache key actually is
The plan cache key wraps the entire normalized BatchQuery. The only
normalization applied is that planner_options are resolved to their
effective values — so leaving an option unset and setting it explicitly to its
default produce the same key and will not cause a spurious miss.
Everything else in the request is part of the key, including:
given_features— the exact set of input features the caller provides- target features (
optional_target_features/required_target_features) tags,required_resolver_tagsmax_staleness_overridesand staleness/recompute flagsquery_name_and_version- and the remaining
BatchQueryfields
So any difference in these forces a new plan. The single most common
real-world cause is a varying given_features set — the caller sometimes
supplies more or fewer input features (e.g. raw *_original inputs, or derived
inputs like email_username) than a previous call. A query that provides a
different set of inputs genuinely needs a different plan (different resolvers
must run), so it cannot reuse the cached one.
The two log lines that signal a miss
The engine emits these on the planning path:
-
Ad-hoc miss warning (
batch_online_query_service.py):Ad-hoc query request with query name/version '(<name>, <version>)' missed the plan cache. The query was: <BatchQuery ...> -
Plan computation — emitted on every miss (ad-hoc or named), and the most useful line for debugging (
local_plan_factory.py):Computed plan for [<target features>] given [<given features>] with planner options <NonNullPlannerOptions(...)>; query.query_name_and_version=(<name>, <version>)
There are also statsd counters if you want to trend rather than read lines:
chalk.engine.planner.python_plan_cache_miss vs
chalk.engine.planner.python_plan_cache_hit.
Finding misses with chalk logs
chalk logs searches the environment's engine logs. Point it at the right
environment/context first (chalk config / the usual project or --environment
selection), then query by log fields.
Search for the miss warnings over the last day:
chalk logs --query 'message:"missed the plan cache"' \
--start-time "24h ago" --end-time "now"
Search for plan computations for one query name (best for diffing):
chalk logs --query 'message:"Computed plan for" message:sv5_idplus_features_parallel' \
--start-time "24h ago" --end-time "now"
Trend miss volume over time in 10-minute buckets:
chalk logs --aggregate --window-period 10m \
--query 'message:"missed the plan cache"' --start-time "6h ago"
Tail misses live while reproducing:
chalk logs --follow --query 'message:"missed the plan cache"'
Query syntax
field:value pairs, space-separated (implicit AND). Quote any value that
contains a space or a . — e.g. message:"missed the plan cache",
query_name:"my.query". Useful fields:
message— substring match on the log message (use this for the two lines above)component— Chalk component, e.g.component:enginequery_name— filter to a specific named queryoperation_id— internal query id;correlation_id— caller-supplied query idtrace_id,pod_name,resource_group,deployment,appall_filter— match across multiple fields at once
Time flags: --start-time / --end-time accept '1h ago', 'now', or
ISO-8601 ('2024-01-01T00:00:00Z'). Add --tui for an interactive viewer.
Debugging workflow
-
Confirm misses are happening and how often:
chalk logs --query 'message:"missed the plan cache"' --start-time "24h ago"If this is empty, plan caching is working — look elsewhere for the latency.
-
Pull the
Computed plan forlines for the affected query name:chalk logs --query 'message:"Computed plan for" message:<query_name>' \ --start-time "24h ago" -
Take the two most recent computations and diff them field by field: compare the
given [...]lists first (most common culprit), then the target-feature lists, then theplanner optionsblock. Whatever differs is the cause of the miss. Because planner options are normalized before keying, a difference there is a real difference — not a set-vs-default artifact. -
Attribute the difference to the caller. A varying
givenset almost always means an upstream service is sending an inconsistent input schema (sometimes including raw/derived inputs, sometimes not). The fix is to make the caller send a stablegiven_featuresset on every request so the key is stable and subsequent identical queries hit the cache.
Notes
- Two byte-identical requests should hit the cache on the second call. If they don't, that's a distinct issue from a field mismatch (e.g. cache not being populated/shared) — call it out separately rather than blaming the request.
- Fixing a miss only helps if all callers converge on one request shape. If
some callers still send a different
givenset, those keep missing. - This skill reads logs and reasons about the request; it does not mutate any environment. Any remediation (changing what a caller sends) happens in the caller's code, not here.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/chalk-ai/chalk-ai-plugins/debug-plan-cache-misses">View debug-plan-cache-misses on skillZs</a>