# MCP Inspect

> Contract intelligence for MCP servers. Track tools, schemas and descriptions, compare contracts, and watch for compatibility changes in the hosted console.

---

# MCP Inspect overview

> Snapshot your MCP server, detect breaking changes in schemas and descriptions, and gate pull requests with deterministic contract checks.

An MCP server is an API contract whose consumers are partly probabilistic. Two
kinds of change break things, and existing API tooling only sees one of them:

- **Structural.** A removed tool, an argument that became required, a narrowed
  enum. A correct caller stops working.
- **Semantic.** A rewritten description, a changed annotation, different server
  instructions. A request can still validate while a model picks a different tool.

MCP Inspect treats both as contract changes.

## The loop

```
develop → inspect → diff → merge → observe → learn → safely evolve
```

| Step         | What it is                                                                                                                                                                             |
| ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Snapshot** | Connect over stdio or streamable HTTP, discover everything the protocol describes, normalise it into a deterministic document with a SHA-256 fingerprint per entity.                   |
| **Diff**     | Compare two surfaces with an MCP-aware engine across [four severities](/docs/severities), with the old value, the new value, an explanation and a suggested migration on every change. |
| **Check**    | The same analysis in CI, with a configurable failure threshold and one pull request comment that stays current.                                                                        |
| **Observe**  | Optional telemetry, emitted by your own process, that records which argument _names_ were supplied — never their values.                                                               |

## What it is not

It is **not a proxy or a gateway**. Nothing here sits in your request path. The
inspector connects the way any MCP client would, on demand, and telemetry is
emitted asynchronously by your own process and is allowed to be dropped. A total
outage of the hosted service cannot affect your server.

## Start here

Start in the [hosted console](https://app.mcprobe.dev/signup). The
[getting started guide](/docs/getting-started) explains scheduled contract
captures and approved preview access to the local CLI. Public CLI and Action
releases are pending; the npm package named `mcp-inspect` is a different project.

The CLI works with no account and no network beyond the server being inspected once
installed. Initialize its configuration and capture a baseline:

```bash
mcp-inspect init --snapshot
# Make your change and rebuild the server, then:
mcp-inspect check --baseline-file .mcp-inspect/surface.json
```

[Getting started](/docs/getting-started) walks through the first comparison.
[MCP breaking changes](/docs/mcp-breaking-changes) explains which contract
changes can affect callers and where integration tests still matter.

---

# Skills and llms.txt

> Install MCP Inspect agent skills for setup, compatibility-check triage, deprecation, telemetry and surface review. Read the docs as Markdown through llms.txt.

Machine-readable surfaces, generated from the same source as the pages you are
reading. llms.txt and skill discovery are opt-in proposals; their presence does
not guarantee that a search engine or agent will use them.

| URL                                                                            | What it is                                                                                                                                                             |
| ------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| [`/llms.txt`](/llms.txt)                                                       | Every page, titled and linked, in the [llms.txt](https://llmstxt.org) format.                                                                                          |
| [`/.well-known/agent-skills/index.json`](/.well-known/agent-skills/index.json) | Skill discovery manifest, per the [Agent Skills Discovery RFC](https://github.com/cloudflare/agent-skills-discovery-rfc). Each entry carries a `sha256` of its source. |
| [`/sitemap.xml`](/sitemap.xml)                                                 | The usual.                                                                                                                                                             |

## The skills

Five, each a procedure rather than a summary of these pages. A skill earns its
place only if it tells an agent something it would otherwise get wrong.

| Skill                                                               | Use it for                                                     |
| ------------------------------------------------------------------- | -------------------------------------------------------------- |
| [`mcp-inspect-setup`](/skills/mcp-inspect-setup/SKILL.md)           | Adding MCP Inspect to a repository: config, baseline, CI.      |
| [`mcp-inspect-triage`](/skills/mcp-inspect-triage/SKILL.md)         | Reading a failing check and deciding what to do about it.      |
| [`mcp-inspect-deprecate`](/skills/mcp-inspect-deprecate/SKILL.md)   | Retiring a tool or an argument without breaking callers.       |
| [`mcp-inspect-instrument`](/skills/mcp-inspect-instrument/SKILL.md) | Adding usage telemetry, or exporting to OpenTelemetry instead. |
| [`mcp-surface-review`](/skills/mcp-surface-review/SKILL.md)         | Reviewing a surface before anyone depends on it.               |

## Using them

These skills describe the local and CI workflows available through approved
preview access. Install the preview CLI before asking an agent to run those
steps. Public CLI and Action package releases are pending. The plugin's GitHub
marketplace source is private; use the public skill files below instead of a
repository-based installation.

An agent that supports the discovery RFC finds them from the manifest. Otherwise
fetch a `SKILL.md` directly, or vendor it:

```bash
mkdir -p .claude/skills/mcp-inspect-setup
curl -sL https://mcprobe.dev/skills/mcp-inspect-setup/SKILL.md \
  -o .claude/skills/mcp-inspect-setup/SKILL.md
```

The manifest's `sha256` lets a client cache a skill and know when it has changed.

Every docs page has a `.md` companion advertised by an HTML `rel="alternate"`
link. For example, fetch [this page as Markdown](/docs/agents.md). Prefer a
specific page to [the full documentation](/llms-full.txt) when possible.

Every HTML page links to `/llms.txt` with `rel="describedby"` and advertises the
skill manifest with `rel="agent-skills"`. The latter is a discovery hint for
clients that recognize it; the RFC's well-known URL remains the entry point.

## Why the rule reference is generated

[Every rule id](/docs/rules) on this site is generated from the diff engine's own
registry. Rule ids are a public API — `.mcp-inspect.yml` silences by id, the JSON
report carries them, and the SARIF output emits a `helpUri` per rule that points
at a heading on that page.

A hand-written reference would drift from the engine the first time someone added
a rule in a hurry, and the symptom would be a link in a customer's CI annotation
that describes a different rule. So the page is derived, and a check in CI fails
when it is stale.

## What an agent should not conclude

The one thing worth stating to a machine as plainly as to a person: **usage data
proves the absence of observed usage, never the absence of all possible
consumers.** Nothing in this product says anything is safe to remove, and an
agent reading its output should not either. The strongest available phrasing is
"no recorded calls in the requested N-day window", with coverage limits. Client
software/version cohorts are not consumers. Argument absence is unknown without
complete collection and exact contract attribution; estimated volume cannot
increase confidence.

For automation, read the [JSON check report](/docs/cli), including on exit 1.
A `status: "skipped"` / `compared: false` report means no comparison ran even
though the process exits 0. Multi-server checks contain named reports in a
single JSON document; select `--server <name>` for the single-report shape.

---

# CLI

> MCP Inspect CLI reference for snapshot, diff, check, watch and push. Compare MCP contracts locally or in CI without an account.

```bash
mcp-inspect <command> [options]
```

## Commands

### init

```bash
mcp-inspect init --snapshot
mcp-inspect init --server search --snapshot
mcp-inspect init --workflow
```

Discover a server and capture its baseline in one step with `--snapshot`.
Use `--server` to select a named server when a client config contains several.
`--workflow` generates a repository-appropriate GitHub workflow without replacing
an existing workflow. Review its build command and configure server secrets in CI.

A capture establishes a baseline; it does not check compatibility. After changing
and rebuilding your server, compare locally with
`mcp-inspect check --baseline-file .mcp-inspect/surface.json`.

### snapshot

Capture the MCP surface.

```bash
mcp-inspect snapshot
mcp-inspect snapshot --stdio "node dist/server.js"
mcp-inspect snapshot --url https://api.acme.com/mcp --header "Authorization: Bearer $T"
mcp-inspect snapshot --out surface.json --pretty --push
```

Only `server/discover` and the `*/list` methods are ever called — all read-only
by specification. **The inspector never invokes a tool.** Pagination is
exhaustive: on hitting a page cap it fails rather than recording a truncated
surface, because a truncated list reads as a mass removal on the next diff.

An HTTP server that answers `401` starts the MCP authorization flow: resource
metadata, authorization server metadata, PKCE, a browser tab, a redirect back
to a listener on `127.0.0.1`. It runs when a terminal is attached (or with
`--oauth`); in CI a `401` is an error that names the alternatives, such as
`--header "Authorization: Bearer …"` or `--oauth-client-id`. **The token stays in
memory for one process and is never written to a snapshot, a config, or a file.**

Use `--out - --raw capture.json` to pipe the normalized JSON while retaining raw
protocol evidence in a file. Explicit `--out` or `--raw` paths require
`--server <name>` when capturing several configured servers.

### diff

```bash
mcp-inspect diff old.json new.json --format markdown --fail-on POTENTIALLY_BREAKING
```

Formats: `text` (default), `markdown`, `json`, `sarif`. The JSON report is a
supported integration surface and is versioned.

### check

Snapshot the working tree's server, resolve a baseline, diff, apply ignores, exit.

```bash
mcp-inspect check --baseline origin/main
mcp-inspect check --format markdown --output comment.md
```

Baseline resolution, in order: `--baseline-file`, then a git ref
(`git show <ref>:<path>`), then the hosted API, then nothing. A missing baseline is explicitly reported as **comparison
skipped**, with exit 0 so the setup PR can merge. It is not evidence of a passing
comparison.

For automation, `--format json --output report.json` writes a report even on
exit 1. A missing baseline reports `status: "skipped"` and `compared: false`.
Checks covering several servers emit one JSON document with
`{ formatVersion: 1, reports: [{ server, report }] }`; inspect each named outcome.
SARIF contains one run per server with `runs[].properties.server` and retains
skipped-comparison metadata. Selecting `--server <name>` gives the existing
single-report shape. The process exits with the worst code across the checks;
a skipped entry still needs a baseline before its compatibility can be assessed.

Under the summary, one advisory line compares the declared version with what
the diff found — `breaking changes suggest a major bump to 3.0.0` — when both
surfaces declare semver. It never changes the verdict.

### watch

```bash
mcp-inspect watch --path src
```

The develop loop: snapshot, diff against the baseline, print, then re-run on
every save under the watched paths. What it prints is what `check` prints in
CI. Ctrl-C stops it.

### doctor

Config, server reachability, what it advertises, baseline, hosted status. The
command to run when `check` did something unexpected.

### keys

```bash
mcp-inspect keys list
mcp-inspect keys create --kind api --name github-actions
mcp-inspect keys rotate key_01k4…
mcp-inspect keys revoke key_01k4…
```

The plaintext goes to **stdout alone** and everything else to stderr, so
`mcp-inspect keys create … > key.txt` captures the key and nothing else.

Rotation mints a replacement and leaves the previous key working for a week. A
rotation that breaks every build the instant it happens is one nobody performs
twice. `revoke` stops a key immediately.

### push and usage

`push` records a snapshot as a deployment. `usage` prints observed usage for a
tool, including argument presence and the observation window. Both need an
account; `push` exits 0 even when the upload fails, because losing a snapshot
must never break a build.

## Exit codes

| Code | Meaning                                                              |
| ---- | -------------------------------------------------------------------- |
| `0`  | Passed, or there was nothing to compare.                             |
| `1`  | The change exceeded the fail threshold.                              |
| `2`  | The command could not run: server unreachable, bad config, bad JSON. |

`2` is never `0`. A snapshot that could not be taken is not "no changes".

| Command    | `0`                                       | `1`      | `2`                                        |
| ---------- | ----------------------------------------- | -------- | ------------------------------------------ |
| `snapshot` | captured                                  | —        | unreachable, bad config                    |
| `diff`     | at or below `fail_on`                     | above it | missing, corrupt, or wrong `formatVersion` |
| `check`    | at or below `fail_on`, **or no baseline** | above it | unreachable, bad config                    |
| `push`     | uploaded, **or the upload failed**        | —        | the snapshot could not be read             |
| `keys`     | done                                      | —        | no token, or a missing argument            |
| `doctor`   | every check passed or warned              | —        | a check failed                             |

## Configuration

`.mcp-inspect.yml`, resolved from the working directory upward.
`.mcp-inspect.json` is accepted identically.

```yaml
server:
  command: node dist/server.js # or: url: https://api.acme.com/mcp
  env:
    DATABASE_URL: ${DATABASE_URL} # ${VAR} interpolates; an unset one is an error
  # headers:
  #   Authorization: Bearer ${MCP_TOKEN}
  timeout_ms: 30000

baseline:
  source: git # git | file | api
  ref: origin/main
  path: .mcp-inspect/surface.json

fail_on: BREAKING

ignore:
  - rule: tool.icons.changed
severity_overrides:
  tool.annotations.safetyWeakened: BREAKING

project: acme/search-mcp # hosted; everything above works without it
```

Several servers in one repository: replace `server:` with a `servers:` map,
one named entry each, and every command runs against each in turn with its own
baseline at `.mcp-inspect/<name>.json`; `--server <name>` picks one.

```yaml
servers:
  api:
    command: node dist/api.js
  internal:
    url: https://internal.example/mcp
    oauth:
      client_id: ${INTERNAL_OAUTH_CLIENT} # optional; `oauth: false` disables the browser flow
```

An unset `${VAR}` is a hard error rather than an empty string. An empty auth
header produces a smaller surface, which the next diff reports as a mass
removal — a loud failure is the only safe behaviour.

A top-level key within two edits of a real one (`sever:`, `failon:`) is an error
naming what you probably meant. A genuinely unrelated key is ignored, so a config
written for a newer version still loads.

---

# Getting started

> Capture your contract, compare a change, then protect the next pull request. Start in the hosted console, or use an approved preview CLI build.

## Start in the hosted console

MCP Inspect is a privately operated hosted product. Open
[app.mcprobe.dev/signup](https://app.mcprobe.dev/signup) and create an account
with an email address and a password of at least 12 characters. An organization
and a first project are created for you. GitHub sign-in is not offered until
its OAuth integration is configured.

After creating an account, paste the server's final public endpoint URL into
**Capture your first contract** and choose **Capture contract**. No CLI, API key
or GitHub connection is needed for hosted HTTP capture. Expand authorization only
if your server requires a header. If authorization fails, add a current header
and retry in the same screen.

Capture progress stays visible until a recorded snapshot opens your server. The
first snapshot is a baseline, not a compatibility verdict. Browse the captured
entities and input paths, then choose **Capture again** to compare the next
version. New watches check hourly and alert on breaking changes; adjust the
interval and threshold in **Settings → Watched servers**. Discovery does not
invoke tools and never sits in the server's request path.

Use the server's **Changes** view to review snapshots, compatibility findings
and migration hints. A baseline establishes history; the next distinct contract
provides a comparison. A failed capture is shown as a failure, not an unchanged
contract. Watched servers cannot target private or loopback addresses.

## Local and CI preview access

The CLI, instrumentation and standalone Action are not publicly released yet.
Use only builds supplied through approved preview access. The source repository
is private; cloning it is not a public installation path. The npm package named
`mcp-inspect` belongs to another project and must not be used for this product.

Once the preview CLI is installed, the local workflow below works without an
account. Hosted usage queries, nightly rollups, account email delivery and paid
checkout remain pending. Self-host deployments are not supported.

## 1. Capture your baseline

From your server repository, with its normal build and environment ready:

```bash
mcp-inspect init --snapshot
```

This imports an existing MCP configuration, connects to the server and saves
`.mcp-inspect/surface.json`. It discovers the advertised surface without calling
any tools. A successful capture is your **baseline**, not a compatibility verdict.

If several servers are configured, select the one you maintain with
`--server <name>`. If none is found, fill in the generated config's `server`
command or URL, then run `mcp-inspect snapshot`. Use
`mcp-inspect doctor` if the connection fails; it is a troubleshooting step,
not a required extra capture.

Already have a configuration? Run `mcp-inspect snapshot` before editing the
server. Already have two snapshots? Go straight to:

```bash
mcp-inspect diff before.json after.json
```

## 2. Compare your change locally

Make your intended server change and rebuild it if needed. Keep the initial
snapshot unchanged, then run:

```bash
mcp-inspect check --baseline-file .mcp-inspect/surface.json
```

The check captures the current server and compares it with your saved baseline.
It does not overwrite that baseline. Read the changed entity, compatibility
severity and migration hint together.

### Example report

This is an illustrative report; your own result depends on your server changes.

```text
BREAKING (1)
  ! search
    `limit` changed from optional to required.
    → Accept calls that omit it for at least one release, applying the
      previous default, then require it.

POTENTIALLY_BREAKING (1)
  ~ create_issue
    Description changed. Agents may alter tool-selection behavior.
    → Review the wording and re-run your tool-selection evaluations.
```

A breaking finding exits 1 by default. A potentially breaking finding remains
visible but only blocks when you choose `fail_on: POTENTIALLY_BREAKING`.
Prose classification is a deterministic heuristic, not a measurement of model
behavior. [Understand severities](/docs/severities).

An unchanged contract produces no changes. That is a completed comparison.
**“Comparison skipped — baseline needed” means no comparison happened**, even
though the command exits 0 so a setup PR can merge.

## 3. Protect the next pull request

The standalone GitHub Action release is pending. The generated workflow currently
references an unreleased Action location. Do not enable it as a working check yet.
See [GitHub Action details](/docs/github-action) for preview access requirements.

After a verified Action release, generate its workflow:

```bash
mcp-inspect init --workflow
```

Review the generated `.github/workflows/mcp-compatibility.yml`, including your
build command and any environment secrets the server needs. Commit it together
with `.mcp-inspect.yml` and the baseline snapshot. Merge the baseline into your
PR base branch before expecting CI to compare against it.

The first setup PR may report a skipped comparison. The next PR should show an
actual comparison; the Action exposes `compared=true` when one ran.

After reviewing an intentional contract change, run `mcp-inspect snapshot`
and commit the updated snapshot with the change. CI reads its baseline from the
base branch, so it can still check that PR while the updated snapshot becomes
the baseline after merge. [GitHub Action details](/docs/github-action).

Do not hand-edit snapshots: their content and fingerprints are verified on read.

## 4. Add shared history when you need it

The local check is fully usable without signup. To keep deployment history,
open the hosted console and follow **Add your first server**: select a project,
create an API key and run the supplied upload command. Setup is complete when
the console receives the snapshot. A successful CLI exit alone does not confirm
upload, because hosted failures are warnings.

In `.mcp-inspect.yml`, set `endpoint: https://api.mcprobe.dev` explicitly when
using a source revision with the previous hosted default. The console is at
[app.mcprobe.dev](https://app.mcprobe.dev). Usage queries and nightly rollups
are pending the deployment's analytics read credential; account email delivery,
GitHub sign-in and paid checkout are not configured yet.

[Telemetry](/docs/telemetry) is a separate, optional step. You can capture,
compare, gate pull requests and keep shared history without installing the SDK.

---

# GitHub Action

> Check MCP server compatibility in GitHub pull requests. Configure baselines, failure thresholds, migration hints and PR comments.

## Preview access required

The standalone Action has not been released at `mcp-inspect/action@v1`.
The implementation repository is private. Use only a reviewed CLI build or
Action supplied through approved preview access; there is no public source
checkout or installable Action release yet. The reference below describes the
implemented behavior, not a working public distribution.

A CI check captures your server in the runner and compares it with a baseline
committed to the pull request's base branch. The hosted service is optional for
that local comparison. [Getting started](/docs/getting-started) covers the
hosted console and preview prerequisites.

## Standalone Action reference (release pending)

The following describes the implemented Action, whose release is still pending.
Its `uses` location is a placeholder and cannot currently be copied into CI.

```yaml
name: MCP Compatibility
on: pull_request

permissions:
  contents: read
  pull-requests: write

jobs:
  mcp-check:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with: { fetch-depth: 0 } # the git baseline needs history
      - uses: actions/setup-node@v4
        with: { node-version: 22 }
      - run: npm ci && npm run build
      - uses: mcp-inspect/action@v1
        with:
          fail-on: BREAKING
```

## Inputs

| Input               | Default              | Meaning                                        |
| ------------------- | -------------------- | ---------------------------------------------- |
| `config`            | `.mcp-inspect.yml`   | Config path.                                   |
| `baseline`          | the PR's base branch | Git ref to compare against.                    |
| `baseline-file`     | —                    | Explicit snapshot file; wins over `baseline`.  |
| `fail-on`           | `BREAKING`           | Threshold. `never` makes the Action advisory.  |
| `comment`           | `true`               | Post or update a pull request comment.         |
| `annotate`          | `true`               | Emit `::error` / `::warning` annotations.      |
| `api-key`           | —                    | Hosted key. Enables usage evidence and `push`. |
| `push`              | `false`              | Record this run as a deployment.               |
| `working-directory` | `.`                  | For monorepos.                                 |

Outputs: `compared`, `passed`, `breaking`, `potentially-breaking`, `summary`, `report`,
`fingerprint`.

## The comment

```markdown
### MCP Compatibility

**1 breaking, 2 potentially breaking, 1 non-breaking**

🚫 **`search`** `.limit`
`limit` changed from optional to required.

> Accept calls that omit it for at least one release, then require it.

⚠ **`create_issue`** `.description`
Description changed substantially. Agents may alter tool-selection behavior.

➕ **`archive_issue`**
New tool.
```

One comment, found by a hidden marker, updated in place. A pull request with
twenty pushes has one current comment rather than twenty stale ones.

## Usage evidence

With an `api-key`, observed usage is folded in where it changes the decision:

> 🚫 **`legacy_search`** — removed
> `legacy_search` accounted for 0.03% of calls in the last 90 days, but **4
> distinct clients** called it, most recently 2 days ago. Removing it will break
> them.

That sentence is the product. It is also why the Action takes a key but does not
require one: the compatibility check is useful alone.

## Failure modes

The job summary is written **before** the comment is attempted, because a fork
pull request has a read-only token and the summary is the only output that always
survives. Failing to reach the hosted service is a warning, never a failure. The
Action exits non-zero only when the threshold was exceeded.

## Know when protection is active

With no baseline, the job exits 0 but reports **Comparison skipped — baseline
needed** and `compared=false`. `passed=true` alone means the job did not block.
A real passing comparison requires both `compared=true` and `passed=true`.

Merge the initial snapshot into the PR base branch, then run a check on a later
PR. After reviewing an intentional change, regenerate and commit the snapshot
with that change to keep the baseline current.

---

# Detect MCP breaking changes before merge

> Compare MCP tool schemas, descriptions and capabilities with a deterministic contract check in your pull request.

An MCP server exposes more than JSON Schema. Tools, resources, resource templates,
prompts, capabilities and server instructions all form its contract. An agent
uses descriptions and annotations to decide what to call, so a change can alter
behavior even when an existing request still validates.

## Which changes can break an MCP caller?

| Change                                      | Why it matters                                    | MCP Inspect severity |
| ------------------------------------------- | ------------------------------------------------- | -------------------- |
| Remove a tool                               | A caller can still request its old name.          | BREAKING             |
| Make an optional input required             | Previously valid calls can omit it.               | BREAKING             |
| Narrow an input enum                        | A caller can send a previously accepted value.    | BREAKING             |
| Rewrite a tool description substantially    | An agent may change its tool selection.           | POTENTIALLY_BREAKING |
| Change an annotation such as `readOnlyHint` | The host or agent may treat the tool differently. | POTENTIALLY_BREAKING |
| Add a tool                                  | Existing tool contracts remain available.         | NON_BREAKING         |

These are examples, not an exhaustive checklist. The [rule reference](/docs/rules)
documents every emitted rule id and its migration hint. Input and output schemas
have different compatibility directions. An unmodeled schema keyword is reported
as breaking rather than silently ignored.

Description classification is a deterministic heuristic. It identifies changes
worth reviewing; it does not predict a particular model’s behavior. Re-run your
tool-selection evaluations when the meaning changes.

## Compare before and after

Run this in your MCP server repository before editing the contract:

```bash
mcp-inspect init --snapshot
```

This writes configuration and captures `.mcp-inspect/surface.json`. The inspector
discovers the advertised contract over stdio or streamable HTTP without invoking
tools. A baseline capture by itself is not a compatibility verdict.

After making your change and rebuilding the server, compare against that baseline:

```bash
mcp-inspect check --baseline-file .mcp-inspect/surface.json
```

The check captures the current surface without overwriting your baseline. By
default a breaking change exits 1; a connection or capture failure exits 2.
Set `fail_on: POTENTIALLY_BREAKING` in `.mcp-inspect.yml` if substantial prose and
annotation changes should also block CI. [Severity details](/docs/severities).

For two existing captures, use `mcp-inspect diff before.json after.json`.
The verdict needs no hosted account or LLM API key.

## Gate a pull request

Run `mcp-inspect init --workflow` and review the generated workflow, server
build command and secrets. Commit the configuration and baseline to the base
branch. The [GitHub Action](/docs/github-action) compares against that branch and
updates one PR comment with findings and migration hints.

A first setup PR can report a skipped comparison. That is not a clean comparison;
the Action’s `compared=true` output tells you whether analysis actually ran.
Once an intentional change is reviewed, capture and commit the updated baseline
with it. The PR still compares against the previous baseline in the base branch.

## What contract checks do not measure

A snapshot captures the surface a particular caller can discover. It does not
execute handlers, test business logic or prove an agent completed its task.
Keep integration tests and tool-selection evaluations alongside the contract check.

Optional [telemetry](/docs/telemetry) adds observed tool and argument-path usage.
It records no argument values. Observation windows and collection limits matter:
absence of observed usage is not proof that removing a tool will affect nobody.

If you consume a third-party server, the hosted console can schedule captures
under watched-server settings and alert when its contract changes. Local
`mcp-inspect watch` instead rechecks your own server on source-file changes;
these are separate workflows.

---

# Privacy

> MCP Inspect telemetry records argument paths, never argument values. Learn what is collected, what is rejected and what usage evidence cannot prove.

Privacy here is a product feature, not a compliance checkbox. The people using
this are instrumenting servers that carry other people's data, and "what exactly
does this send?" is the first question any of them asks. The answer has to be
short, checkable, and enforced.

| Recorded                                       | Never recorded                       |
| ---------------------------------------------- | ------------------------------------ |
| That `search` was called                       | What was searched for                |
| That `query` and `limit` were present          | The value of either                  |
| That it took 42 ms and succeeded               | What it returned                     |
| That the caller called itself `claude-desktop` | Who the end user was                 |
| An anonymous per-process session id            | Any stable user or device identifier |

## How it is enforced

**One chokepoint.** `summarizeArguments()` is the only function ever handed raw
arguments, and it returns field paths. It walks into nested objects —
`filters.owner` is a field a schema change can break — but not into array
elements, because `items[3].owner` is data.

**The server refuses values.** Ingest returns `400` for a payload containing
`arguments`, `input`, `output`, `result` or `content`. The SDK runs on your
machine, so trusting it would make the guarantee unverifiable.

**Paths must look like paths.** Each entry must match a property-path shape or it
is dropped, so a sentence or an address cannot arrive labelled as a "path". That
is a shape check, not a secret detector, and is not presented as one — a token
shaped exactly like an identifier is indistinguishable from a property name. The
real guarantee is the chokepoint above, which structurally cannot emit a value.

**Nothing is inferred.** No IP geolocation, no fingerprinting, no cross-server
correlation of "the same user".

## Snapshots

A surface snapshot is your server's public interface — the tool names,
descriptions and schemas every client already receives. It is stored verbatim.

An MCP surface may legitimately vary by authorization, so a snapshot taken with a
privileged credential can describe tools most callers never see. Snapshots record
the `cacheScope` of each list and warn when one was `private`, and the console
labels such a surface as authorization-scoped. **Credentials used to take a
snapshot are never stored** — they are read from the environment at capture time,
and the snapshot records only that a header was set, not its value.

## Retention

| Data              | Retention                                  |
| ----------------- | ------------------------------------------ |
| Surface snapshots | Life of the project. They are the product. |
| Invocation stream | ~90 days, then rolled up                   |
| Daily rollups     | Your plan's window. Counts only.           |
| Client records    | Name, version, first and last seen         |

Deleting a project deletes its rows and its stored documents.

---

# Rule reference

> Every change MCP Inspect can report, with its default severity and what to do about it. Generated from the diff engine, so it cannot drift.

Each rule has a stable id. `.mcp-inspect.yml` silences by id, the JSON and
SARIF reports carry it, and every heading on this page is a link target — so
an annotation in your CI can point straight at the rule that produced it.

Renaming a rule would break customers' configs, so a classification change
adds a new id and keeps the old one emitting until a major version.

```yaml
# .mcp-inspect.yml
ignore:
  - rule: tool.icons.changed
  - rule: tool.description.changed
    entity: legacy_*
    reason: "Legacy tools are being rewritten; churn is expected."

severity_overrides:
  tool.annotations.safetyWeakened: BREAKING
```

## Index

**Server and protocol** — [`server.capability.added`](#server.capability.added) · [`server.capability.changed`](#server.capability.changed) · [`server.capability.removed`](#server.capability.removed) · [`server.extension.added`](#server.extension.added) · [`server.extension.changed`](#server.extension.changed) · [`server.extension.removed`](#server.extension.removed) · [`server.identity.changed`](#server.identity.changed) · [`server.instructions.changed`](#server.instructions.changed) · [`server.protocolVersion.changed`](#server.protocolVersion.changed) · [`server.surface.scopeNarrowed`](#server.surface.scopeNarrowed) · [`server.version.added`](#server.version.added) · [`server.version.dropped`](#server.version.dropped)

**Tool input schemas** — [`tool.input.added`](#tool.input.added) · [`tool.input.constraint.changed`](#tool.input.constraint.changed) · [`tool.input.constraint.loosened`](#tool.input.constraint.loosened) · [`tool.input.constraint.tightened`](#tool.input.constraint.tightened) · [`tool.input.default.changed`](#tool.input.default.changed) · [`tool.input.deprecated.added`](#tool.input.deprecated.added) · [`tool.input.deprecated.removed`](#tool.input.deprecated.removed) · [`tool.input.description.changed`](#tool.input.description.changed) · [`tool.input.enum.narrowed`](#tool.input.enum.narrowed) · [`tool.input.enum.widened`](#tool.input.enum.widened) · [`tool.input.header.changed`](#tool.input.header.changed) · [`tool.input.property.added`](#tool.input.property.added) · [`tool.input.property.addedRequired`](#tool.input.property.addedRequired) · [`tool.input.property.removed`](#tool.input.property.removed) · [`tool.input.removed`](#tool.input.removed) · [`tool.input.required.added`](#tool.input.required.added) · [`tool.input.required.removed`](#tool.input.required.removed) · [`tool.input.schema.changed`](#tool.input.schema.changed) · [`tool.input.type.changed`](#tool.input.type.changed) · [`tool.input.type.narrowed`](#tool.input.type.narrowed) · [`tool.input.type.widened`](#tool.input.type.widened)

**Tool output schemas** — [`tool.output.added`](#tool.output.added) · [`tool.output.constraint.changed`](#tool.output.constraint.changed) · [`tool.output.constraint.loosened`](#tool.output.constraint.loosened) · [`tool.output.constraint.tightened`](#tool.output.constraint.tightened) · [`tool.output.default.changed`](#tool.output.default.changed) · [`tool.output.deprecated.added`](#tool.output.deprecated.added) · [`tool.output.deprecated.removed`](#tool.output.deprecated.removed) · [`tool.output.description.changed`](#tool.output.description.changed) · [`tool.output.enum.narrowed`](#tool.output.enum.narrowed) · [`tool.output.enum.widened`](#tool.output.enum.widened) · [`tool.output.header.changed`](#tool.output.header.changed) · [`tool.output.property.added`](#tool.output.property.added) · [`tool.output.property.addedRequired`](#tool.output.property.addedRequired) · [`tool.output.property.removed`](#tool.output.property.removed) · [`tool.output.removed`](#tool.output.removed) · [`tool.output.required.added`](#tool.output.required.added) · [`tool.output.required.removed`](#tool.output.required.removed) · [`tool.output.schema.changed`](#tool.output.schema.changed) · [`tool.output.type.changed`](#tool.output.type.changed) · [`tool.output.type.narrowed`](#tool.output.type.narrowed) · [`tool.output.type.widened`](#tool.output.type.widened)

**Tools** — [`tool.added`](#tool.added) · [`tool.annotations.added`](#tool.annotations.added) · [`tool.annotations.changed`](#tool.annotations.changed) · [`tool.annotations.removed`](#tool.annotations.removed) · [`tool.annotations.safetyWeakened`](#tool.annotations.safetyWeakened) · [`tool.description.changed`](#tool.description.changed) · [`tool.icons.changed`](#tool.icons.changed) · [`tool.meta.changed`](#tool.meta.changed) · [`tool.removed`](#tool.removed) · [`tool.renamed`](#tool.renamed) · [`tool.title.changed`](#tool.title.changed) · [`tool.ui.resourceUri.added`](#tool.ui.resourceUri.added) · [`tool.ui.resourceUri.changed`](#tool.ui.resourceUri.changed) · [`tool.ui.resourceUri.removed`](#tool.ui.resourceUri.removed) · [`tool.ui.visibility.narrowed`](#tool.ui.visibility.narrowed) · [`tool.ui.visibility.widened`](#tool.ui.visibility.widened)

**Resource templates** — [`resourceTemplate.added`](#resourceTemplate.added) · [`resourceTemplate.annotations.changed`](#resourceTemplate.annotations.changed) · [`resourceTemplate.description.changed`](#resourceTemplate.description.changed) · [`resourceTemplate.meta.changed`](#resourceTemplate.meta.changed) · [`resourceTemplate.mimeType.changed`](#resourceTemplate.mimeType.changed) · [`resourceTemplate.name.changed`](#resourceTemplate.name.changed) · [`resourceTemplate.removed`](#resourceTemplate.removed) · [`resourceTemplate.title.changed`](#resourceTemplate.title.changed)

**Resources** — [`resource.added`](#resource.added) · [`resource.annotations.changed`](#resource.annotations.changed) · [`resource.description.changed`](#resource.description.changed) · [`resource.meta.changed`](#resource.meta.changed) · [`resource.mimeType.changed`](#resource.mimeType.changed) · [`resource.name.changed`](#resource.name.changed) · [`resource.removed`](#resource.removed) · [`resource.title.changed`](#resource.title.changed) · [`resource.ui.csp.narrowed`](#resource.ui.csp.narrowed) · [`resource.ui.csp.widened`](#resource.ui.csp.widened) · [`resource.ui.domain.changed`](#resource.ui.domain.changed) · [`resource.ui.permissions.added`](#resource.ui.permissions.added) · [`resource.ui.permissions.removed`](#resource.ui.permissions.removed) · [`resource.ui.widgetDescription.changed`](#resource.ui.widgetDescription.changed)

**Prompts** — [`prompt.added`](#prompt.added) · [`prompt.argument.added`](#prompt.argument.added) · [`prompt.argument.addedRequired`](#prompt.argument.addedRequired) · [`prompt.argument.description.changed`](#prompt.argument.description.changed) · [`prompt.argument.nowOptional`](#prompt.argument.nowOptional) · [`prompt.argument.nowRequired`](#prompt.argument.nowRequired) · [`prompt.argument.removed`](#prompt.argument.removed) · [`prompt.description.changed`](#prompt.description.changed) · [`prompt.removed`](#prompt.removed) · [`prompt.title.changed`](#prompt.title.changed)

## Server and protocol

### server.capability.added

**Default severity:** `NON_BREAKING`

A capability is now advertised.

### server.capability.changed

**Default severity:** `POTENTIALLY_BREAKING`

e.g. `listChanged` or `subscribe` flipped.

> [!NOTE]
> Check that clients relying on the previous flag still behave correctly.

### server.capability.removed

**Default severity:** `BREAKING`

`capabilities.tools` / `resources` / `prompts` / … withdrawn.

> [!NOTE]
> Clients gate whole feature sets on this capability. Restore it, or announce the removal a release ahead.

### server.extension.added

**Default severity:** `NON_BREAKING`

A new extension is advertised.

### server.extension.changed

**Default severity:** `POTENTIALLY_BREAKING`

An extension's settings object changed.

> [!NOTE]
> Extension settings are part of the negotiated contract; version the extension id if this is not backwards compatible.

### server.extension.removed

**Default severity:** `BREAKING`

An advertised extension is gone; clients negotiated for it.

> [!NOTE]
> A client that negotiated this extension will lose the behaviour it asked for. Restore it, or fall back to core behaviour explicitly.

### server.identity.changed

**Default severity:** `DOCUMENTATION`

`serverInfo.name` / `version` / `title` changed.

### server.instructions.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

The `instructions` string changed. It is in the model's context.

> [!NOTE]
> Instructions are in the model's context on every conversation. Re-run any evaluation that depends on how this server is used.

### server.protocolVersion.changed

**Default severity:** `POTENTIALLY_BREAKING`

The negotiated revision changed. Different rules now apply.

> [!NOTE]
> Confirm clients on the previous revision can still reach this server, and keep the old version in supportedVersions until they have moved.

### server.surface.scopeNarrowed

**Default severity:** `POTENTIALLY_BREAKING`

A list result went `cacheScope: public` → `private`: the surface now varies by authorization, so this snapshot describes one caller's view, not the server's.

> [!NOTE]
> This surface now varies by authorization, so a snapshot describes one caller's view. Capture baselines with a consistent credential or this diff will report phantom changes.

### server.version.added

**Default severity:** `NON_BREAKING`

A new protocol version is supported.

### server.version.dropped

**Default severity:** `BREAKING`

A previously supported protocol version is gone.

> [!NOTE]
> Restore the version, or confirm no client still negotiates it.

## Tool input schemas

### tool.input.added

**Default severity:** `NON_BREAKING`

An absent `outputSchema` appeared.

### tool.input.constraint.changed

**Default severity:** `BREAKING`

`pattern`, `format`, `const`, `multipleOf`, `$ref` or a `oneOf` branch set changed: not orderable.

> [!NOTE]
> The two constraints are not orderable, so compatibility cannot be established structurally. Verify by hand.

### tool.input.constraint.loosened

**Default severity:** `NON_BREAKING`

A bound moved outward.

> [!NOTE]
> Consumers validating results against the old bound will start rejecting them.

### tool.input.constraint.tightened

**Default severity:** `BREAKING`

A bound moved inward (`minLength`, `maximum`, `additionalProperties: false`, `uniqueItems`, `allOf` branch added, `anyOf` branch removed).

> [!NOTE]
> Input that was valid is now rejected. Loosen the bound, or stage the change.

### tool.input.default.changed

**Default severity:** `POTENTIALLY_BREAKING`

A `default` was added, removed or changed. The schema still validates; the behavior does not match.

> [!NOTE]
> The schema still validates, but calls that omit this argument now behave differently. Check telemetry for how often it is omitted before shipping.

### tool.input.deprecated.added

**Default severity:** `NON_BREAKING`

`deprecated: true` appeared on a property.

> [!NOTE]
> Check usage before removing it in a later release.

### tool.input.deprecated.removed

**Default severity:** `NON_BREAKING`

A property is no longer marked deprecated.

### tool.input.description.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

A property `description` **or** `title` changed. Models read these when filling arguments.

> [!NOTE]
> Models read property descriptions when filling arguments. A changed one can change what gets passed.

### tool.input.enum.narrowed

**Default severity:** `BREAKING`

Values removed from an `enum`.

> [!NOTE]
> Callers passing a removed value will now fail. Accept the old values and map them for one release.

### tool.input.enum.widened

**Default severity:** `NON_BREAKING`

Values added to an `enum`.

> [!NOTE]
> Consumers switching exhaustively on this enum will not handle the new values.

### tool.input.header.changed

**Default severity:** `BREAKING`

`x-mcp-header` added, removed or changed: it alters how the call is transmitted over HTTP.

> [!NOTE]
> x-mcp-header changes how the call is transmitted over HTTP, so intermediaries routing on the mirrored header will stop matching.

### tool.input.property.added

**Default severity:** `NON_BREAKING`

A new optional property.

### tool.input.property.addedRequired

**Default severity:** `BREAKING`

A new property that is immediately required.

> [!NOTE]
> Every existing call omits this property and will now fail validation. Add it as optional with a default, then require it a release later.

### tool.input.property.removed

**Default severity:** `BREAKING`

A property is gone.

> [!NOTE]
> Accept and ignore the property for one release so existing callers keep working, then remove it.

### tool.input.removed

**Default severity:** `BREAKING`

A declared `outputSchema` disappeared.

> [!NOTE]
> A client that generated types from this schema loses them. Restore the schema, or release the removal as a major version.

### tool.input.required.added

**Default severity:** `BREAKING`

Optional → required.

> [!NOTE]
> Accept calls that omit it for at least one release, applying the previous default, then require it.

### tool.input.required.removed

**Default severity:** `NON_BREAKING`

Required → optional.

> [!NOTE]
> Consumers may assume this field is always present in a result. Keep emitting it, or release the change as a major version.

### tool.input.schema.changed

**Default severity:** `BREAKING`

**The completeness net.** A difference the structured walk does not model. See §9.

> [!NOTE]
> This schema keyword is not modelled by the diff engine, so the change is reported as breaking until a human confirms otherwise.

### tool.input.type.changed

**Default severity:** `BREAKING`

Type sets are not comparable.

> [!NOTE]
> The type sets are not comparable, so no existing caller is guaranteed to still work.

### tool.input.type.narrowed

**Default severity:** `BREAKING`

The `type` set shrank.

> [!NOTE]
> Keep accepting the removed types and coerce them for one release.

### tool.input.type.widened

**Default severity:** `NON_BREAKING`

The `type` set grew.

> [!NOTE]
> Consumers typed against the narrower schema will not handle the new types.

## Tool output schemas

### tool.output.added

**Default severity:** `NON_BREAKING`

An absent `outputSchema` appeared.

### tool.output.constraint.changed

**Default severity:** `BREAKING`

`pattern`, `format`, `const`, `multipleOf`, `$ref` or a `oneOf` branch set changed: not orderable.

> [!NOTE]
> The two constraints are not orderable, so compatibility cannot be established structurally. Verify by hand.

### tool.output.constraint.loosened

**Default severity:** `BREAKING`

A bound moved outward.

> [!NOTE]
> Consumers validating results against the old bound will start rejecting them.

### tool.output.constraint.tightened

**Default severity:** `NON_BREAKING`

A bound moved inward (`minLength`, `maximum`, `additionalProperties: false`, `uniqueItems`, `allOf` branch added, `anyOf` branch removed).

> [!NOTE]
> Input that was valid is now rejected. Loosen the bound, or stage the change.

### tool.output.default.changed

**Default severity:** `POTENTIALLY_BREAKING`

A `default` was added, removed or changed. The schema still validates; the behavior does not match.

> [!NOTE]
> The schema still validates, but calls that omit this argument now behave differently. Check telemetry for how often it is omitted before shipping.

### tool.output.deprecated.added

**Default severity:** `NON_BREAKING`

`deprecated: true` appeared on a property.

> [!NOTE]
> Check usage before removing it in a later release.

### tool.output.deprecated.removed

**Default severity:** `NON_BREAKING`

A property is no longer marked deprecated.

### tool.output.description.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

A property `description` **or** `title` changed. Models read these when filling arguments.

> [!NOTE]
> Models read property descriptions when filling arguments. A changed one can change what gets passed.

### tool.output.enum.narrowed

**Default severity:** `NON_BREAKING`

Values removed from an `enum`.

> [!NOTE]
> Callers passing a removed value will now fail. Accept the old values and map them for one release.

### tool.output.enum.widened

**Default severity:** `BREAKING`

Values added to an `enum`.

> [!NOTE]
> Consumers switching exhaustively on this enum will not handle the new values.

### tool.output.header.changed

**Default severity:** `BREAKING`

`x-mcp-header` added, removed or changed: it alters how the call is transmitted over HTTP.

> [!NOTE]
> x-mcp-header changes how the call is transmitted over HTTP, so intermediaries routing on the mirrored header will stop matching.

### tool.output.property.added

**Default severity:** `NON_BREAKING`

A new optional property.

### tool.output.property.addedRequired

**Default severity:** `NON_BREAKING`

A new property that is immediately required.

> [!NOTE]
> Every existing call omits this property and will now fail validation. Add it as optional with a default, then require it a release later.

### tool.output.property.removed

**Default severity:** `BREAKING`

A property is gone.

> [!NOTE]
> Accept and ignore the property for one release so existing callers keep working, then remove it.

### tool.output.removed

**Default severity:** `BREAKING`

A declared `outputSchema` disappeared.

> [!NOTE]
> A client that generated types from this schema loses them. Restore the schema, or release the removal as a major version.

### tool.output.required.added

**Default severity:** `NON_BREAKING`

Optional → required.

> [!NOTE]
> Accept calls that omit it for at least one release, applying the previous default, then require it.

### tool.output.required.removed

**Default severity:** `BREAKING`

Required → optional.

> [!NOTE]
> Consumers may assume this field is always present in a result. Keep emitting it, or release the change as a major version.

### tool.output.schema.changed

**Default severity:** `BREAKING`

**The completeness net.** A difference the structured walk does not model. See §9.

> [!NOTE]
> This schema keyword is not modelled by the diff engine, so the change is reported as breaking until a human confirms otherwise.

### tool.output.type.changed

**Default severity:** `BREAKING`

Type sets are not comparable.

> [!NOTE]
> The type sets are not comparable, so no existing caller is guaranteed to still work.

### tool.output.type.narrowed

**Default severity:** `NON_BREAKING`

The `type` set shrank.

> [!NOTE]
> Keep accepting the removed types and coerce them for one release.

### tool.output.type.widened

**Default severity:** `BREAKING`

The `type` set grew.

> [!NOTE]
> Consumers typed against the narrower schema will not handle the new types.

## Tools

### tool.added

**Default severity:** `NON_BREAKING`

A new tool appears.

### tool.annotations.added

**Default severity:** `POTENTIALLY_BREAKING`

An annotation appeared. Agents use these to decide whether to call.

> [!NOTE]
> Confirm clients that filter tools by annotation still see this one.

### tool.annotations.changed

**Default severity:** `POTENTIALLY_BREAKING`

An annotation's value changed.

> [!NOTE]
> Confirm clients that filter or rank tools by annotation still behave as intended.

### tool.annotations.removed

**Default severity:** `POTENTIALLY_BREAKING`

An annotation disappeared.

> [!NOTE]
> A client filtering on this annotation will no longer match this tool.

### tool.annotations.safetyWeakened

**Default severity:** `POTENTIALLY_BREAKING`

`readOnlyHint`/`idempotentHint` left `true`, or `destructiveHint`/`openWorldHint` left `false`. The tool withdrew a promise. Escalate this rule to `BREAKING` in config if you gate on annotations.

> [!NOTE]
> The tool withdrew a behavioural promise. Clients that auto-approve read-only or idempotent tools will now prompt, or worse, will have already auto-approved a tool that is no longer read-only. Set severity_overrides.tool.annotations.safetyWeakened to BREAKING if you gate on annotations.

### tool.description.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

`description` changed. This is the primary tool-selection signal.

> [!NOTE]
> The description is the primary tool-selection signal. Re-run your evaluations: a model may now pick this tool where it did not, or stop picking it where it did.

### tool.icons.changed

**Default severity:** `DOCUMENTATION`

`icons` changed.

### tool.meta.changed

**Default severity:** `NON_BREAKING`

`_meta` changed. Emitted for every `_meta` change; the MCP Apps keys are also interpreted by the `tool.ui.*` rules below (§7.1).

### tool.removed

**Default severity:** `BREAKING`

A tool is gone.

> [!NOTE]
> Keep the tool deprecated for at least one release. Review observed usage and unmeasured consumers before removal; no observed calls is not proof of no dependency.

### tool.renamed

**Default severity:** `BREAKING`

A removal and an addition share an identical schema fingerprint, one candidate each way. Reported as one change instead of two.

> [!NOTE]
> Keep the old name as an alias that forwards to the new one for at least one release. Agents have the old name in their context and in cached tool lists.

### tool.title.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

`title` changed. Clients display it; models read it.

### tool.ui.resourceUri.added

**Default severity:** `NON_BREAKING`

The tool now links a `ui://` template, so hosts render its results with a UI.

### tool.ui.resourceUri.changed

**Default severity:** `POTENTIALLY_BREAKING`

The tool links a different template. Hosts may have prefetched and cached the old one.

> [!NOTE]
> Hosts may have prefetched and cached the old template. Keep serving the old URI for a release, and make sure the new template accepts the structuredContent this tool returns.

### tool.ui.resourceUri.removed

**Default severity:** `POTENTIALLY_BREAKING`

The tool no longer links a UI template. Calls succeed; hosts fall back to text and users lose the UI.

> [!NOTE]
> Hosts fall back to plain text, so calls still succeed, but users lose the UI they had. If this is intentional, say so in the release notes; otherwise restore `_meta.ui.resourceUri`.

### tool.ui.visibility.narrowed

**Default severity:** `BREAKING`

The tool lost the `"model"` audience (hosts drop it from the agent's tool list) or the `"app"` audience (hosts reject the UI's calls).

> [!NOTE]
> Without "model" the host takes the tool out of the agent's tool list, which is a removal for every agent; without "app" the host rejects calls from the UI. Keep the previous visibility for a release, or add a replacement first.

### tool.ui.visibility.widened

**Default severity:** `NON_BREAKING`

The tool gained the `"model"` or `"app"` audience.

## Resource templates

### resourceTemplate.added

**Default severity:** `NON_BREAKING`

A new `uriTemplate`.

### resourceTemplate.annotations.changed

**Default severity:** `POTENTIALLY_BREAKING`

—

### resourceTemplate.description.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

—

### resourceTemplate.meta.changed

**Default severity:** `NON_BREAKING`

`_meta` changed.

### resourceTemplate.mimeType.changed

**Default severity:** `BREAKING`

—

> [!NOTE]
> Consumers parsing the body by its declared type will fail.

### resourceTemplate.name.changed

**Default severity:** `POTENTIALLY_BREAKING`

—

### resourceTemplate.removed

**Default severity:** `BREAKING`

A `uriTemplate` is gone.

> [!NOTE]
> Clients constructing URIs from this template will stop being able to.

### resourceTemplate.title.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

—

## Resources

### resource.added

**Default severity:** `NON_BREAKING`

A new resource URI.

### resource.annotations.changed

**Default severity:** `POTENTIALLY_BREAKING`

`audience` / `priority` etc. changed.

> [!NOTE]
> Annotations like audience and priority steer which resources a client surfaces.

### resource.description.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

`description` changed.

### resource.meta.changed

**Default severity:** `NON_BREAKING`

`_meta` changed. Emitted for every `_meta` change, alongside any `resource.ui.*` rule.

### resource.mimeType.changed

**Default severity:** `BREAKING`

A consumer parsing the body will fail.

> [!NOTE]
> Consumers parsing the body by its declared type will fail.

### resource.name.changed

**Default severity:** `POTENTIALLY_BREAKING`

Models select resources by name.

> [!NOTE]
> Models select resources by name; a rename changes what gets read.

### resource.removed

**Default severity:** `BREAKING`

A resource URI is gone.

> [!NOTE]
> Clients holding this URI will start getting errors. Keep it, or redirect it.

### resource.title.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

`title` changed.

### resource.ui.csp.narrowed

**Default severity:** `BREAKING`

An origin left a UI template's CSP list, so the host blocks the template's requests to it.

> [!NOTE]
> The host builds the template's Content-Security-Policy from these lists, so requests to a removed origin are blocked. Remove the origin only after the template stops using it.

### resource.ui.csp.widened

**Default severity:** `NON_BREAKING`

An origin was added to a UI template's CSP list.

### resource.ui.domain.changed

**Default severity:** `POTENTIALLY_BREAKING`

The UI template's dedicated origin was set, unset or changed. Anything keyed on the origin must follow.

> [!NOTE]
> The template now runs on a different origin. Update anything keyed on the old one (OAuth redirect URIs, CORS and API-key allowlists) before releasing.

### resource.ui.permissions.added

**Default severity:** `NON_BREAKING`

The template requests a new sandbox permission.

### resource.ui.permissions.removed

**Default severity:** `POTENTIALLY_BREAKING`

The template no longer requests a sandbox permission (`camera`, `microphone`, `geolocation`, `clipboardWrite`).

> [!NOTE]
> The host stops granting this browser capability to the template. Confirm the template detects its absence instead of failing.

### resource.ui.widgetDescription.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

`openai/widgetDescription` changed. ChatGPT puts it in the model's context when the template loads.

> [!NOTE]
> ChatGPT puts this summary in the model's context when the template loads. Re-run any evaluation that depends on how the model talks about the UI.

## Prompts

### prompt.added

**Default severity:** `NON_BREAKING`

A new prompt.

### prompt.argument.added

**Default severity:** `NON_BREAKING`

A new optional argument.

### prompt.argument.addedRequired

**Default severity:** `BREAKING`

A new required argument.

> [!NOTE]
> Every existing invocation omits it. Add it as optional first.

### prompt.argument.description.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

—

### prompt.argument.nowOptional

**Default severity:** `NON_BREAKING`

Required → optional.

### prompt.argument.nowRequired

**Default severity:** `BREAKING`

Optional → required.

> [!NOTE]
> Accept invocations that omit it for at least one release.

### prompt.argument.removed

**Default severity:** `BREAKING`

An argument is gone.

> [!NOTE]
> Accept and ignore the argument for one release, then remove it.

### prompt.description.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

—

### prompt.removed

**Default severity:** `BREAKING`

A prompt is gone.

> [!NOTE]
> Clients invoking this prompt by name will start getting errors.

### prompt.title.changed

**Default severity:** `DOCUMENTATION` or `POTENTIALLY_BREAKING`

—

## Severities

| Severity | Means | CI default |
| --- | --- | --- |
| `BREAKING` | A correct existing caller can stop working. | fail |
| `POTENTIALLY_BREAKING` | Behaviour or model-visible semantics changed. | warn |
| `NON_BREAKING` | Strictly additive or widening. | pass |
| `DOCUMENTATION` | Prose changed too little to plausibly move a model. | pass |

Severity is assigned by pure code. No model participates in it.

---

# Severities

> Understand breaking, potentially breaking, non-breaking and documentation changes in MCP contracts, with deterministic CI failure thresholds.

| Severity               | Means                                                                             | CI default |
| ---------------------- | --------------------------------------------------------------------------------- | ---------- |
| `BREAKING`             | A correct existing caller can stop working.                                       | fail       |
| `POTENTIALLY_BREAKING` | Behaviour or model-visible semantics changed. Callers may drift without erroring. | warn       |
| `NON_BREAKING`         | Strictly additive or widening for the consumer of that side of the call.          | pass       |
| `DOCUMENTATION`        | Prose changed too little to plausibly move a model.                               | pass       |

No model participates in assigning these. An optional `--semantic` pass can
attach an advisory note to a change the engine already classified; it cannot
raise, lower, add or remove one.

## Variance: why input and output differ

An `inputSchema` is **written** by the caller, so it is contravariant: loosening
it keeps existing calls valid, tightening it breaks them.

An `outputSchema` is **read** by the caller, so it is covariant: the rules
invert. A server that starts returning `string | number` where it promised
`string` breaks every consumer that generated types from the old schema.

So `tool.input.type.widened` is `NON_BREAKING` while
`tool.output.type.widened` is `BREAKING`. This is the most-questioned pair in
the engine; it is not backwards.

## Prose is part of the contract

Tool descriptions, titles, property descriptions, annotations and server
instructions steer tool selection. They are diffed and classified, never dropped
as "just docs".

The question the classifier answers is _"could this change move a model's tool
selection?"_ — not _"is this text different?"_. In order:

1. Normalise: lowercase, collapse whitespace, strip trailing punctuation.
   Identical → `DOCUMENTATION`.
2. One side empty and the other not → `POTENTIALLY_BREAKING`. Adding or removing
   a description is never a typo fix.
3. A **semantic marker** differs → `POTENTIALLY_BREAKING`, whatever the
   similarity. Markers are negations (`not`, `never`, `without`), modality
   (`must`, `should`, `only`, `required`, `optional`) and any number.
   `limit defaults to 10` → `limit defaults to 100` is a two-character edit that
   changes behaviour.
4. **Same words, different order** → `POTENTIALLY_BREAKING`. An instruction's
   order is its meaning: _use search before create_ is not _use create before
   search_, and every similarity score calls those identical.
5. Otherwise compare character-level edit similarity and token-set overlap.
   `DOCUMENTATION` when one score is high and neither is low.

Two scores, because each alone has a known failure: character similarity calls a
rewritten sentence of the same length "similar", and token overlap calls one
typo in a four-word description "different".

This deliberately errs toward `POTENTIALLY_BREAKING`. The default threshold only
warns on it, so a false positive costs a line in a comment while a false
negative silently ships a tool-selection regression.

## Unmodelled means breaking

A JSON Schema keyword the engine does not model is reported as `BREAKING`.

Two independent implementations compare schemas: a structured walk that produces
named rules with old and new values, and a variance-aware differ that reports
_everything_ it does not recognise. The stricter one always wins. A keyword added
to JSON Schema next year is reported on the day it appears in your schema rather
than silently ignored until someone notices.

Duplicate reporting is the acceptable failure mode here. Silence is not.

## Changing a classification for your project

```yaml
severity_overrides:
  tool.annotations.safetyWeakened: BREAKING
  tool.output.enum.widened: POTENTIALLY_BREAKING

ignore:
  - rule: tool.icons.changed
  - rule: tool.description.changed
    entity: legacy_*
    reason: "Legacy tools are being rewritten; churn is expected."
```

A suppressed change is not deleted — reports print the count and the reason, so a
silenced rule stays visible in review. When _everything_ is suppressed the
summary says so rather than claiming there were no changes.

---

# Telemetry

> Add optional MCP usage telemetry or OpenTelemetry export. Record tool outcomes and argument paths without recording values or awaiting network delivery.

Schema diffing tells you what changed. Telemetry tells you whether anything
depends on it.

Hosted preview status: ingest is deployed at `https://ingest.mcprobe.dev`.
Usage queries and nightly rollups are pending an analytics read credential.
Package publication and paid checkout are pending; the plan table below
describes the implemented plans.

```ts
import { instrumentServer } from "@mcp-inspect/instrumentation";

instrumentServer(server, {
  apiKey: process.env.MCP_INSPECT_KEY,
  endpoint: "https://ingest.mcprobe.dev",
});
```

It patches the handlers for `tools/call`, `resources/read` and `prompts/get` —
the three methods that represent _use_. It never patches `tools/list`: counting
discovery as usage would make every unused tool look used, which is the exact
question this exists to answer.

## The event

```json
{
  "entityKind": "tool",
  "entityName": "search",
  "argumentsPresent": ["limit", "query"],
  "durationMs": 42,
  "success": true
}
```

Everything else — client name and version, protocol version, server version,
surface fingerprint, environment, an anonymous per-process session id — is
optional enrichment. A field the SDK cannot determine is **omitted, never
guessed**: inventing an identity the protocol did not supply makes every
downstream conclusion a lie.

## Guarantees, in order

1. It never throws into your handler.
2. It never delays a response.
3. It never grows without bound.
4. It probably delivers your events.

Events are buffered in memory (bounded, oldest dropped at capacity) and flushed
on a timer, at a batch size, or on `flush()`. On failure a batch is **dropped**,
not retried — a telemetry backlog that survives in memory is a memory leak in
your production server.

## OpenTelemetry

If you would rather keep the data:

```ts
import { otelExporter } from "@mcp-inspect/instrumentation/otel";

instrumentServer(server, { emit: otelExporter(tracer) });
```

Attributes follow the published conventions (`mcp.tool.name`,
`mcp.request.outcome`, `mcp.protocol.version`), so the spans are readable by
tooling that has never heard of this product. MCP reserves `traceparent`,
`tracestate` and `baggage` in `_meta` as an explicit exception to its prefix
rules; the bridge reads and writes those rather than inventing a correlation id,
so a span joins whatever trace the caller already had.

## What you get back

> `search.limit` has never been supplied in 2,712,441 observed calls over 90
> days. Telemetry shows the absence of observed usage, not the absence of all
> possible consumers.

That caveat travels with every number. There is no `safeToRemove` anywhere in
the product: the strongest available phrasing is "candidate for deprecation",
always beside its observation window and confidence.

## Quotas

Whatever the plan, going over an allowance **never returns an error your SDK
might act on**. A quota that can take your server down is an outage with a
billing explanation attached. What differs is what happens to the events.

| Plan           | Included per month | Past the allowance                  | Retention |
| -------------- | ------------------ | ----------------------------------- | --------- |
| Free           | 1,000,000 events   | dropped, and reported to you        | 90 days   |
| Team — $100/mo | 50,000,000 events  | billed at $2 per additional million | 365 days  |

On the free plan events past the cap are dropped and the count is shown in the
console, so you can see exactly what you did not keep.

On a paid plan nothing is dropped. Paying for a limit _and_ losing data at it
would be the worst of both models, so the allowance becomes a billing boundary
rather than a wall: past 50 million events you are charged $2 per additional
million. A partial million is not billed — 50.4M events costs $100, not $102.

Usage is metered out of band from a nightly job, never from the ingest request,
so a billing provider being down cannot affect whether your telemetry is
accepted. The console shows the running total, and the invoice comes from the
billing portal.

