Can an MCP tool description change break an agent?

A tool can accept the same arguments while a rewritten description changes when an agent selects it. Review prose changes alongside schema changes.

Changing an MCP tool description can change tool selection without changing the input schema. The request may still validate, but an agent can choose a different tool or use it for a different purpose. That makes descriptions part of the model-visible contract.

This is a reason to review prose and run evaluations, not a claim that every wording edit breaks every model.

#A schema-compatible change worth reviewing

Before:

Search the issue tracker.

After:

Find tickets across every connected workspace.

The tool remains named search and its arguments stay identical. Its stated scope changes from one tracker to every connected workspace. MCP Inspect emits tool.description.changed at POTENTIALLY_BREAKING for this example. See the actual word redline and engine output.

Description classification uses a deterministic prose heuristic. The finding is reproducible from the two captures; it does not measure selection probability or predict a particular model's response. Small edits can be classified as DOCUMENTATION, and even then your own evaluations remain useful.

#Which model-visible fields matter?

FieldReview question
Tool description and titleHas the stated purpose, scope or preferred use changed?
Input property descriptionsWill a model supply a different meaning or unit for the same property?
Server instructionsDoes guidance change how the server's tools should be combined or selected?
Tool annotationsDid read-only, destructive or idempotent hints change?
Prompt descriptions and declared argumentsDid a caller's declared interface or guidance change?

Tool annotations are untrusted hints, not proof of safety. A withdrawn hint can still change host or agent behavior. Contract review must not substitute for authorization or approval of the tool's actual effects. The MCP tools specification defines these tool fields; the rule reference documents how MCP Inspect classifies changes that it captures. A snapshot captures prompt declarations, not every dynamically rendered prompt result.

#How to evaluate a description rewrite

  1. Keep the before and after contracts and identify the intended meaning change.
  2. Collect representative tasks, including ambiguous cases involving neighboring tools and cases where no tool should be called.
  3. Run both versions with the same model, host settings and available tool set. Record the selected tool, argument shape and task result. Repeat probabilistic runs rather than treating one result as conclusive.
  4. Review selection changes against the intended behavior. Keep handler and permission tests separate from the selection evaluation.
  5. Ship the wording with a reviewed contract change and retain its baseline.

No benchmark score or model behavior is inferred from a diff alone. MCP Inspect does not run these evaluations for you.

#Track prose changes in your own server

Capture a public HTTP MCP endpoint with email and password, then capture again after a wording change. Open Changes to review the old and new values beside schema findings. No telemetry, API key or GitHub connection is needed for this hosted capture workflow.

For approved-preview CLI/CI workflows, fail_on: POTENTIALLY_BREAKING includes substantial prose and annotation findings in the failure threshold. The default BREAKING threshold still displays those findings but does not fail solely for them. Read the severity settings.

Continue with input versus output compatibility or scheduled monitoring.