Severities

Understand breaking, potentially breaking, non-breaking and documentation changes in MCP contracts, with deterministic CI failure thresholds.

SeverityMeansCI default
BREAKINGA correct existing caller can stop working.fail
POTENTIALLY_BREAKINGBehaviour or model-visible semantics changed. Callers may drift without erroring.warn
NON_BREAKINGStrictly additive or widening for the consumer of that side of the call.pass
DOCUMENTATIONProse changed too little to plausibly move a model.pass

No model participates in assigning these. An optional --semantic pass can attach an advisory note to a change the engine already classified; it cannot raise, lower, add or remove one.

#Variance: why input and output differ

An inputSchema is written by the caller, so it is contravariant: loosening it keeps existing calls valid, tightening it breaks them.

An outputSchema is read by the caller, so it is covariant: the rules invert. A server that starts returning string | number where it promised string breaks every consumer that generated types from the old schema.

So tool.input.type.widened is NON_BREAKING while tool.output.type.widened is BREAKING. This is the most-questioned pair in the engine; it is not backwards.

#Prose is part of the contract

Tool descriptions, titles, property descriptions, annotations and server instructions steer tool selection. They are diffed and classified, never dropped as "just docs".

The question the classifier answers is "could this change move a model's tool selection?" — not "is this text different?". In order:

  1. Normalise: lowercase, collapse whitespace, strip trailing punctuation. Identical → DOCUMENTATION.
  2. One side empty and the other not → POTENTIALLY_BREAKING. Adding or removing a description is never a typo fix.
  3. A semantic marker differs → POTENTIALLY_BREAKING, whatever the similarity. Markers are negations (not, never, without), modality (must, should, only, required, optional) and any number. limit defaults to 10 → limit defaults to 100 is a two-character edit that changes behaviour.
  4. Same words, different order → POTENTIALLY_BREAKING. An instruction's order is its meaning: use search before create is not use create before search, and every similarity score calls those identical.
  5. Otherwise compare character-level edit similarity and token-set overlap. DOCUMENTATION when one score is high and neither is low.

Two scores, because each alone has a known failure: character similarity calls a rewritten sentence of the same length "similar", and token overlap calls one typo in a four-word description "different".

This deliberately errs toward POTENTIALLY_BREAKING. The default threshold only warns on it, so a false positive costs a line in a comment while a false negative silently ships a tool-selection regression.

#Unmodelled means breaking

A JSON Schema keyword the engine does not model is reported as BREAKING.

Two independent implementations compare schemas: a structured walk that produces named rules with old and new values, and a variance-aware differ that reports everything it does not recognise. The stricter one always wins. A keyword added to JSON Schema next year is reported on the day it appears in your schema rather than silently ignored until someone notices.

Duplicate reporting is the acceptable failure mode here. Silence is not.

#Changing a classification for your project

severity_overrides:
  tool.annotations.safetyWeakened: BREAKING
  tool.output.enum.widened: POTENTIALLY_BREAKING

ignore:
  - rule: tool.icons.changed
  - rule: tool.description.changed
    entity: legacy_*
    reason: "Legacy tools are being rewritten; churn is expected."

A suppressed change is not deleted — reports print the count and the reason, so a silenced rule stays visible in review. When everything is suppressed the summary says so rather than claiming there were no changes.