Severities
Understand breaking, potentially breaking, non-breaking and documentation changes in MCP contracts, with deterministic CI failure thresholds.
| Severity | Means | CI default |
|---|---|---|
BREAKING | A correct existing caller can stop working. | fail |
POTENTIALLY_BREAKING | Behaviour or model-visible semantics changed. Callers may drift without erroring. | warn |
NON_BREAKING | Strictly additive or widening for the consumer of that side of the call. | pass |
DOCUMENTATION | Prose changed too little to plausibly move a model. | pass |
No model participates in assigning these. An optional --semantic pass can
attach an advisory note to a change the engine already classified; it cannot
raise, lower, add or remove one.
#Variance: why input and output differ
An inputSchema is written by the caller, so it is contravariant: loosening
it keeps existing calls valid, tightening it breaks them.
An outputSchema is read by the caller, so it is covariant: the rules
invert. A server that starts returning string | number where it promised
string breaks every consumer that generated types from the old schema.
So tool.input.type.widened is NON_BREAKING while
tool.output.type.widened is BREAKING. This is the most-questioned pair in
the engine; it is not backwards.
#Prose is part of the contract
Tool descriptions, titles, property descriptions, annotations and server instructions steer tool selection. They are diffed and classified, never dropped as "just docs".
The question the classifier answers is "could this change move a model's tool selection?" — not "is this text different?". In order:
- Normalise: lowercase, collapse whitespace, strip trailing punctuation.
Identical →
DOCUMENTATION. - One side empty and the other not →
POTENTIALLY_BREAKING. Adding or removing a description is never a typo fix. - A semantic marker differs →
POTENTIALLY_BREAKING, whatever the similarity. Markers are negations (not,never,without), modality (must,should,only,required,optional) and any number.limit defaults to 10→limit defaults to 100is a two-character edit that changes behaviour. - Same words, different order →
POTENTIALLY_BREAKING. An instruction's order is its meaning: use search before create is not use create before search, and every similarity score calls those identical. - Otherwise compare character-level edit similarity and token-set overlap.
DOCUMENTATIONwhen one score is high and neither is low.
Two scores, because each alone has a known failure: character similarity calls a rewritten sentence of the same length "similar", and token overlap calls one typo in a four-word description "different".
This deliberately errs toward POTENTIALLY_BREAKING. The default threshold only
warns on it, so a false positive costs a line in a comment while a false
negative silently ships a tool-selection regression.
#Unmodelled means breaking
A JSON Schema keyword the engine does not model is reported as BREAKING.
Two independent implementations compare schemas: a structured walk that produces named rules with old and new values, and a variance-aware differ that reports everything it does not recognise. The stricter one always wins. A keyword added to JSON Schema next year is reported on the day it appears in your schema rather than silently ignored until someone notices.
Duplicate reporting is the acceptable failure mode here. Silence is not.
#Changing a classification for your project
severity_overrides:
tool.annotations.safetyWeakened: BREAKING
tool.output.enum.widened: POTENTIALLY_BREAKING
ignore:
- rule: tool.icons.changed
- rule: tool.description.changed
entity: legacy_*
reason: "Legacy tools are being rewritten; churn is expected."
A suppressed change is not deleted — reports print the count and the reason, so a silenced rule stays visible in review. When everything is suppressed the summary says so rather than claiming there were no changes.