# Severities

> Understand breaking, potentially breaking, non-breaking and documentation changes in MCP contracts, with deterministic CI failure thresholds.

| Severity               | Means                                                                             | CI default |
| ---------------------- | --------------------------------------------------------------------------------- | ---------- |
| `BREAKING`             | A correct existing caller can stop working.                                       | fail       |
| `POTENTIALLY_BREAKING` | Behaviour or model-visible semantics changed. Callers may drift without erroring. | warn       |
| `NON_BREAKING`         | Strictly additive or widening for the consumer of that side of the call.          | pass       |
| `DOCUMENTATION`        | Prose changed too little to plausibly move a model.                               | pass       |

No model participates in assigning these. An optional `--semantic` pass can
attach an advisory note to a change the engine already classified; it cannot
raise, lower, add or remove one.

## Variance: why input and output differ

An `inputSchema` is **written** by the caller, so it is contravariant: loosening
it keeps existing calls valid, tightening it breaks them.

An `outputSchema` is **read** by the caller, so it is covariant: the rules
invert. A server that starts returning `string | number` where it promised
`string` breaks every consumer that generated types from the old schema.

So `tool.input.type.widened` is `NON_BREAKING` while
`tool.output.type.widened` is `BREAKING`. This is the most-questioned pair in
the engine; it is not backwards.

## Prose is part of the contract

Tool descriptions, titles, property descriptions, annotations and server
instructions steer tool selection. They are diffed and classified, never dropped
as "just docs".

The question the classifier answers is _"could this change move a model's tool
selection?"_ — not _"is this text different?"_. In order:

1. Normalise: lowercase, collapse whitespace, strip trailing punctuation.
   Identical → `DOCUMENTATION`.
2. One side empty and the other not → `POTENTIALLY_BREAKING`. Adding or removing
   a description is never a typo fix.
3. A **semantic marker** differs → `POTENTIALLY_BREAKING`, whatever the
   similarity. Markers are negations (`not`, `never`, `without`), modality
   (`must`, `should`, `only`, `required`, `optional`) and any number.
   `limit defaults to 10` → `limit defaults to 100` is a two-character edit that
   changes behaviour.
4. **Same words, different order** → `POTENTIALLY_BREAKING`. An instruction's
   order is its meaning: _use search before create_ is not _use create before
   search_, and every similarity score calls those identical.
5. Otherwise compare character-level edit similarity and token-set overlap.
   `DOCUMENTATION` when one score is high and neither is low.

Two scores, because each alone has a known failure: character similarity calls a
rewritten sentence of the same length "similar", and token overlap calls one
typo in a four-word description "different".

This deliberately errs toward `POTENTIALLY_BREAKING`. The default threshold only
warns on it, so a false positive costs a line in a comment while a false
negative silently ships a tool-selection regression.

## Unmodelled means breaking

A JSON Schema keyword the engine does not model is reported as `BREAKING`.

Two independent implementations compare schemas: a structured walk that produces
named rules with old and new values, and a variance-aware differ that reports
_everything_ it does not recognise. The stricter one always wins. A keyword added
to JSON Schema next year is reported on the day it appears in your schema rather
than silently ignored until someone notices.

Duplicate reporting is the acceptable failure mode here. Silence is not.

## Changing a classification for your project

```yaml
severity_overrides:
  tool.annotations.safetyWeakened: BREAKING
  tool.output.enum.widened: POTENTIALLY_BREAKING

ignore:
  - rule: tool.icons.changed
  - rule: tool.description.changed
    entity: legacy_*
    reason: "Legacy tools are being rewritten; churn is expected."
```

A suppressed change is not deleted — reports print the count and the reason, so a
silenced rule stays visible in review. When _everything_ is suppressed the
summary says so rather than claiming there were no changes.
