Security & Access Control
MCP Tool Annotations: Hints, Defaults and Trust
By MCP Beast·
MCP tool annotations describe what a server says a tool does. They can help a client present an operation and reason about its effects, but they do not enforce permissions or prove that the implementation behaves as advertised. Treat an annotation as a claim to evaluate alongside the server's identity, the tool contract and independent controls.
For an operator, the practical question is not whether a tool has a reassuring label. It is whether the intended identity can perform only the approved operation, against the approved destination, with a result you can verify. This guide uses MCP 2026-07-28 and separates the standard hints from a proposed review workflow.
Read the hints and defaults together
The ToolAnnotations schema defines these optional fields. Destructive and idempotent hints are meaningful when readOnlyHint is false. The schema explicitly warns against basing tool-use decisions on annotations from untrusted servers.
| Field | Meaning of true | Default when omitted |
|---|---|---|
readOnlyHint | Does not modify its environment | false |
destructiveHint | May make destructive changes | true |
idempotentHint | Repeating identical arguments adds no further effect | false |
openWorldHint | May interact with external entities | true |
These are interpretation defaults, not a measured assessment of the tool. An omitted read-only hint does not prove that an operation writes, and an explicit true value does not make a write impossible. Keep the distinction visible in review notes: “publisher declares read-only” and “test identity cannot write” are different statements backed by different evidence.
Review a small tool definition
Use the MCP tool schema reviewer for a local structural pass over sanitized definitions before reviewing actual behavior. It checks selected schema and annotation shapes; it does not establish authorization or safety.
This fictional definition is a structural excerpt, not a complete MCP response or a runnable server. It describes a lookup against a bounded synthetic catalog:
{
"name": "lookup_fixture",
"description": "Read a synthetic record from the test catalog.",
"inputSchema": {
"type": "object",
"properties": {
"record_id": { "type": "string" }
},
"required": ["record_id"],
"additionalProperties": false
},
"annotations": {
"readOnlyHint": true,
"openWorldHint": false
}
}
The definition is easy to read, but the implementation could still be wrong. A handler might update a record as part of a lookup, use a broadly privileged downstream account or return another workspace's object. None of those properties is settled by the JSON above.
For this example, propose three independent checks: inspect the handler's intended behavior, run it with a credential that lacks write permission, and compare the test destination before and after the call. Use synthetic records and record which implementation version you inspected. The Inspector workflow gives a repeatable way to organize those success and failure cases.
Separate four decisions before execution
First, identify the server and exact tool. A friendly title is presentation data. Preserve the server-qualified identity so a similarly named tool from a different connector does not inherit a previous review.
Second, validate the arguments against the current schema. A bounded record identifier is a different contract from an unrestricted query or arbitrary command. Review destination fields, paths and filters that determine what the operation touches. A valid schema describes acceptable structure; it does not decide whether this caller is allowed to use those values.
Third, enforce authorization at the operation boundary. Restrict the downstream credential, check the user's access to the target object and apply the approval rule for the actual effect. OWASP's excessive-agency guidance supports limiting functionality and permissions rather than relying on model judgment alone.
Fourth, verify the outcome. A successful response should correspond to the intended destination state. Keep enough redacted evidence to investigate a discrepancy without storing full sensitive arguments. These decisions work together; a correct annotation does not replace any of them.
Do not turn idempotency into a blanket retry rule
Imagine a fictional tool that creates a support note and returns its identifier. The first attempt reaches the downstream system, but the network fails before the client receives the result. The client now has an uncertain outcome. Retrying immediately could create a second note unless the implementation has a suitable duplicate-prevention contract.
An idempotency claim needs a scope. Ask whether it applies to identical arguments, a specific operation key, one caller or a bounded time window. Ask what happens when two requests arrive concurrently and whether the downstream system participates in deduplication. These are implementation questions; do not invent a standard MCP idempotency-key field to answer them.
In a test environment, repeat the approved operation and inspect the resulting records. Then test an interrupted response if your harness can do so safely. Record what was observed and the limits of the test. For a real uncertain write, inspect the destination before repeating it. The timeout troubleshooting guide explains why a missing response is not proof that no work occurred.
Use this narrow repeat-operation evidence card for one reviewed tool and release. The observations are blank; a declaration alone cannot fill them.
| Evidence field | What to establish | Observed value / restricted evidence reference |
|---|---|---|
| Declared hint | Exact idempotentHint declaration and server-qualified tool/release | |
| Business effect | Which records, messages, charges or other effects one successful call produces | |
| Duplicate-protection owner | Component and operator responsible for preventing repeated effects, if any | |
| Caller / key / time scope | Documented identity, argument or application-key binding and any expiry; unknown where unverified | |
| Identical repeat result | Destination state after a sequential repeat with identical arguments under the same caller | |
| Concurrent repeat result | Destination state when matching requests overlap in the controlled test | |
| Interrupted-response reconciliation | Safe destination lookup or operation-status evidence after the response is lost | |
| Untested cases | Other callers, changed arguments, expired windows, releases or failure paths not evaluated |
The annotation's claim concerns repeated calls with identical arguments. A separate application key contract may provide duplicate protection only for a documented caller, key and time scope; it does not automatically prove the broader annotation claim. No key field in this card is a standard MCP parameter. Use the recorded contract and destination evidence to justify a retry, or keep the outcome unresolved until reconciliation is possible.
Treat additive and closed-world operations carefully
A tool that only adds information can still send an unwanted message, incur a charge or disclose data. “Not destructive” is not the same as harmless. Review the business effect and destination instead of reducing every decision to whether an existing record is deleted.
Likewise, a bounded interaction domain does not prove that returned content is safe to follow as instructions. A private document can contain misleading text, and an internal database can contain user-generated content. Keep retrieved content in its data role and apply the same authorization boundary when a later action is proposed.
The current tools specification supplies the discovery and invocation contract. It does not turn descriptive metadata into a security boundary. If you need to constrain file, network or account access, implement and test those constraints in the runtime and downstream permissions that actually govern them.
Revisit the review when the catalog changes
Record the reviewed server identity, release, schema and annotations. A tool can retain its name while its description, arguments or behavior change. Compare the current catalog against the reviewed baseline before carrying an earlier approval forward. The server versioning guide describes that comparison separately from protocol compatibility.
When a hint disagrees with observed behavior, preserve a redacted reproduction and suspend any policy shortcut that depended on the hint. Correct the declaration or implementation at its source, then repeat the relevant test. Do not silently rewrite the metadata in a client and assume every other consumer is now protected.
A useful final review note states the claim, the independent controls, the tested behavior and the remaining uncertainty. For example: “Lookup declared read-only; test credential lacks writes; synthetic record unchanged after one valid and one denied request; other server versions not assessed.” That is more useful than a generic safe badge because another operator can understand exactly what the evidence covers.
Frequently Asked Questions
Do MCP tool annotations enforce security?
No. They are descriptive hints. Enforce access and operation limits independently, and review the server and implementation behind the claims.
What happens when a tool omits annotations?
The schema defines defaults: readOnlyHint and idempotentHint are false, while destructiveHint and openWorldHint are true. These defaults are not observations of actual behavior.
Does idempotentHint mean every failed call can be retried?
No. Verify the implementation's repeat-operation contract and the destination state after an uncertain outcome before deciding whether a retry is safe.
Does destructiveHint false mean an operation is harmless?
No. An additive operation can still send information, create unwanted records or incur costs. Review the intended effect, destination, permissions and approval requirements.
When evaluating an MCP routing setup, bring one tool whose annotations matter to your workflow. Review its declaration, downstream permissions and a synthetic success-and-denial test together.