Operations & Gateways
MCP Server Logging: Safe Diagnostics and Correlation
By MCP Beast·
Useful MCP server logging tells an operator where an operation failed, which request it belonged to, and what can safely happen next. Start with structured events at meaningful boundaries, keep secrets and unnecessary content out, and verify that the events reach the intended reader. Logging complete tool arguments by default is usually a poor substitute for choosing useful diagnostic fields.
Protocol Logging is deprecated in MCP 2026-07-28. The current specification advises new implementations not to adopt it and recommends stderr for stdio or OpenTelemetry for structured observability. Existing support remains a compatibility topic; an old logging/setLevel example should not be presented as the current setup procedure.
Decide which question each signal answers
Do not treat every emitted message as the same kind of evidence. A startup diagnostic helps explain a failed launch. A trace links timed work across components. An audit record serves a defined accountability purpose. They may share correlation identifiers, but their completeness, access rules and retention can differ.
| Signal | Useful question | Limit to remember |
|---|---|---|
| Local stderr | Why did this subprocess fail or exit? | Host capture and presentation vary |
| Structured application log | Which handler or dependency failed? | Fields and delivery are implementation-defined |
| Trace/span data | Where did time accumulate across the request? | Collection and sampling affect visibility |
| Protocol log notification | What diagnostic did a compatible server expose to the client? | Deprecated feature and revision-specific behavior |
| Audit event | Who attempted or completed a controlled action? | Requires its own evidence and retention design |
For reconstructing consequential actions, use the audit-evidence guide. This article focuses on operational diagnosis: locating a failing boundary without turning every tool payload into a permanent record.
Keep stdio diagnostics off the protocol channel
The current stdio binding reserves stdout for valid MCP messages and permits diagnostic text on stderr. A client may capture, forward or ignore stderr; a line there does not automatically mean the operation failed.
Test the actual host's behavior. A server may write a helpful startup error that nobody sees because the host does not persist its stderr. Conversely, a wrapper may merge stderr into stdout and make valid diagnostics look like corrupt protocol traffic.
Use one controlled failure, such as a missing synthetic configuration value, to verify the path. Confirm that the operator can find the event and that the client does not attempt to parse it as JSON-RPC. The connection-closed guide covers launch and stream failures that this distinction helps diagnose.
Choose fields from a specific failure question
Consider a fictional lookup_article operation that cannot reach its approved content source. An operator needs to know which boundary failed, which release handled it, whether another attempt is appropriate, and how to locate related events. The entire article query or returned document may add little to that diagnosis.
This is an original application-log example, not an MCP message, OpenTelemetry schema or MCP Beast export:
{
"event": "dependency_request_finished",
"request_ref": "demo-request-17",
"component": "article_lookup",
"release": "example-release-a",
"dependency": "approved_content_source",
"outcome": "timeout",
"elapsed_ms": 2400,
"attempt": 1,
"retry_decision": "not_attempted",
"arguments_recorded": false
}
The duration and identifiers are fictional. The record distinguishes a dependency timeout from a client deadline and says no retry was attempted. It does not claim the external operation had no effect. If that distinction matters, record a separate verified outcome or an explicit unresolved state.
Use this dictionary to interpret that fictional event consistently. These are application-defined meanings, not OpenTelemetry semantic conventions or an MCP wire schema. “Authoritative” identifies the component that observed or assigned the field; it does not make the log a complete audit history.
| Field | Authoritative emitter | Precise meaning in this example | Meaning if absent | Allowed disclosure |
|---|---|---|---|---|
request_ref | Request handler assigns the reference; dependency wrapper propagates it | Correlates this dependency attempt with one accepted application request; not an authenticated identity | This event cannot be joined confidently without other evidence; do not guess from nearby timestamps | Opaque reference in access-controlled diagnostics; no embedded user data or credentials |
outcome | Dependency wrapper observing the attempt | timeout means its wait expired; it does not establish downstream success, failure or absence of effects | Attempt outcome is unreported/unknown, never implicit success | Approved bounded category, not raw provider text or payload |
elapsed_ms | Dependency wrapper's timer | Duration from this dependency attempt's start to the observed timeout; excludes earlier handler work and later response delivery | Duration unknown, not zero | Numeric duration for this boundary; no unrelated timing trace required |
attempt | Component coordinating dependency attempts | One-based attempt index within this request's dependency operation; 1 is the initial attempt | Attempt ordering unknown; do not infer that no retry occurred | Integer index with its request/dependency context |
retry_decision | Retry coordinator at event emission | not_attempted means no retry has been started as of this event; it is not a promise that no later retry will occur | Retry disposition unknown, not an instruction to retry | Approved category; restricted policy/evidence reference if explanation is needed |
Here, 2400 describes only the dependency wait, not end-to-end tool latency. A missing completion event is likewise a collection or outcome gap, not proof of success. If a later retry or reconciliation changes the known state, correlate a new event rather than reading that later knowledge into this earlier record.
Redact before data reaches broad log access
OWASP's logging guidance discusses excluding sensitive data, sanitizing event data, controlling access and testing failure conditions. OpenTelemetry's sensitive-data guidance emphasizes minimizing collection and removing sensitive attributes where appropriate.
Apply that to your own event design. Prefer a known operation name and an approved classification over a raw free-text query. Record that a credential was absent or rejected rather than recording its value. Avoid full authorization headers, cookies, access tokens and unrestricted exception objects.
A practical redaction exercise uses synthetic marker values in several locations: an argument, a nested field, an exception message, and an upstream error response. Cause the controlled failure, then inspect every destination that receives the event. Verify the markers are absent where the policy says they must be absent. These are proposed tests, not a claim that one regular expression can identify every secret.
Do not assume hashing makes arbitrary sensitive data anonymous. Low-entropy identifiers can remain guessable, and stable hashes can still link activity. Choose identifiers and retention according to the diagnostic need, with the relevant data owner involved.
Correlate requests without confusing identity
Use a request or trace reference to join events from the client-facing handler, routing layer and downstream call. Distinguish the authenticated principal from self-reported client metadata. A display name in protocol metadata helps debugging but does not prove who was authorized to act.
For concurrent requests, check that context propagation does not mix their events. A useful synthetic test runs two operations with distinct references and different controlled failures, then reconstructs each timeline independently. If one event cannot be assigned confidently, improve correlation before increasing log volume.
Keep retries visible as attempts of a logical operation where your design supports that relationship. Otherwise an operator may mistake three attempts for three independent user actions. The timeout guide shows why that matters when a late result and a retry overlap.
Make debug capture temporary and bounded
Normal logs should answer common operational questions without enabling verbose mode. When deeper capture is necessary, scope it to the relevant component, operation or approved test environment, and record who owns the capture. Set an end condition so temporary diagnostic detail does not become permanent collection.
Bound event size and volume. A large tool result or recursively serialized error can consume the log budget needed to preserve the first useful failure. Consider truncation indicators, dropped-event counters and separate restricted attachments when detailed evidence is justified. Do not silently truncate the only field that explains the outcome.
Also test a slow or unavailable log destination. Decide whether the application buffers, drops optional diagnostics or blocks a particular operation under its evidence policy. Diagnostic availability and mandatory audit evidence may call for different behavior; make that decision explicit instead of letting a library default decide it accidentally.
Maintain legacy compatibility deliberately
If an existing client relies on protocol Logging, record its revision and required behavior. The current deprecated feature uses request-scoped log metadata and notifications; older lifecycle examples may use different controls. Test the specific supported pairing while planning its replacement.
A migration to structured observability should preserve useful event meaning and correlation, not merely change the export destination. Verify that operators can still find launch failures, authorization denials and dependency errors. Keep historical queries understandable when field names or outcome categories change.
Add the logging destination, access owner and diagnostic procedure to the server-management record. During an incident, the operator should not need to guess whether the relevant evidence lives in a desktop client, service runtime or collector.
Frequently Asked Questions
Is MCP protocol Logging deprecated?
Yes. Revision 2026-07-28 deprecates the feature and recommends stderr for stdio or OpenTelemetry for structured observability. Existing compatibility should be documented separately from a new implementation.
Where should a stdio MCP server write diagnostics?
Use stderr or the supported diagnostic destination. Stdout is reserved for valid MCP protocol messages, and ordinary log text there can break parsing.
Should I log every tool argument and result?
Not by default. Select fields that answer a diagnostic question, minimize sensitive content, and test redaction across the complete collection path.
Are diagnostic logs sufficient for an audit trail?
Not automatically. Audit evidence has its own requirements for identity, action, outcome, access, completeness and retention. Correlation can connect the two, but their purposes differ.
Choose one common failure and confirm that a teammate can locate its cause from the retained fields. Then run a synthetic redaction test before increasing diagnostic detail.