Four ways an MCP server fails
MCP server monitoring is not the same as website uptime. A remote MCP server sits between an AI agent and real tools, so the failures that matter are the ones that change what the agent can do — or what it is told to do. There are four:
| Failure | What the agent sees | What catches it |
|---|---|---|
| Down or not speaking MCP | Connection errors, or an HTML error page | An initialize + tools/list check |
| Auth broken | Every call refused | The same check, with the credential the agent uses |
| Tool definitions changed | Different tools, arguments or behaviour than you approved | Fingerprint pinning |
| Poisoned descriptions | Hidden instructions inside a tool’s text | A scan of every new or changed definition |
The first two are availability. The last two are integrity, and they are the reason MCP monitoring deserves its own design: the protocol puts server-written text straight into the model’s context, and the specification says so plainly.
For trust & safety and security, clients MUST consider tool annotations to be untrusted unless they come from trusted servers.
Uptime: check the protocol, not the port
A load balancer answering 200 on / says nothing about whether the MCP endpoint works. A real check does what a client does at connect time: send initialize, send notifications/initialized, then tools/list, and confirm tools come back. That exercises the transport, the session handling, the auth layer and the tool registry in one pass. The protocol’s own ping is for keeping an established session alive, not for an external monitor that connects fresh each time.
#!/usr/bin/env bash
# Synthetic MCP check: initialize, then tools/list, then fingerprint the definitions.
URL="https://mcp.example.com/mcp"
H=(-H "Content-Type: application/json" -H "Accept: application/json, text/event-stream")
SID=$(curl -s -m 20 -D - -o /dev/null "${H[@]}" "$URL" \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"uptime-check","version":"1.0"}}}' \
| awk -F': ' 'tolower($1)=="mcp-session-id" {print $2}' | tr -d '\r')
S=(-H "MCP-Protocol-Version: 2025-06-18"); [ -n "$SID" ] && S+=(-H "Mcp-Session-Id: $SID")
curl -s -m 20 "${H[@]}" "${S[@]}" "$URL" -d '{"jsonrpc":"2.0","method":"notifications/initialized"}' > /dev/null
# A streamable-HTTP server may answer as JSON or as an event stream: keep the JSON either way.
curl -s -m 20 "${H[@]}" "${S[@]}" "$URL" -d '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
| sed -n -e 's/^data: //' -e '/^{/p' > tools.json
jq -e '.result.tools | length > 0' tools.json > /dev/null || { echo "DOWN or no tools"; exit 1; }
jq -S '[.result.tools[] | {name, description, inputSchema, annotations}] | sort_by(.name)' tools.json \
| sha256sum | cut -d' ' -f1 > fingerprint.new
cmp -s fingerprint.new fingerprint.pinned || echo "TOOL DEFINITIONS CHANGED: review before re-pinning"Two details that make the difference between a monitor you trust and one you mute. A 401 is not always down: a server that wants credentials is alive — it becomes a failure only when you gave it credentials and they were refused. And follow pagination: tools/list can return a nextCursor, and a check that reads only the first page will report a “removed” tool every time the server grows past it.
Tool-definition drift: pin fingerprints
When you connect an agent to a server you are implicitly approving its tools as they are today. Nothing in the protocol stops the server from changing a description, adding a parameter or introducing a new tool tomorrow — and a notifications/tools/list_changed message only reaches clients holding an open session. A monitor has to compare for itself.
What goes into the hash decides what you catch. Pin at least the name, description and input schema, plus annotations if your client relies on them; hash a canonical form (sorted keys) so formatting noise does not alert. Be clear about what a monitor you did not build pins: ours fingerprints each tool’s name and description (reading up to 60 tools and the first 240 characters of each description), which catches renamed and rewritten tools but not a schema-only change — pin the full definition in your own CI as well if schema drift matters to you.
Prompt injection in tool descriptions
Tool poisoning hides instructions for the model inside a tool’s description: an <IMPORTANT> block telling it to read a key file first, a line asking it to pass the whole conversation as an argument, or a note to “use this instead of send_email”. The user never sees the text; the model reads all of it. The rug pull is the same attack delivered later — a server that looked clean when you connected rewrites a description after you approved it. The attack side is covered in depth in MCP prompt injection; monitoring is how you catch it in a server you already use.
| Pattern | Example shape | Severity |
|---|---|---|
| Hidden instruction tags | <IMPORTANT>, <system>, <secret> blocks in a description | Blocking |
| Exfiltration requests | “Include the full conversation / system prompt in the notes argument” | Blocking |
| Tool shadowing | “Use this instead of another_server_tool” | Blocking |
| Credential reads | References to key files, env files or SSH directories | Blocking |
| Secrets in arguments | “Pass your API key in the token field” — sometimes legitimate | Warning |
Run the scan on the baseline and then only on tools that are new or changed, so one bad definition alerts once rather than on every check forever. And be honest about the limit: this is a text screen. It does not run the tools or read the server’s code, so a clean scan never proves a server is safe — it proves the text you can see is not obviously hostile. The controls that bound the damage are in MCP security best practices.
curl -X POST https://businessmcp.com/api/hub/v1/mcp/scan \
-H "Authorization: Bearer $BUSINESSMCP_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://mcp.example.com/mcp"}'Alert design: page on state changes
An MCP monitor that pages on every failed check gets muted within a week. The design that survives is built on state changes:
- Down after 2 consecutive failures. One dropped packet on an hourly check should never wake anyone; 2 in a row is an outage.
- Recovery only after a down alert. A “recovered” message for an outage nobody was told about is noise.
- Each tool change exactly once. Alert on added, removed and changed tools, then move the pins.
- Blocking scan findings immediately, but only for new or changed tools.
- Auto-pause a dead monitor. A server that fails for a week should stop being checked (and billed): after 168 failures on an hourly monitor, 28 on a six-hourly one or 7 on a daily one, ours pauses itself and says so.
Send alerts where the owner of the agent integration will see them. Our monitors deliver each alert in-app, by email, to Slack when it is connected, and to an alert.raised webhook, and keep an event history you can read back:
curl https://businessmcp.com/api/hub/v1/mcp/monitors/events?limit=5 \
-H "Authorization: Bearer $BUSINESSMCP_API_KEY"The golden-signal thinking from Google’s SRE book applies unchanged: alert on symptoms a user would notice, keep every page actionable, and put everything else in a log.
Cadence and cost
How often to check depends on what a failure costs you. An agent in a customer-facing workflow justifies hourly checks; a research tool you use weekly is fine daily. Creating a monitor runs the baseline scan for 5 credits, and each scheduled check costs 1 credit:
| Interval | Checks a month | Credits a month | List price |
|---|---|---|---|
| hourly | 720 | 720 | $1.44 |
| 6h | 120 | 120 | $0.24 |
| daily | 30 | 30 | $0.06 |
These are first-party tools, so they draw on the 1,000 free monthly credits first — one hourly monitor fits inside them on its own. Each check is charged before the probe runs, with a request id unique to its time slot, so an overlapping run never charges twice; if a check cannot be charged the monitor pauses instead of running unpaid. A workspace can hold up to 25 monitors, and listing, pausing and deleting them is free.
curl -X POST https://businessmcp.com/api/hub/v1/mcp/monitors \
-H "Authorization: Bearer $BUSINESSMCP_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://mcp.example.com/mcp","interval":"hourly"}'The same tools are exposed over MCP, so an agent can set up and read its own monitors — see the trust and safety API for the full list.
Frequently asked questions
How do I check if an MCP server is up?
Run the handshake a client runs: send initialize, then notifications/initialized, then tools/list, and confirm tools come back. An HTTP 200 from the root path does not prove the MCP endpoint, its session handling or its auth work.
What is an MCP rug pull?
A server that looked safe when you connected it later changes a tool description or definition, for example to add hidden instructions for the model. Pinning a fingerprint of every tool and alerting on any change is how you catch it.
Does a clean prompt-injection scan mean an MCP server is safe?
No. A scan reads tool names and descriptions and flags known hostile patterns. It does not run the tools or audit the server code, so it reduces risk rather than proving safety. Pair it with least-privilege scoping and approval gates.
How often should I monitor an MCP server?
Hourly for servers in customer-facing or automated workflows, every six hours or daily for tools used occasionally. Alert only after consecutive failures so a single dropped request never pages anyone.
How much does MCP server monitoring cost on BusinessMCP?
Creating a monitor costs 5 credits for the baseline scan, and each scheduled check costs 1 credit, so an hourly monitor uses about 720 credits in a 30-day month and a daily one about 30. The first 1,000 first-party credits each month are free.
Sources
BusinessMCP Team
Every guide is written from running BusinessMCP on its own platform — the match rates, reply rates, and deliverability lessons are from our own data, not recycled blog folklore. About BusinessMCP
Turn your business into one AI-ready MCP server
Connect your tools, install one tracking script, and expose your unified data to any AI agent through a single secure endpoint.
Get started free