// Journal · Oct 05, 2026 · 7 min read
How to build an MCP server for an internal tool

To build an MCP server for an internal tool, you wrap the API that tool already has in a handful of typed functions, describe each one in plain English, and serve them over the Model Context Protocol. With FastMCP that is about 30 lines of Python. The server is the easy part. What takes the time is deciding which five operations to expose, how the server authenticates, and what it is allowed to change.
We build these alongside the internal tools we ship, for the same reason those tools get a web interface: people — and now the agents working next to them — need a way in that isn’t a database login.
What an MCP server actually is
MCP is an open protocol for connecting LLM applications to tools and data over JSON-RPC 2.0. A host (a chat app, a coding agent, your own production agent) runs a client that connects to servers, and each server offers some mix of tools (functions the model can call), resources (data) and prompts (templated workflows).
The current specification revision is 2026-07-28. Two transports are standard: stdio, for a server the client launches as a subprocess on the same machine, and Streamable HTTP, where each message is an HTTP POST to a single endpoint. Worth knowing before you design anything: in this revision requests are stateless and self-contained, with per-request capability negotiation, and servers never initiate JSON-RPC requests.
The payoff is that you build the interface once. An MCP server over your orders system works with any MCP client — a desktop assistant, a coding agent, or a production workflow — without a bespoke integration each time. The ecosystem signal is real: the reference modelcontextprotocol/servers repository sits at roughly 91,000 GitHub stars as of October 2026, and FastMCP alone records about 52 million PyPI downloads in a 30-day window (a figure that includes CI and mirror traffic, so read it as a trend, not a user count).
How to build an MCP server for an internal tool in about 30 lines
Here is a read-only server over an internal orders API. It wraps HTTP calls the internal tool already serves to its own front end:
# server.py — MCP server over the internal orders API
import os
from typing import Annotated
import httpx
from fastmcp import FastMCP
from fastmcp.exceptions import ToolError
from mcp.types import ToolAnnotations
from pydantic import Field
mcp = FastMCP(name="acme-orders", mask_error_details=True)
api = httpx.AsyncClient(
base_url="https://orders.internal.acme.test/v1",
headers={"Authorization": f"Bearer {os.environ['ORDERS_API_TOKEN']}"},
timeout=10,
)
@mcp.tool(annotations=ToolAnnotations(readOnlyHint=True, openWorldHint=False))
async def find_orders(
customer_email: Annotated[str, Field(description="Exact billing email on the order.")],
status: Annotated[str | None, Field(description="open, shipped or cancelled")] = None,
limit: Annotated[int, Field(ge=1, le=50)] = 10,
) -> list[dict]:
"""Find a customer's orders, newest first.
Use this before answering any question about delivery or refunds — the
orders system is authoritative and the CRM copy is often days stale.
"""
params = {"email": customer_email, "limit": limit}
if status:
params["status"] = status
response = await api.get("/orders", params=params)
if response.status_code == 404:
raise ToolError(f"No customer found with email {customer_email}.")
response.raise_for_status()
return response.json()["orders"]
if __name__ == "__main__":
mcp.run(transport="http", port=8000)
That is a complete server. FastMCP derives the tool name, description and JSON schema from the function signature, docstring and type hints, so there is no schema to hand-write.
Pin your version. This is written against the FastMCP 4.0.x API as documented — 4.0.0 landed on 31 August 2026 and twelve releases followed in the next five weeks, up to 4.0.11 on 4 October. We checked the decorator and verifier shapes against the current docs rather than against your environment; a library moving that fast deserves a pinned requirement and a quick read of the tools reference before you copy anything.
Expose operations, not endpoints
The temptation is to generate one tool per API route. Resist it. Tools should match what someone would actually ask for — “find this customer’s orders”, “check stock for this SKU” — not the shape of your REST surface. Five good tools beat fifty wrappers, for reasons the token-cost section below makes concrete.
Write the docstring for the model, not for a developer
The docstring is not documentation. It is the prompt the model reads when deciding whether to call your tool, and it is the single highest-leverage thing in the file. Ours always answer three questions:
- What does this return? “A customer’s orders, newest first.”
- When should it be used? “Before answering any question about delivery or refunds.”
- What is authoritative? “The orders system is; the CRM copy is often days stale.”
Annotations like readOnlyHint and destructiveHint help a client decide whether to ask the user first — but the MCP specification is explicit that annotations from an untrusted server are themselves untrusted. They are labels, not controls. The controls come next.
Lock it down before anything connects
A remote MCP server is an authenticated API, and FastMCP’s docs are blunt that the static and debug token verifiers are for development only, never production. For a real deployment, verify a JWT:
from fastmcp import FastMCP
from fastmcp.server.auth.providers.jwt import JWTVerifier
verifier = JWTVerifier(
jwks_uri="https://auth.acme.internal/.well-known/jwks.json",
issuer="https://auth.acme.internal",
audience="acme-orders-mcp",
)
mcp = FastMCP(name="acme-orders", auth=verifier, mask_error_details=True)
Three rules we apply to every one of these:
- The server account gets the least privilege that works. If the service token can issue refunds, then anyone who reaches the server can issue refunds, through a model that was persuaded by an email.
- Writes go through a human. Read tools can be direct; anything that moves money, mails a customer, or deletes a record goes into a review queue and waits for a person.
- Errors don’t leak internals.
mask_error_details=Truemeans only messages you raise asToolErrorreach the client — stack traces and upstream URLs stay on your side.
Keep the tool catalogue small — it costs tokens before anyone asks anything
Every connected server’s tool definitions sit in the model’s context on every request, before it does any work. Anthropic’s engineering write-up on code execution with MCP (November 2025) puts numbers on it: agents connected to thousands of tools process hundreds of thousands of tokens before reading the request, and rewriting one Google Drive → Salesforce workflow so the agent wrote code against its tools instead of calling them one at a time took that task from 150,000 tokens to 2,000 — a 98.7% reduction.
So a fat catalogue is not free surface area; it is a bill and an accuracy problem at once. This is the same discipline as context engineering for AI agents: what’s in the window right now decides how good the answer is.
Version tool schemas like a public API
Once a second team’s agent depends on find_orders, its name and schema are a contract. Renaming a parameter breaks workflows you can’t see. FastMCP’s tool decorator takes a version identifier (documented from v3.0.0), and the convention that survives contact with reality is simple: add new tools rather than changing old ones, deprecate before deleting, and announce schema changes to the teams whose agents call them.
When you don’t need an MCP server at all
If the process is deterministic and scheduled — the same transform, on the same data, every night — a cron job and an integration are cheaper, faster and more auditable than any agent holding a tool. An MCP server earns its place when the questions are open-ended and the data is in a system a model otherwise can’t reach. That test is the same one in our five-question checklist for automating a process, applied to the interface rather than the work.
Havoric is an AI automation and web development agency: we automate repetitive manual processes and build the web and mobile apps around them. These days every internal tool we ship through our AI automation work gets an MCP server next to its web interface — small, read-biased, and behind the same auth as everything else.
Common questions
- How do you build an MCP server for an internal tool?
- Wrap the API the tool already has in a handful of typed functions, give each one a docstring written for a model rather than a developer, and serve them over the Model Context Protocol — roughly 30 lines of Python with FastMCP. The work that actually takes time is choosing which five operations to expose, authenticating the server, and deciding what it is allowed to change.
- How many tools should an MCP server expose?
- As few as will do the job. Every tool definition sits in the model’s context on every request, before any work happens, so a sprawling catalogue costs money and accuracy at the same time. Five well-named tools that match what people actually ask for beat fifty endpoint wrappers.