MCPtrove interactive resource
MCP Tool Schema Token Estimator — Measure tools/list Payload Size
Find oversized tool descriptions and schemas with a formula you can inspect and cite.
The reusable resource
A transparent 3–5 characters-per-token range plus per-tool serialized size and schema review notes.
MCP Tool Schema Token Estimator measures the serialized size of a tools/list payload and provides transparent character-based token ranges without uploading tool definitions.
Contents
- What the tool does
- How to use it
- How the analysis works
- How to interpret the output
- Limitations and privacy
- Practical next steps
- Related MCPtrove resources
- FAQ
What the tool does
An MCP server exposes tools through a tools/list response. Each definition can contain a name, description, and inputSchema describing accepted arguments. Across several connected servers, these definitions can become a substantial part of the context sent to an MCP client or model.
The MCP Tool Schema Token Estimator measures that definition layer before you change a server or tool stack. Paste supported JSON, and it returns aggregate measurements plus a per-tool breakdown.
The report includes:
- Number of tools detected.
- Minified UTF-8 byte size.
- Character count for the formatted representation.
- Total characters in tool descriptions.
- Total characters in JSON Schemas.
- Serialized character size for each tool.
- A character-based estimated token range.
- Review notes for missing
name,description, orinputSchemafields.
The estimator measures definition size only. It does not determine whether a tool is useful, whether a model will select it correctly, or what a particular client will spend at runtime. A large schema may be necessary for a complex operation, while a compact definition may still be ambiguous.
How to use it
Provide the JSON you want to inspect. The estimator accepts a complete tools/list result object, a bare array of tool objects, or any object containing a tools array.
A valid payload shape is:
{
"tools": [
{
"name": "lookup_item",
"description": "Find an item by identifier.",
"inputSchema": {
"type": "object",
"properties": {
"id": {
"type": "string"
}
},
"required": ["id"]
}
}
]
}
Paste the JSON into the estimator; the result updates immediately. If the text cannot be parsed, the tool returns a parse error. Check for unmatched braces, missing commas, invalid quotation marks, or comments inside JSON before trying again.
After a successful run, start with the summary. It shows the total tool count and aggregate size. Then inspect the per-tool table, which is sorted by serialized character count. That ordering helps you find definitions that are unusually large within the same payload.
You can copy the report or download it as Markdown. This supports configuration reviews, team discussions, and before-and-after records when you refine a tool stack.
How the analysis works
The estimator separates several measurements that answer different questions.
It serializes the supplied tool definitions in minified form and reports the resulting UTF-8 byte count. It also reports the character count of the formatted representation, including indentation and line breaks. The two values describe different views of the same payload: encoded compact size and readable text size.
The report separately totals description characters and JSON Schema characters. Description characters show how much prose the definitions contribute. Schema characters show how much structure is devoted to parameters, properties, constraints, and nested objects. Per-tool serialized characters provide a consistent comparison across definitions.
The displayed token estimate uses this formula:
estimated tokens = characters ÷ 5 through characters ÷ 3
This range assumes approximately three to five characters per token. It is not produced by a model-specific tokenizer and should not be treated as an exact token count.
The estimator also checks expected metadata fields. If a tool lacks a name, description, or input schema, it adds a review note. The omission is not automatically treated as a fatal error because the report is intended to describe the payload you supplied.
How to interpret the output
Use the measurements to decide where to investigate, not to produce an automatic quality verdict.
| Metric | What it tells you | Useful question |
|---|---|---|
| Tool count | Number of definitions | Is this the complete stack I intended to review? |
| Minified UTF-8 bytes | Compact encoded size | How large is the serialized JSON? |
| Formatted characters | Readable text size | How much visible content does the payload contain? |
| Description characters | Space used by prose | Are descriptions concise and distinct? |
| JSON Schema characters | Space used by parameters | Which schemas need closer review? |
| Per-tool characters | Relative definition size | Is one tool much larger than its peers? |
| Estimated token range | Approximate context footprint | What range is suitable for rough planning? |
A definition with an unusually large count may contain long prose, many properties, nested objects, or extensive constraints. Review whether that detail is necessary, understandable, and aligned with the tool’s purpose. A large description may provide important context, but it may also repeat information or combine unrelated responsibilities. A large schema may reflect legitimate complexity or indicate that an operation could be narrower.
Missing-field notes deserve similar attention. A missing description can make a tool harder to distinguish during review. A missing schema can make the input contract less explicit. These notes identify metadata to confirm; they do not prove that a tool is unusable or incorrectly implemented.
Limitations and privacy
Tokenization varies by model, vocabulary, Unicode content, wrapper messages, and client behavior. The same JSON can therefore produce different token counts in different environments. The estimator’s character-based range is a planning aid, not a precise billing figure or exact context calculation.
The report measures tool definitions, not the complete request lifecycle. It does not include every message a client may add around a tools/list response, later tool calls, returned results, conversation history, or client-specific processing.
Payload size does not establish tool quality, selection accuracy, security, or runtime cost. An unusually large definition is a review signal, not proof of compromise or malicious behavior. A parse error means the estimator could not parse the supplied JSON; it does not, by itself, diagnose the system that produced the text.
The tool is designed to provide its analysis without uploading tool definitions. Nevertheless, inspect the content before placing it into any workflow. Descriptions and schemas can contain internal names, URLs, operational details, or other information your organization restricts. Do not include secrets, credentials, or private values in tool metadata, and keep downloaded reports within the access boundaries appropriate for the source configuration.
Practical next steps
Use the estimator when evaluating a new server, reviewing an existing stack, or investigating why a collection of definitions is difficult to manage.
A focused review sequence is:
- Capture the relevant
tools/listdata. - Run it through the estimator.
- Record the aggregate measurements and estimated range.
- Review the largest individual definitions first.
- Check descriptions and schemas for repetition, unnecessary breadth, or unclear boundaries.
- Confirm any missing-field notes.
- Make changes only when the tool’s purpose and contract support them.
- Re-run the estimator to document the result.
The goal is not to minimize every character. It is to keep the tool surface understandable, appropriately scoped, and easier to reason about.
Related MCPtrove resources
For an actual MCP configuration review, use MCP Config Doctor. It is the better next step when you need to examine a working stack rather than one pasted payload.
For broader directory context, see the MCP Tool Count Benchmark. It places tool counts alongside other directory-level observations, while this estimator remains focused on serialized definitions.
If a review reveals more breadth than your use case requires, browse Best MCP Servers for focused alternatives. Choose based on the capabilities and contract you need, not size alone.
FAQ
Does the estimator return an exact token count?
No. It returns a transparent range based on approximately three to five characters per token. Exact tokenization depends on the model, tokenizer, Unicode content, wrapper messages, and client behavior.
What input formats can I paste?
You can provide a tools/list result object with a tools array, a bare array of tool objects, or another object containing a tools array. Invalid JSON produces a parse error.
Why is one tool much larger than the others?
It may contain longer descriptions, more properties, nested objects, or additional constraints. Treat the per-tool size as a review signal and decide whether that detail is necessary.
Are missing fields treated as errors?
Missing name, description, or inputSchema fields are flagged as review notes rather than automatically treated as fatal errors. Confirm whether the omission is intentional.
Does payload size tell me which tools a model will choose?
No. Size does not establish selection accuracy, tool quality, or runtime cost. It describes serialized definitions and supplies a rough character-based token range.