MCP Directory

MCPtrove original research

MCP Tool Count Benchmark 2026: The Median Server Exposes 10 Tools

A tool-count benchmark for planning MCP portfolios: median, quartiles, long-tail servers and the limits of using count as a quality signal.

MCPtrove·September 19, 2026·6 min read· Download raw CSV

The citable finding

The median server exposes 10 tools, the average is 13.2, and the 90th percentile is 26.

Across 520 MCP server listings, the median server exposes 10 tools, the average is 13.2, and the 90th percentile is 26.

Contents

Methodology and limitations

This benchmark measures the normalized tool count attached to each listing in MCPtrove’s 520-listing directory snapshot. The dataset contains 6,882 tools across 520 listings, using the directory’s tool_count field as the measurement unit.

A listing is a directory record, not necessarily a one-to-one representation of a deployed process. One listing can expose many tools, and one user configuration can combine several listings. That distinction matters: the useful unit for reviewing an MCP setup is the active tool portfolio, not simply the number of servers installed.

The category, authentication, and transport labels used alongside this benchmark reflect MCPtrove’s normalized directory fields. They are useful for consistent comparison, but they are not claims that every upstream project uses identical terminology or packaging.

Percentiles use nearest-rank-style positions on the sorted 520-row tool_count field, as generated by MCPtrove. In plain language, MCPtrove sorts the listings by normalized tool count and selects the value at each requested percentile position. The calculation does not smooth values between listings or estimate an unseen population.

This is a descriptive benchmark of the directory snapshot. It is not a survey of MCP users, a census of every MCP server in existence, or a performance test. Tool count does not tell us whether tools are reliable, frequently used, well documented, fast, safe, or valuable. It also does not establish how many tools are simultaneously active in a particular client configuration.

The benchmark should therefore be read as a practical map of directory breadth. It shows how tool portfolios are distributed among the listings MCPtrove captured, while leaving runtime behavior and user experience to separate analysis.

Main findings

The distribution is broad but concentrated around compact portfolios. The 25th percentile is 5 tools, the median is 10, and the 75th percentile is 16. Half of the listings therefore sit between the lower and upper quartiles, while the upper tail stretches substantially beyond that central range.

The average is higher than the median: 13.2 tools per listing. That gap is consistent with a right-skewed distribution, where a smaller group of large portfolios pulls the average upward. The 90th percentile is 26 tools, and the 95th percentile is 39 tools.

The snapshot includes a small number of unusually expansive listings. The maximum normalized count is 94 tools, while 25 listings exceed 40 tools. MCPtrove uses 40 tools as a practical review threshold: it marks the point where a listing becomes unusual enough to deserve deliberate inspection, not a protocol or client maximum.

The highest normalized counts in the snapshot are:

ListingNormalized tool count
Aseprite MCP Tools94
Salesforce DX MCP Server89
Bitrise MCP Server81
Safari MCP79
Crypto Indicators MCP Server78
JavaLens75
Discord MCP73
figma-mcp-go73
PagerDuty MCP71
Confluent MCP Server71

The presence of 4 listings with zero normalized tools is also a useful reminder that directory metadata can represent different stages of listing completeness. A zero in this field should not automatically be interpreted as proof that an implementation has no callable capability; it means MCPtrove recorded zero normalized tools for that listing in this snapshot.

Interpretation

The editorial takeaway is straightforward: compact servers are typical, but configuration complexity can arrive through combination.

A median of 10 tools does not mean a user experiences only that many tools. A configuration with several ordinary-sized listings can create a much larger active portfolio. Similar capabilities may appear under different names, while a broad server may include tools that are irrelevant to the task at hand. The resulting noise can affect discoverability, prompt selection, review effort, and the likelihood that a user understands what an action-capable tool can do.

That is why server count is a weak proxy for operational complexity. The better question is: how many tools are active, overlapping, permissioned, and available to the model at the same time?

The benchmark also argues against simplistic limit narratives. MCPtrove does not repeat the unsupported claim that Cursor has a universal hard 40-tool limit. The number is useful here only as a review threshold chosen from this directory distribution. A threshold can guide inspection without pretending to be a protocol rule, a client guarantee, or a causal explanation of failures.

Most importantly, this dataset does not show correlation as causation. It does not prove that larger listings cause slower responses, worse tool selection, higher error rates, or reduced reliability. Those questions require runtime measurements and controlled comparisons. The benchmark tells us where to look, not what will happen in every environment.

Practical decision guidance

Use the benchmark as a triage tool.

If a listing sits near the center of the distribution, inspect its tool names, descriptions, authentication requirements, and transport before deciding whether it belongs in a configuration. A moderate count is not automatically good or bad. Relevance and clarity matter more than raw breadth.

If a listing crosses 40 tools, treat it as a prompt for review. Group the tools by task, remove capabilities you do not need, check for overlapping functions, and confirm that the authentication scope matches the intended workflow. The threshold is deliberately practical: it identifies uncommon breadth in this snapshot.

When several listings are combined, review the resulting portfolio rather than evaluating each server in isolation. Look for duplicate search, browsing, file, database, or project-management capabilities. Compare names and descriptions, then keep the smallest set that supports the work. Fewer active tools can make a configuration easier to understand, audit, and maintain, even when no formal client limit is involved.

For a hands-on next step, paste your configuration into the MCPtrove configuration doctor. Then use the MCPtrove configuration generator to assemble a smaller, more intentional set of servers and tools.

How to cite and download the data

Download the complete raw dataset from mcp-tool-count-benchmark-2026.csv. The CSV is the source for the original figures on this page, including the distribution statistics, zero-tool listings, threshold count, and highest-count listings.

When reusing a figure, link the number directly to the CSV where practical. Keep the scope attached to the claim: these results describe MCPtrove’s directory snapshot, not the entire MCP ecosystem.

MCPtrove, “MCP Tool Count Benchmark 2026,” https://mcptrove.com/research/mcp-tool-count-benchmark, accessed September 19, 2026.

For the client-limit context behind the review threshold, read Cursor Tool Limit Math.

For adjacent directory views, compare MCP Server Language Statistics and MCP Transport Statistics. Together, these pages help separate tool breadth from implementation language, transport choice, and other normalized directory characteristics.

FAQ

Does the median mean most MCP servers expose the same number of tools?

No. The median is the middle value after sorting the 520 listings. It means half of the observed listings are at or below 10 tools, while the other half are at or above that value. It does not mean that 10 tools is the most common exact count.

Is the review threshold a hard MCP or Cursor limit?

No. 40 tools is MCPtrove’s practical review threshold because 25 listings exceed it in this snapshot. It is not presented as a protocol maximum, a universal client limit, or a guarantee that a configuration below it will behave well.

Does one listing equal one tool?

No. One listing can expose many tools. This benchmark counts the normalized tools recorded for each listing, which is why the active tool portfolio is more informative than server count alone.

Can the benchmark predict runtime performance?

No. Tool count is a breadth measure, not a latency, reliability, safety, or quality score. The benchmark does not establish causation between a larger portfolio and any particular runtime outcome.

Why are some listings recorded with zero tools?

The dataset contains 4 listings with zero normalized tools. That value reflects MCPtrove’s normalized directory field for this snapshot. It should be treated as metadata about the listing record, not as a universal statement about the underlying implementation.

Use the data, then inspect your own stack

Download the source rows for analysis, or check whether your active MCP configuration is focused enough for reliable tool selection.

Related MCPtrove research