I Graded 36 Popular MCP Servers: A Third Fail Agent Usability
I graded 36 popular MCP servers on agent usability. A third got a D or F

I built mcpgrade to test 36 popular MCP servers and found that a third fail basic agent usability checks. Despite passing protocol compliance, many official servers from MongoDB, Notion, and GitHub lack critical parameter descriptions, causing models to hallucinate arguments or refuse tasks. The data reveals that agent reliability depends less on engineering specs and more on disciplined documentation practices that current tools ignore.
Your schema generator is quietly stripping the single most important signal your tools have.
- brookst
> well-documented big catalogs are possible; they're just rare
They’re not that rare, it’s just that flat big catalogs don’t work well.
The most common failure mode for MCP server design is mapping MCP tools to system APIs. This results in a tool to list files, a tool to rename a file, a tool to copy a file, etc.
You can have 100 “tools”, you just need to define 10-12 areas (actual tools in the MCP model) and actions within each one. It creates a hierarchy that models can dive into when needed and avoid context clutter when that area is not needed.
So instead of 15 tools for file operations you have a single “file” tool, with actions like list, copy, rename, etc.
Single biggest arch requirement for any MCP server (well, any that has more than a handful of operations).
- codeonline
Somewhat ironically the link to your results table is broken
- lolive
JSON and CSV [and a massive amount of custom-made spaghetti code] gave the impression that semantics was totally unnecessary in dataflows. Now that we live in a world where dataflows build themselves, we will need a description language that is much more robust.
Probably generated from chatbot interactions, but formally defining how the consumer can discover content and mix them.