I graded the official MCP servers. Two got an A, one got an F — and the most useful finding came from a layer nobody else measures.
mcp-vitals (open source) grades an MCP server's reliability A–F across three layers: static schema quality, behavioral tests run against the live server, and — the novel part — agent-usability: hand an LLM the tool list and a real task, and see if it picks the right tool and builds a valid call.
Results on the official reference servers:
• memory — A (94)
• everything — A (92)
• filesystem — C (71)
• sequential-thinking — F (57)*
The finding that matters came from the agent-usability layer. On the filesystem server, a model given a "read this file" task couldn't tell read_file from read_text_file and picked the tool that doesn't exist. A security scanner or a static grader never sees that — but it silently breaks agents in production.
*The F is honest: sequential-thinking's single tool needs a payload my schema-only generator can't synthesize, so the behavioral layer couldn't exercise it. A conservative grader that won't vouch for what it can't verify — I document that plainly in the report.
Security scanners tell you a server is safe. mcp-vitals tells you if it's any good.
Use AI for the boring 60% of your MVP: boilerplate, CRUD screens, forms, API wiring. Save your own brain for the 40% that decides if it works: the core loop, edge cases and what NOT to build:)
Saw this - been testing Seedance 2.5 too. Output is good but still needs human clean-up, esp label warp. I disclose it as AI-assisted. Human eye still matters.
Is AI finally good enough to shoot a product demo on its own, or human touch is still needed?
I generated a 20s product demo with Seedance 2.5 in a single take providing it with one product image reference and an example video transcript (posted in the comments).
Do you disclose...