Bending a Public MCP Server Without Breaking It — Nimrod Hauser, Baz
By AI Engineer
Key Concepts
- MCP (Model Context Protocol) Server: A standard for connecting AI agents to external tools and data sources.
- Agentic Tools: Callable functions wrapped in natural language descriptions that allow LLMs to interact with external systems.
- Context Engineering: The practice of refining tool descriptions and selection to optimize LLM performance.
- Deterministic Guardrails: Hard-coded logic that prevents agents from performing unauthorized or unsafe actions.
- Multimodal Agents: AI systems capable of processing both text (tickets) and visual data (Figma designs/screenshots).
1. The Challenge of Third-Party Tools
The speaker, Nimrod Hower, highlights that while third-party tools (like the Playwright MCP server) are powerful, they are often too generic for specific enterprise use cases. Using them "out of the box" leads to:
- Unpredictability: Agents may hallucinate or fail to navigate complex systems.
- Performance Degradation: Suboptimal tool selection or incorrect usage.
- Security Risks: In multi-tenant architectures, agents without proper guardrails may inadvertently leak data or access unauthorized schemas.
2. Use Case: The "Spec Reviewer"
Buzz’s "Spec Reviewer" is an agentic workflow designed to:
- Collect Requirements: Read tickets (Jira/Linear) and visual designs (Figma).
- Execute Verification: Use Playwright to navigate the system, check branches, and verify implementation.
- Verdict & Evidence: Provide a pass/fail result with a screenshot as proof.
3. Framework: Five Best Practices for Agentic Tools
To move from a failing system to a robust one, the speaker proposes five methodologies:
I. Curating Tools
Reduce the "noise" in the agent's context window by removing unnecessary tools.
- Method: Use list comprehension to filter out irrelevant functions (e.g., browser resizing or code execution) that the specific use case does not require.
- Result: A smaller, more focused toolset leads to better decision-making by the LLM.
II. Wrapping Tools with Enhanced Descriptions
Generic descriptions (e.g., "press a key") are insufficient.
- Method: Create a
ToolWrapperclass that maps original tool names to custom, verbose descriptions. - Actionable Insight: Provide "chain-of-thought" guidance within the description (e.g., "Before calling
click, first callaccessibility_snapshot").
III. Implementing Deterministic Guardrails
Prevent agents from going "rogue" by enforcing logic outside the LLM's decision-making process.
- Method: Intercept tool calls (e.g.,
take_screenshot) to validate parameters like file paths. - Mechanism: If an agent attempts to save a file outside the designated
screenshots_root, the system raises an exception with a helpful error message, prompting the agent to retry with a valid path.
IV. Composing New Tools
Build specialized tools by combining existing ones.
- Method: Create an "Evidence Tool" that wraps the standard screenshot function but adds specific requirements, such as forcing the inclusion of a ticket number in the filename.
- Benefit: This allows the agent to distinguish between a "generic" screenshot and an "evidence" screenshot, applying different logic to each.
V. Treating Tools as Deterministic Functions
Not every action needs to be agentic.
- Method: Execute complex, repetitive, or sensitive tasks (like logging into a system with JWT tokens) as standard code before the agent starts its flow.
- Benefit: This unburdens the agent from "clunky" tasks, reducing latency and potential failure points.
4. Notable Quotes
- "Agents are already non-deterministic, unpredictable things. You give them tools and you get unpredictability at scale."
- "Sometimes there are aspects of your tasks that are just too sensitive to leave at the hands of the agents."
- "The world is your oyster and knock yourselves out." (Regarding the flexibility of composing new tools).
5. Synthesis and Conclusion
The transition from a failing MCP server to a functional one demonstrates that agentic success is not about the tools themselves, but how they are curated and constrained. By moving from a "vanilla" implementation to a "tailored" one—using filtering, enhanced descriptions, deterministic guardrails, and pre-processing—developers can significantly improve the reliability and security of AI agents. The ultimate takeaway is that developers must act as "architects" of the agent's environment, balancing the flexibility of LLMs with the precision of deterministic code.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

Is there a Chinese cyber threat to EU solar energy? | DW News
DW News

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English

The Bubble is Bursting... (Emergency Update)
Bravos Research

The AI Bubble Just Ended - Without Popping
Heresy Financial