Grounded Reasoning Systems for Cloud Architecture - Iman Makaremi

AI EngineerAbout 4 min readJun 4, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Grounded reasoning systems, cloud architecture, multi-agent orchestration, AI copilot, semantic enrichment, graph-enhanced search, requirement understanding, architecture identification, architecture recommendation, complex reasoning scenarios, semantic context, graph context, role-specific agents, structured message format, memory management, human evaluation, relevance, visibility, clarity.

1. Introduction: The Need for Reasoning in Cloud Architecture

  • Main Point: Cloud architecture is becoming increasingly complex, requiring reasoning capabilities beyond simple automation.
  • Details: The complexity stems from the growing number of users, developers, tools, constraints, and rising expectations.
  • Argument: Traditional automation tools are insufficient to handle the diversity of decisions in cloud architecture. Systems need to understand, debate, justify, and plan.
  • Quote: "Cloud architecture needs reasoning, not just automation."

2. Challenges in Applying AI to Architecture Design

  • Main Point: Applying AI to architecture design involves three key challenges.
  • Challenge 1: Requirement Understanding:
    • Details: Understanding the source, format, scope (global vs. specific), and importance of requirements.
  • Challenge 2: Architecture Identification:
    • Details: Understanding the function of different components within an architecture and their interrelationships.
  • Challenge 3: Architecture Recommendation:
    • Details: Providing recommendations to match requirements or improve the architecture based on best practices.
  • AI-Related Challenges:
    • Semantic and Graph Context: Bridging the gap between textual requirements and graph-based architecture data.
    • Complex Reasoning Scenarios: Handling vague and broad questions by breaking them down and planning solutions.
    • Evaluation and Feedback: Developing methods to evaluate and provide feedback to large AI systems with many moving parts.

3. Grounding Agents in Specific Context

  • Main Point: AI agents need proper context about architecture to reason effectively.
  • Challenge: Translating natural language into meaningful architecture retrieval tasks is difficult, especially at speed.
  • Approaches:
    • Semantic Enrichment of Architecture Data: Adding relevant semantic information to each component to improve searchability in vector search.
    • Graph-Enhanced Component Search: Using graph algorithms to retrieve relevant information from the architecture graph.
    • Early Score Enrichment of Requirement Documents: Scoring requirement documents based on important concepts to speed up retrieval.
  • Learnings:
    • Semantic grounding improves reasoning but has limitations.
    • Design is critical in soft grounding (telling the agent what to focus on).
    • Graph memory supports continuity (understanding connections) and not just accuracy.
  • Example: Initial design for architecture retrieval involved breaking down architecture into components, converting JSON data to natural language, enriching with connection data, embedding in vector DB. Later shifted towards graph-based searches.
  • Example: Initial design for requirement understanding involved requirement templates with specific structure to extract relevant information.

4. Complex Reasoning Scenarios and Multi-Agent Orchestration

  • Main Point: Architecture design involves conflicting goals, trade-offs, and debates, requiring agents that can collaborate, argue, and converge on justified recommendations.
  • Approach: Building a multi-agent orchestration system with role-specific agents.
  • Key Elements:
    • Structured message format for communication between agents.
    • Conversation management to isolate conversations and manage memory.
    • Cloning of agents for parallel processing.
  • Learnings:
    • Structured outputs improve clarity and control.
    • Multi-agent systems can resolve trade-offs dynamically.
    • Successful orchestration requires control flow.

5. Multi-Agent System for Architecture Recommendation

  • Main Point: A multi-agent system is used to generate architecture recommendations.
  • System Components:
    • Chief Architect: Oversees and coordinates higher-level tasks.
    • Staff Architects (10): Specialized in specific domains (infrastructure, API, IAM, etc.).
    • Requirement Retriever: Accesses requirement data.
    • Architecture Retriever: Understands the current architecture state.
  • Workflow:
    1. List Generation: Generate a list of possible recommendations.
    2. Conflict Resolution: Chief Architect resolves conflicts and redundancies in the list.
    3. Design Proposal: Generate a full design proposal for each recommendation.
  • Process Details:
    • Chief Architect requests recommendations from Staff Architects.
    • Staff Architects request information from Requirement and Architecture Retrievers in parallel.
    • Staff Architects provide a list of recommendations to the Chief Architect.
    • After conflict resolution, Staff Architects are cloned to generate design proposals in parallel, with each clone having access to the past history.

6. Evaluation and Feedback

  • Main Point: Closing the loop with human scoring and structured feedback is crucial for improving the system.
  • Challenge: Determining if a recommendation is good in a complex multi-agent system.
  • Approach: Building an internal human evaluation tool ("Eagle Eye") to review cases, architecture, requirements, agent conversations, and generated recommendations.
  • Evaluation Metrics: Relevance, visibility, clarity.
  • Learnings:
    • Confidence is not correctness.
    • Human feedback is essential early on.
    • Evaluation must be baked into system design from the start.
  • Example: The evaluation tool helped identify cases of hallucination, such as a Staff Architect scheduling a workshop with specific dates.

7. Conclusion: Reasoning Systems as AI Copilots

  • Main Point: Building an AI copilot is about designing a system that can reason, not just generate answers.
  • Key Requirements:
    • Ability to process large amounts of data (architecture components, documents).
    • Support for various stakeholders (developers, CTOs).
    • Roles, workflows, memories, and structure.
  • Future Directions:
    • Experimentation to identify patterns that work best based on data.
    • Increasing importance of graphs in designs.
    • Optimizing agent interactions and autonomy.
    • Using LangGraph for building agent workflows and Slide for higher-level management.
  • Quote: "We believe this is how AI will design software."

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.