Overview of Web IQ

By John Savill's Technical Training

Share:

Key Concepts

  • Web IQ: A service designed to provide AI applications with real-time, external, non-organizational data from the internet.
  • Retrieval Augmented Generation (RAG): A framework where external knowledge is retrieved and appended to a prompt to improve the accuracy and relevance of an AI model's output.
  • Tokenomics: The economic impact of token usage in LLMs; minimizing input tokens is critical for cost reduction and performance.
  • Grounding: The process of anchoring an AI’s response in verifiable, external data to prevent hallucinations.
  • Latency: The time delay between a request and a response; critical for AI agents that perform iterative lookups.
  • Passages: A feature that extracts only the most relevant segments of a webpage rather than the entire document, optimizing token usage.

1. The Problem: Knowledge Gaps in AI Models

Generative AI models are limited by their training process. The speaker identifies three primary "knowledge gaps":

  • Finite Knowledge: Models cannot store the entirety of human knowledge; they are limited by their parameter count.
  • Knowledge Cut-off: Models are static snapshots of data at a specific point in time and lack awareness of events occurring after their training.
  • Lack of Organizational Context: Pre-trained models have no access to private, internal data (e.g., meetings, artifacts, or proprietary workflows).

Key Argument: AI is often "mile-wide but inch-deep." Without external grounding, models can be "confidently wrong," leading to hallucinations. As AI evolves from simple chat assistants to complex agents, these errors cascade, eroding user trust.

2. The Solution: Web IQ

Web IQ acts as the bridge between AI applications and the real world. It provides a structured way to fetch current, relevant information to ground AI responses.

Core Features and Benefits:

  • Freshness: Ensures the AI has access to the latest market data, news, and regulations.
  • Low Latency: Boasts a P95 latency of 164ms, which is 2.5x faster than alternatives. This is vital for "iterative loops" where an AI agent performs multiple sequential lookups.
  • Token Efficiency: By using the "Passages" feature, the system returns only the most pertinent paragraph of a site, significantly reducing the number of input tokens sent to the LLM.
  • Security and Risk Management: Provides a safe "Browse" capability. Instead of an AI agent visiting potentially malicious URLs directly, Web IQ fetches and sanitizes the content, protecting the application from prompt injection or web-based attacks.
  • Responsible AI Grounding: Allows publishers to specify if their content can be used for AI grounding, helping businesses avoid IP infringement and legal issues.

3. Methodologies and Frameworks

The speaker categorizes organizational knowledge into three existing "IQ" components, with Web IQ serving as the fourth pillar for external data:

  1. Work IQ: Context on people, collaborations, and workflows.
  2. Fabric IQ: Context on business entities and systems of record.
  3. Foundry IQ: Context on policies and authoritative documents.
  4. Web IQ: Context on the internet (news, images, video, and general web data).

The RAG Workflow:

  1. Prompt: The user submits a query.
  2. Retrieval: The AI application calls the Web IQ API to fetch relevant, current data.
  3. Augmentation: The retrieved data (e.g., a specific passage) is appended to the prompt.
  4. Generation: The model uses the augmented prompt to generate a grounded, accurate response.

4. Technical Demonstration

The speaker demonstrated the flexibility of the Web IQ API:

  • Content Formats: Data can be retrieved as HTML, plain text, or markdown, depending on the application's needs.
  • Search Modes:
    • Web Search: General information retrieval.
    • News: Real-time updates on current events.
    • Video/Images: Returns rich metadata (captions, duration, source, views).
    • Browse: Securely extracts content from a specific URL provided by the user.

5. Synthesis and Conclusion

The primary takeaway is that token efficiency is the key to scalable AI. While Web IQ may have a cost per call, it drastically reduces the total cost of ownership by minimizing the number of tokens sent to the LLM. By providing high-quality, relevant, and secure data, Web IQ enables developers to build AI agents that are not only faster and cheaper but also more trustworthy and accurate.

Note: Web IQ is currently in limited access General Availability (GA); users are encouraged to contact their account teams for access.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video