Anthropic Just Replaced Claude Code With New Claude Tag

By AI Revolution

Share:

Key Concepts

  • Claude Tag: An enterprise-focused, team-based AI agent system integrated into Slack.
  • Extended Thinking: A feature where models perform internal reasoning; currently, users receive a summary rather than the raw chain of thought.
  • Cryptographic Reasoning Blocks: Encrypted JSON blobs containing model state, which researchers found to be vulnerable to replay attacks and side-channel analysis.
  • Fugu (Sakana AI): A "dispatcher" or routing model that intelligently delegates tasks to the best-performing frontier models (Claude, GPT, Gemini).
  • Cross-Domain Alignment: OpenAI research suggesting that beneficial behaviors (e.g., truthfulness, correctability) can be generalized across unrelated domains through Reinforcement Learning (RL).

1. Anthropic’s Claude Tag

Anthropic has launched Claude Tag, an evolution of Claude Code designed for organizational infrastructure.

  • Functionality: Operates within Slack, allowing teams to collaborate in public channels. It integrates with GitHub, Jira, Linear, and CRM systems.
  • Key Features:
    • Shared Context: Tasks are visible to the entire team, preventing the need for repetitive explanations.
    • Ambient Mode: Proactively surfaces stalled discussions or unresolved decisions.
    • Asynchronous Execution: Allows the agent to run long-term tasks in the background.
    • Claude Identities: Isolated instances per team with specific token budgets and audit logs.
  • Strategic Goal: To capture "tacit organizational knowledge" by embedding AI directly into team workflows.

2. The "Thinking" Transparency Controversy

A significant debate has emerged regarding how Anthropic and OpenAI handle "thinking" processes.

  • The Discovery: Developer Patrick McKenna found that Claude Code logs contained empty "thinking" blocks replaced by 600-character cryptographic signatures.
  • The Reality: Anthropic provides a summary of the reasoning, not the actual chain of thought. The raw reasoning is encrypted, and Anthropic holds the key.
  • Cryptographic Vulnerabilities: Professor Matt Green (Johns Hopkins) analyzed these reasoning blocks and found:
    • Global Key Usage: Both companies appear to use a single global encryption key for all users, allowing for "replay attacks" where a reasoning block from one account can be injected into another.
    • Side-Channel Leaks: By observing the length of reasoning blocks or wall-clock response times, an attacker can reconstruct hidden data (e.g., bits of a secret) even without decrypting the content.
  • Implications: These findings suggest that "zero data retention" modes may be less secure than advertised, as reasoning state is effectively escrowed under a single, non-rotated key.

3. Sakana AI’s Fugu Model

Sakana AI released Fugu, a routing model that does not possess its own base weights but acts as an intelligent dispatcher.

  • Methodology: It classifies incoming requests (e.g., math, coding, research) and routes them to the most capable model (Claude Opus, GPT, or Gemini).
  • Performance: Benchmarks show Fugu Ultra outperforming individual models, though this is largely due to the "aggregated ceiling" of the underlying frontier models it accesses.
  • Business Risk: The model relies on the cooperation of the very companies it competes with (OpenAI, Anthropic, Google), creating a precarious long-term business model.

4. OpenAI Research: Persistent Alignment

OpenAI published a paper on using Reinforcement Learning (RL) to instill "broadly and persistently beneficial" traits.

  • The Problem: "Emergent misalignment," where training a model to be safe in one domain (e.g., coding) fails to prevent deceptive behavior in another (e.g., general conversation).
  • The Experiment: Researchers trained a model on a dataset containing 15 beneficial traits (e.g., truthfulness, risk-aware planning) across 12 domains.
  • Key Finding: Cross-domain transfer. A model trained on beneficial behaviors in health contexts showed improved alignment in non-health domains (17 out of 19 evaluations).
  • Significance: This suggests that alignment can be scaled by teaching general "beneficial traits" rather than patching individual failure modes one by one.

Synthesis and Conclusion

The AI landscape is currently defined by a tension between utility and transparency. While companies like Anthropic are successfully embedding AI into the fabric of enterprise collaboration (Claude Tag), they are simultaneously obscuring the "reasoning" processes of these models behind opaque cryptographic layers. The research by Matt Green highlights that these "black box" reasoning implementations currently suffer from significant security flaws, such as global key usage and side-channel leakage. Meanwhile, the emergence of routing models like Fugu signals a shift toward "model-agnostic" workflows, and OpenAI’s latest research offers a promising, scalable path toward making AI models inherently more aligned across diverse, real-world applications.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video