Spanner: The always-on, virtually unlimited scale database

By Google Cloud Tech

Share:

Key Concepts

  • Multimodal Database: A database that supports multiple data models (relational, key-value, graph, vector) within a single system.
  • Globally Consistent: Ensures that data is consistent across all replicas and regions, regardless of where it is accessed.
  • Virtually Unlimited Scale: The ability to handle massive amounts of data and query loads without performance degradation.
  • Generative AI (Gen AI): Artificial intelligence that can create new content, such as text, images, or code.
  • Vector Search: A search technique that uses vector embeddings to find semantically similar data.
  • Graph Query Language (GQL): A standard language for querying graph databases.
  • ACID Transactions: Atomicity, Consistency, Isolation, and Durability – properties that guarantee reliable transaction processing.
  • Synchronous Replication: Data is written to multiple replicas simultaneously to ensure consistency.
  • Geo Partitioning: Distributing data at the row level across different geographical regions.

Spanner: A Next-Generation Multimodal Database

In today's demanding environment, data is a crucial differentiator for understanding users and tailoring products. Applications must handle explosive data growth, scale on demand, comprehend the semantic meaning of data for generative AI, and consistently deliver exceptional user experiences, all while adhering to complex data regulatory requirements. Traditional databases often struggle to meet these demands, which is where Spanner emerges as a solution.

Spanner is presented as an "always on" globally consistent multimodal database with virtually unlimited scale. It's built on Google's unique infrastructure, aiming to eliminate the complexities of traditional databases and allow developers to focus on application innovation rather than infrastructure management.

Multimodal Capabilities

A key differentiator of Spanner is its multimodal nature, unifying relational, key-value, graph, and vector search workloads within a single database. This unification enables the development of "truly intelligent applications" by:

  • Unlocking diverse data models: Developers can leverage different ways to structure and access their data.
  • Understanding intricate relationships: The database facilitates the comprehension of complex connections within the data.
  • Leveraging advanced search technologies: Integrated search capabilities enhance data discovery.

This unified approach offers several benefits:

  • Data consistency across all models: Ensures a single source of truth.
  • Reduced operational complexity: Simplifies management by consolidating workloads.
  • Shrunk attack surface: Enhances security by reducing the number of systems to manage.

Spanner Graph: Unlocking Data Relationships

Spanner Graph is highlighted for its ability to leverage data relationships and provide accurate, relevant answers at scale. It features:

  • Graph query interface: Compatible with the ISO GQL standard.
  • Combined SQL and GQL: Integrates established SQL capabilities with the expressiveness of GQL for graph pattern matching.
  • Graph-optimized storage and query enhancement: Designed for high performance with graph data.
  • High performance at scale: Maintains efficiency even with vast amounts of graph data.

Built-in Vector Search for Gen AI

Spanner incorporates built-in vector search, eliminating the need for separate vector databases. This capability is crucial for powering Gen AI applications and enables features like:

  • Searching vector embeddings using natural language: Users can query for semantically similar data through intuitive language.
  • Eliminating cost and complexity: Reduces the overhead of managing an additional database system.

Integrated Full-Text Search

Drawing from Google's extensive search experience, Spanner's full-text search goes beyond exact matches and structured fields. It allows for intelligent searching of words, phrases, and numbers across Spanner data using:

  • Machine learning-based processing: Determines the intent behind search queries for more accurate results.

Seamless Integration of Capabilities

The power of Spanner lies in the seamless integration of its various capabilities. For example, a workflow could involve:

  1. Using Spanner's vector and full-text search to identify relevant nodes and edges in the data.
  2. Subsequently, using GQL to explore the connections within the identified graph data.

Robust Relational and Key-Value Support

Beyond its advanced features, Spanner retains strong foundational capabilities:

  • Powerful relational capabilities: Supports traditional structured data.
  • ACID transactions: Guarantees reliable data operations.
  • Strongly consistent secondary indexes: Enhances query performance for structured data.
  • Key-value workloads: Handles simple key-value data storage with ease.

Engineering Excellence and Scalability

Spanner's foundation is built on "solid engineering and unparalleled innovation," offering:

  • Virtually unlimited scalability: The ability to grow with data needs.
  • Industry-leading availability: High uptime guarantees.
  • Strong consistency at global scale: Data integrity across all regions.

Elasticity is a key feature, allowing users to "start small, dream big, and scale effortlessly" without application modifications. This simplifies architecture and operations by automatically handling:

  • Replicas: Copies of data for availability and performance.
  • Sharding: Partitioning data across multiple nodes.
  • Transaction processing: Managing data updates across distributed systems.

Spanner's scalability is demonstrated by its ability to process over 6 billion queries per second at peak and manage more than 17 exabytes of data.

High Availability and Strong Consistency

Spanner eliminates the trade-off between high availability and strong consistency. It offers up to 99.999% availability through synchronous replication between multiple replicas, which can be:

  • Regional: Within a single geographical region.
  • Dual-region or multi-region: Spread across different geographical locations for enhanced resilience and performance.

Geo partitioning further enhances performance by allowing data to be partitioned at the row level across the globe, serving data closer to users for lower latency. Crucially, even with geo partitioning, Spanner maintains all distributed data as a single cohesive table for queries and mutations, enabling compliance with local data regulatory requirements while maintaining global data access.

Price Performance and Cost Management

Spanner also focuses on excellent price performance through features like:

  • Managed autoscaling: Automatically adjusts resources based on demand.
  • Incremental backups: Efficiently backs up only changed data.
  • Tiered storage: Optimizes storage costs by using different storage classes.

Conclusion and Next Steps

In summary, Spanner provides a multimodal experience, industry-leading high availability, strong consistency, and increased operational efficiencies. It empowers developers to build the next generation of intelligent applications while optimizing costs. Spanner is presented not just as a solution for current database challenges but as a foundation for future innovation.

To learn more, users are directed to cloud.google.com/spanner. Spanner also facilitates easy workload migration from databases like MySQL, PostgreSQL, and Cassandra through comprehensive tools. Interested users can get started with a 90-day free trial instance and the Spanner emulator.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video