New York Times' Connections: A Case Study on NLP in Word Games — Shafik Quoraishee, NYT Games

AI EngineerAbout 4 min readJul 6, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Connections Game: A word association game by the New York Times where players group 16 words into four categories of four.
  • Difficulty Structure: Yellow (easiest), Green, Blue, Purple (most difficult due to decoys).
  • System 1 & 2 Thinking: System 1 is fast, intuitive thinking; System 2 is slow, deliberate reasoning.
  • Graph Coloring Problem: Assigning colors to vertices of a graph such that no adjacent vertices share the same color.
  • Semantic Similarity: The degree to which two words are related in meaning.
  • Relational Alignment: A metric that quantifies the relationship between two words.
  • Lexical Databases: Resources like WordNet and ConceptNet that provide information about word relationships.
  • Graph Neural Networks (GNNs): Neural networks that operate on graph-structured data.
  • Reinforcement Learning: A type of machine learning where an agent learns to make decisions in an environment to maximize a reward.
  • Hypergraphs: A generalization of graphs where edges can connect more than two vertices.

Connections Game Overview

The New York Times' Connections game involves grouping 16 words into four categories of four. The game has a difficulty structure: Yellow (easiest), Green, Blue, and Purple (most difficult). The purple category often includes decoys and overlaps, making it challenging. The game's mechanics and puzzles are currently human-made.

AI and Connections: Challenges and Opportunities

Connections presents a unique challenge for AI due to its need for abstract reasoning and the presence of intentional decoys. The game can serve as a reproducible and scalable test bed for AI capabilities. While Large Language Models (LLMs) have shown progressive capability in solving Connections, they are not perfect.

Human Problem-Solving Strategies

Humans solve Connections using a combination of System 1 (fast, intuitive) and System 2 (slow, deliberate) thinking. Effective strategies involve balancing these two modes of thought. Failures can occur due to over-reliance on either system.

Mathematical Analysis of Random Guessing

The probability of winning Connections by random guessing is extremely low. With the initial board, the chance is near 0%. After getting one category correct, the chance is about 1 in 5,000 or 6,000. If stuck on the third and fourth categories, random guessing yields only a 3% chance of winning.

Modeling Connections as a Graph Coloring Problem

Connections can be modeled as an augmented graph coloring problem. Each word is a vertex, and the categories are colors. The goal is to color each word node with one of the four categories such that all four words belong to a specific category receive the same color. Edges represent the strength of the connection between words, creating a search space for algorithms.

Semantic Similarity and Relational Alignment

Semantic similarity alone is insufficient for solving Connections. A tree of word relationships exists, including anagrams, morphology, encyclopedic relationships, and associative relationships. Relational alignment, a metric that associates two words, is crucial. A heat map simulation can visualize relational alignment scores between words.

Computational Analysis of Puzzle Difficulty

Relational alignment scores can help determine puzzle difficulty. Easy puzzles tend to have a higher overall coherence in relational alignment compared to hard puzzles. Time-variant relational alignment scores can be tracked across categories and time to identify patterns.

Multi-Dimensional Relational Alignment

Words can be related in multiple ways, leading to multiple relational alignment scores. A radar chart can map the semantic space of a word across different categories. Multi-dimensional relational alignment distribution is a key factor in analyzing how AI can solve puzzles.

Semantic Distribution Evaluation Framework

A semantic distribution evaluation framework can be used to analyze the distribution of categories (hypernomy, morphology, orthography, etc.) in Connections puzzles over time. This allows for the identification of trends and patterns.

Graph Clustering and Hypergraphs

Adding semantic relationships to the graph coloring problem leads to multi-dimensional hypergraphs. These hypergraphs represent intercluster and intracluster strengths between different nodes and categories. This provides more dimensionality for AI to solve the problem.

Building Semantic Graphs with Lexical Databases

Semantic graphs can be built using lexical databases like WordNet and ConceptNet, as well as word embeddings. These resources provide information about word relationships and allow for the creation of complex conceptual semantic graphs.

Graph Neural Networks and Reinforcement Learning

A graph convolutional neural network (GNN) can be used to take in a graph as an input and output candidate subgraphs that could be solutions. This is combined with a reinforcement learning system to optimize edge and node weights and identify the best candidate solutions.

System Architecture and Visualization

The system combines reinforcement learning agents with a graph-based system using a GNN. The visualization in three-space shows how semantic graphs are constructed and how the reinforcement learning system navigates subclusters.

Results and Future Directions

The solvability rate for a small subset of hard puzzles increases with the described system. Future steps include extending the system to more puzzles and connecting it to the ArcGI benchmark.

Conclusion

The presentation explores the application of AI, particularly graph neural networks and reinforcement learning, to solve the New York Times' Connections game. It emphasizes the importance of semantic relationships, relational alignment, and multi-dimensional analysis in creating an effective AI solver. The research aims to develop a transparent and explainable AI approach, addressing the limitations of LLMs in this context.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.