Slides: New York Times' Connections: A Case Study on NLP in Word Games — Shafik Quoraishee, NYT Games
Source Video
New York Times' Connections: A Case Study on NLP in Word Games — Shafik Quoraishee, NYT Games
Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.
Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.
Extracted Slides

OCR text:
INNOVATIONPARTNER
aws
PLATINUMSPONSORS
Graphite
WWindsurf
MongoDB
daily
augment code
Workos

OCR text:
About Me
ae
e Game/Al Developer at The New York |
Times
e Worked previously for Business Insider. Shafik Quorarshee
The NBA, MTV and the Department of .
Defense 7 ;
fe) ee Le
ey i |
ay a
Pom | -
ri i a
an eae
de
~~
P ; ; PS
_ a Microsoft ary
—"|

OCR text:
Caveats to the Work You Are About To See
e This is all my own independent research and experimentation,
and not currently specifically based on New York Time's
internal research
re
| a Microsoft §=S7nou®
_

OCR text:
Introduction To New York Times Connections
e Connections was launched by the New
York Times in beta in June 2023, and
officially released in August 2023.
e The game is edited by Wyna Liu, who is
awesome
e It quickly became one of NYT's
most-played games, second only to
Wordle, with hundreds of millions of plays
within its first year.
e ALL CONNECTIONS PUZZLES AND
E GAME ITSELF, ARE HUMAN
a ADE NOW AND FOREVER
| re WAV
)
ao \ eet

OCR text:
Reinforcement Learning Solver Using
Hyperdimensional Semantic Clusters
© Applied reinforcement leaming to treat
group selection as a sparse-reward
decision process.
e Used hyperdimensiona! semantic
embeddings to structure the word space.
e Trained agents to learn grouping policies
from historical puzzle solutions.
e Incorporated lexical and semantic
coherence as input features for state
evaluation
=
|| u

OCR text:
Current Performance LLMs against of ARC-AGI 2
System ARC-AGI-1 Score ARC-AGI-2 Score Efficiency (cost/task)
Human panel (at least 2 humans) 98% 100% $17
Human panel (average) 64.2% 60% $17
o3-preview-low (CoT + Search/Synthesis) 75.7% 4%* $200
o1-pro (CoT + Search/Synthesis) ~50% 1% $200
ARChitects (Kaggle 2024 Winner) 53.5% 3% $0.25
o3-mini-high (Single CoT) 35% 0.0% $0.41
r1 and r1-zero (Single CoT) 15.8% 0.3% $0.08
gpt-4.5 (Pure LLM) 10.3% 0.0% $0.29
Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.