Why are we trying to have a single diver do everything?
AI text/layout recreation from video frame; verify against source image.