SO, LET'S SUPPORT 85+ MODELS WITH 35+ RUNTIMES AND INFERENCE BACKENDS
Every architecture is different under the hood
bert_flash
qwen2_flash
modernbert_flash
colbert
cross_encoder
clip_siglip_florence2
splade_flash
sglang
No universal engine
An agentic workflow
Re-implementation
15 of 35 adapters use Flash Attention 2 with variable-length sequence packing
AI text/layout recreation from video frame; verify against source image.