RL Selects the Response, Not the Facts
State:
1. Failure category
2. Risk level
3. Retry count
4. Drift severity
5. Data-quality condition
Action:
retry • coerce • rollback • quarantine • escalate • log
1. Tabular Q-learning
2. Small, interpretable state space
3. Low-memory inference
4. Inspectable Q-values for every decision
TECHNICALLY, THIS IS A SINGLE-STEP CONTEXTUAL DECISION PROBLEM IMPLEMENTED WITH TABULAR Q-LEARNING.