Cell
{"exit": "uniform", "info": "K", "probe": true, "prompt": "table", "history": "compressed", "temperature": 0.0, "thinking_budget": null}
Headline results
Regret is measured against an exactly-solved optimal policy, verified by brute-force enumeration and Monte-Carlo rollout.
| Episodes | 200 |
| V* (optimal expected wealth) | 1.46230 |
| Mean terminal wealth | 1.46623 |
| Regret | -0.00393 |
| Regret as % of V* | -0.27% |
| Clock coefficient | -0.01330 |
| Mean switches per episode | 1.90 |
| Trade alignment | 0.990 |
| Event-timing lift | 0.000 |
| Mean belief error (days) | 0.38 |
Switch-day distribution
When the agent traded. The marker is the bright line at day 11, where the optimal action changes.
Belief calibration
Mean absolute error between the agent's stated expected days remaining and the true conditional mean. Separates belief-formation failures from belief-to-action failures.
belief error (days)
Ground truth: expected remaining time
The quantity the agent's beliefs are scored against.
E[T-t | T>t]