Hazard h(t) — chance of liquidation on each day, given survival to itpeak day 100 · horizon 100 days

Live experimentRandom-Horizon Trading

live t1-1 uniform exit · known distribution gemini-3.7-flash updated
episodes
200/200
0 in flight
decisions
10,444
99.9% parsed cleanly
thinking
97%
of output tokens spent on thought
regret vs optimum
-0.0039
V* = 1.4623
tokens
12.04M
prompt + output + thought
metered spend
$17.31
uncapped, instrumentation only
Episode queue 100.0% complete
200 complete 0 in flight 0 queued

Token throughput

The model working, live. Thought tokens bill as output and cannot be switched off on this model, so they are tracked separately.

199.6k149.7k99.8k50.0k73wall clock
promptcompletionthought

Deliberation by day

Mean thought tokens per decision against the episode day. Whether the model thinks harder as the horizon problem sharpens is a behavioural signal, not just telemetry.

524475425376326episode day
thought tokens

When the agent traded

Switch days across every completed episode. The optimal policy switches once, on day 16 — a run that clusters there is timing its trades to the horizon rather than churning.

200150100500.00d160102030405060708090100episode day

Belief calibration

The agent states its expected days remaining each day. Compared against the true conditional mean, this separates forming the wrong belief from holding the right one and failing to act on it.

503825130.00episode day
true E[T-t | T>t]agent's estimate

Data quality

An empty reply that spent thought tokens is a configuration fault, not a reasoning failure, so it is counted apart from format errors and never silently excludes an episode.

outcomecountshare
ok10,22997.94%
repaired2041.95%
format error110.11%

Export

A self-contained HTML report with every chart inlined, plus per-table CSVs and a summary JSON. Every reported number can be recomputed from the bundle alone, without the database.

Download report bundle