Results, stated plainly.
P&L Performance is a real judging category, and the number is reported here in full. But even a genuine 60%-edge agent only beats a coin flip 69% of the time over 20 trades — measured, not assumed — so a one-week P&L sits closer to noise than proof. Calibration and attribution, below, are the honest instruments for the harder question: did the agent actually know what it was doing?
Equity over time
Calibration.
Whether stated confidence matched observed frequency — a harder, more honest question than "did it make money." Brier score and the Murphy decomposition answer it directly.
145 forecast(s), 11.6 effective (8% of face value - the sample is concentrated in a few names, so it says less about NEW ones than the count suggests)
14 of 21 position(s) haven't reached their thesis horizon yet.
Attribution.
Was the view right, and was the way it was expressed right — scored separately, so a profit on a wrong view (bottom-right) is excluded from what lets the agent size up.
The competence ladder.
Position size is earned, not chosen — four rungs, gated on resolved theses, calibration reliability, and attribution rate.
The starting allocation, while the record is too thin for Kelly to mean anything.
5 resolved theses. Kelly engages, capped at 10% of the calculated fraction.
15 resolved theses, 60% attributable (view actually explicable, not just profitable).
40 resolved theses, reliability <0.04, 70% attributable — strictly enforced.
Currently Establish — 145 resolved theses, 57% attributable, Kelly ×0.17.
The rungs above are what each tier earns. An operator risk appetite of ×1.75 is applied on top, so the book cap actually enforced is 26.3% and the exploration floor 3.9%. It scales size, not selectivity.
Book risk.
Beta-weighted, because names are not exposures.