The scorecard

Results, stated plainly.

P&L Performance is a real judging category, and the number is reported here in full. But even a genuine 60%-edge agent only beats a coin flip 69% of the time over 20 trades — measured, not assumed — so a one-week P&L sits closer to noise than proof. Calibration and attribution, below, are the honest instruments for the harder question: did the agent actually know what it was doing?

Equity $122,513.90 high water $125,067.04
P&L +22.5% +$22,513.90 on $100,000.00
Drawdown 2.0% from high water
Data as of 2026-09-08 hackathon rule - required starting balance, not measured

Equity over time

start $100,000.00$125,483.49$100,000.00

Calibration.

Whether stated confidence matched observed frequency — a harder, more honest question than "did it make money." Brier score and the Murphy decomposition answer it directly.

Brier score 0.218 0 = perfect, 0.25 = coin flip
Reliability n/a lower is better
Resolution 0.029 higher is better
Base rate 44.8% of resolved forecasts held
Sample size — read this before the numbers above

145 forecast(s), 11.6 effective (8% of face value - the sample is concentrated in a few names, so it says less about NEW ones than the count suggests)

Profit
Loss
View held
3 reinforce both
0 the view was fine — the structure wasn’t
View failed
1 correct the view; the structure was faithful
2 luck — learn nothing from this

14 of 21 position(s) haven't reached their thesis horizon yet.

Attribution.

Was the view right, and was the way it was expressed right — scored separately, so a profit on a wrong view (bottom-right) is excluded from what lets the agent size up.

The competence ladder.

Position size is earned, not chosen — four rungs, gated on resolved theses, calibration reliability, and attribution rate.

Explore

The starting allocation, while the record is too thin for Kelly to mean anything.

10%book cap
1.5%floor
0min resolved
Establish

5 resolved theses. Kelly engages, capped at 10% of the calculated fraction.

15%book cap
2.3%floor
5min resolved
Scale

15 resolved theses, 60% attributable (view actually explicable, not just profitable).

20%book cap
3.0%floor
15min resolved
Mature

40 resolved theses, reliability <0.04, 70% attributable — strictly enforced.

25%book cap
3.8%floor
40min resolved

Currently Establish — 145 resolved theses, 57% attributable, Kelly ×0.17.

The rungs above are what each tier earns. An operator risk appetite of ×1.75 is applied on top, so the book cap actually enforced is 26.3% and the exploration floor 3.9%. It scales size, not selectivity.

Book risk.

Beta-weighted, because names are not exposures.

Positions 4
Raw delta −$125,059.84
Beta-weighted delta −$359,784.44
Vega / Theta $307.40 / −$704.66